Whatever happened to BABA is AI from 2024? [D]
A 2024 paper presented at the ICML conference by MIT and Virginia Tech researchers found that state-of-the-art multi-modal large language models like GPT-4o, Gemini-1.5-Pro, and Gemini-1.5-Flash "fail dramatically" when generalization requires manipulating and combining game rules. This research, potentially important for benchmarks like ARC-AGI-4, suggests current LLMs struggle with complex, rule-based puzzles, though some believe agentic swarms could solve them.
This report highlights a specific failure mode for state-of-the-art LLMs like GPT-4o and Gemini 1.5, unlike other benchmarks that focus on general capabilities.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 9, 2026, 01:00 UTC
- Ingested
- Oct 9, 2026, 01:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →