ARC-AGI-3
The first interactive reasoning benchmark designed to measure human-like intelligence in AI agents.
Interactive reasoning benchmark and evaluation environment for AI agents, consisting of novel turn-based game-style environments where agents must explore, infer goals, build world models, plan actions, and adapt without natural-language instructions.

Recent stories
0 linked stories
No linked stories yet.