The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
ARC-AGI-3 benchmarkgoal acquisitionaction efficiencyexploration vs exploitationlanguage modelsplanning in AIcoding agentsrequirements engineeringcore knowledge priorsabstraction synthesisreinforcement learningAI safetybitter lessonneural guided searchLLM reasoningsoftware engineering AIbenchmark designgeneral intelligence
Tim Scarfe interviews the Tufa Labs ARC-AGI-3 team to dissect their winning approach on the ARC-AGI-3 benchmark, focusing on how their system discovers goals and balances exploration with action efficiency. The episode explores the challenges of goal acquisition, the role of language models in reasoning and planning, and the evolving design of ARC-AGI benchmarks to better test general intelligence.