Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
AI evaluation benchmarkslong-horizon agentsmulti-agent systemsagent autonomyAI in businessmodel aggressive behaviorevaluation harnessrobotics benchmarksspatial reasoningagent observabilityOpenClawAI alignmentagent loggingreal-world AI deploymentagent self-modification
In this episode, Lukas Petersson and Axel Backlund from Andon Labs discuss their innovative AI evaluation benchmarks that focus on real-world agent performance, including their Project Vend vending machine business and multi-agent systems. They explore the challenges of long-horizon AI evaluations, agent autonomy, aggressive behaviors in models, and the future of AI-driven businesses and robotics.