The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]
AI evaluationScalable oversightBenchmark limitationsData contaminationReward hackingTime Horizon metricHuman baseliningAgentic AIAI capabilitiesAI timelinesSoftware engineering automationAI riskDeceptive alignmentRecursive self-improvementAI monitoring
This episode features Beth Barnes and David Rein from METR discussing their 'Time Horizon' graph, a unified metric for measuring AI progress based on human task completion time. They explain how this benchmark addresses the limitations of traditional evaluations by focusing on real-world task difficulty and generalization, rather than just accuracy. The conversation also delves into the implications of AI capabilities for timelines, the future of software engineering, and the complex issues of AI agency, reward hacking, and self-improvement.