No Priors: Artificial Intelligence | Technology | Startups · 26 June 2026 · 36 min

Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown

AI model evaluationTest-time computeAI benchmarksModel capability scalingAI safety evaluationsRecursive self-improvementMulti-agent AI systemsAI reasoningLatent AI capabilitiesResearch taste in AI

OpenAI research scientist Noam Brown discusses how traditional AI benchmarks are failing to accurately evaluate modern models due to their increasing reliance on large-scale test-time compute. He argues that model capabilities are now a function of the computational budget allocated during inference, impacting everything from performance metrics to AI safety evaluations and the pace of research. Brown advocates for evaluating models by plotting performance against test-time compute or setting explicit budget limits.

Listen on Hopper →