Dwarkesh Podcast · 19 June 2026 · 12 min

The data black hole at the center of AI

Sample efficiencyData distributionReinforcement LearningSynthetic data generationHuman expert trajectoriesAI training costsScaling lawsWhite-collar automationAI research automationEvolutionary pre-trainingMultimodal data

This episode argues that current AI progress is primarily driven by an immense quantity of high-quality, task-specific data, rather than improvements in sample efficiency. The speaker highlights the vast data requirements of frontier models compared to human learning, likening AI to a "data black hole" built from billions of carefully constructed examples. The discussion also addresses common objections regarding evolutionary pre-training, multimodal data, and scaling laws, concluding that AI's data-intensive approach can still automate white-collar work and potentially AI research itself due to its ability to amortize training costs.

Listen on Hopper →