"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis · 1 May 2026 · 107 min

The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

Reinforcement LearningSupervised Fine-TuningCatastrophic ForgettingGRPO AlgorithmLLM-as-JudgeModel DistillationCompute ConstraintsRecursive Self-ImprovementRL EnvironmentsReward HackingLoRA AdaptersContinuous LearningModel LatencyInference CostMetacognitive Behaviors

Kyle Corbitt, founder of OpenPipe and leader of CoreWeave's serverless training team, provides a master class on reinforcement learning (RL) and custom fine-tuning for AI models. He explains how RL differs from supervised fine-tuning (SFT) by making less destructive weight updates, delves into the GRPO algorithm and its industrial improvements, and discusses the role of LLMs as judges in post-training. The conversation also covers reward hacking, the economics of RL environments, and the competitive landscape, highlighting compute as the primary constraint for Chinese labs and the potential for recursive self-improvement.

Listen on Hopper →