The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
Reinforcement LearningSupervised Fine-TuningCatastrophic ForgettingGRPO AlgorithmLLM-as-JudgeModel DistillationCompute ConstraintsRecursive Self-ImprovementRL EnvironmentsReward HackingLoRA AdaptersContinuous LearningModel LatencyInference CostMetacognitive Behaviors
Kyle Corbitt, founder of OpenPipe and leader of CoreWeave's serverless training team, provides a master class on reinforcement learning (RL) and custom fine-tuning for AI models. He explains how RL differs from supervised fine-tuning (SFT) by making less destructive weight updates, delves into the GRPO algorithm and its industrial improvements, and discusses the role of LLMs as judges in post-training. The conversation also covers reward hacking, the economics of RL environments, and the competitive landscape, highlighting compute as the primary constraint for Chinese labs and the potential for recursive self-improvement.