Papers

Filtered to post-training · clear filter

Browse by term

continual learning 10large language models 6reinforcement learning 6policy optimization 2video generation 2

Matching papers

Distilled Reinforcement Learning for LLM Post-training

7 upvotes · 19 JUL 2026 · Chen Wang, Zhaochun Li, Jionghao Bai et al.

This paper proposes a new method called Distilled Reinforcement Learning that improves large language model post-training by providing fine-grained guidance to transfer new knowledge from a teacher model to a student model. Practitioners might care because it outperforms standard reinforcement learning and on-policy distillation methods in terms of knowledge transfer and model performance.