This paper investigates whether giving a language model extra information, such as a worked solution, improves its learning through on-policy self-distillation. A practitioner might care about how to optimize this technique for better performance.
Firehose
Filtered to Papers, tagged “self-distillation” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives