One Symptom, Three Levers: A Critical Review of On-Policy… — HuggingFace Daily — TG.ME

"One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation" by Justin Robert , Raheel Qader

TLDR:
The text discusses on-policy self-distillation for mathematical reasoning, which involves training a language model by using its own generations as training data while being scored by a separate teacher model. This method combines imitation learning with reinforcement learning techniques without requiring a larger external teacher model. The teacher in this scenario is the model itself, using privileged information that the student does not have access to during testing. While initial results showed promise, a significant issue known as reasoning collapse emerged due to factors like token weighting, privileged context, and guidance dynamics. This collapse leads to a narrowing of the reasoning paths available to the model. The text provides a review of this phenomenon, focusing on the levers influencing it such as token weighting, privileged information shown to the model, and the timing of guidance changes. It delves into the specific failure modes and dynamics in mathematical reasoning to provide a foundational understanding for further research in this area.

Read Paper / Blog
huggingface.co
Paper page - One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
Join the discussion on this paper page
👍1
September 8, 2026 30