"Rethinking On-Policy Distillation of Large Language Models II: One Training Example" by Zixuan Fu , Bingxiang He , Yuxin Zuo , Haohuan Huang , Jinqian Zhang , Ruhang Xiao , Cheng Qian , Qinyu Luo , Huan-ang Gao , Yudong Wang , Zhiyuan Liu , Ning Ding , Chaojun Xiao
TLDR:
The text discusses how on-policy distillation (OPD) improves rapidly from a single query, covering teacher states quickly but showing slow student alignment. This suggests that OPD is constrained more by algorithm limitations than by lack of data. The study examines OPD's behavior at the data-minimal limit with just one query, finding that one-shot OPD steadily improves and closely matches full-data OPD performance across various tasks and model types. The focus is on the states visited during training and the rate of student-teacher alignment. State coverage is measured, showing that a single query already reaches a high percentage of the states visited by full-data OPD. Adding more diverse queries further increases coverage and validation accuracy. While alignment slows down regardless of training on one query or the full dataset, the absorption of states takes time, indicating that OPD is data-overfed but algorithm-starved. The study suggests that OPD's success lies in its ability to quickly expose broad supervision, with future research encouraged to explore OPD's efficiency and the mechanisms driving its recent achievements.
Read Paper / Blog
1September 4, 2026 51