Data Science Wire

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

arXiv cs.CL2w4 min read

arXiv:2607.13399v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges en

Read the full story at arXiv cs.CL

More in AI