Data Science Wire

Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

arXiv cs.LG4w4 min read

arXiv:2607.01763v1 Announce Type: new Abstract: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy learning can mitigate forgetting, with on-policy self-distillation emerging as a particularly attractive approach. In this work, we revisit this optimistic view through self-distillation policy optimization (SDPO). Our experiments show that SDPO can accelerate in-domain specialization when teacher signals are stable and well aligned, but it struggles to generalize to out-of-distribution scenarios.

Read the full story at arXiv cs.LG

More in Machine Learning