Recursive Self-Distillation Doubles Qwen3-8B Math Accuracy to 65.97%
A new arXiv paper proposes DCE+SRCL, letting the teacher model co-evolve with its student to lift Qwen3-8B math accuracy by 35 points.
- Paper proposes DCE+SRCL, a recursive on-policy self-distillation method for reasoning models.
- Dynamic Co-Evolution unfreezes the privileged teacher so student gains feed back into supervision.
- Self-Refined Concise Learning trains on verified shorter rewrites to curb verbose self-criticism.
- On Qwen3-8B, reaches 65.97% Average@12, beating OPSD by 35.62 points.
- Cuts mean output length by 7.80% relative to DCE alone across four math benchmarks.
- Authors show frozen teachers bias students toward terminating instead of reflecting on errors.
Recursive self-distillation lifts Qwen3-8B math accuracy from 30.35% to 65.97%
A new arXiv preprint argues that a frozen teacher limits self-distillation because it cannot absorb improvements made by the student. Shangjian Yin and co-authors propose periodically updating the teacher, then training on shorter verified rewrites to control verbosity. On Qwen3-8B, the combined method raises reported competition-math Average@12 accuracy by 35.62 percentage points over the previous self-distillation baseline.
The arXiv preprint, Recursive Self-Improvement via On-Policy Distillation for Reasoning, presents a post-training method for improving reasoning models without a larger external teacher. The results are author-reported and await peer review and independent replication.
Why the frozen teacher goes stale
On-policy distillation starts with responses sampled from the student model. A teacher supplies next-token probability targets for those responses, and the student learns to match the teacher’s distribution at each token. This dense signal provides guidance throughout a solution instead of assigning one scalar reward after the complete response.
On-policy self-distillation, or OPSD, uses a second copy of the student as the teacher. The student receives the problem, while the teacher also receives the ground-truth answer in its context. Both copies begin with the same weights, and the solution context gives the teacher its advantage.
Previous implementations keep the privileged teacher frozen to stabilize training. The authors argue that this fixed target becomes stale as the student develops better reflection and revision behavior. In their probes on evolving trajectories, the frozen teacher increasingly favors ending the sequence over reconsidering or correcting an answer, passing that termination bias back to the student.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.