Despite their vast capabilities, large language models (LLMs) frequently stumble on complex reasoning tasks. According to new research published on ArXiv, these failures are often not the result of global incompetence, but rather "localized reasoning bugs" occurring during intermediate steps of the process.
The Role of Weak Models in Diagnosis
The study demonstrates that these reasoning bugs are frequently repairable. By inserting a short "patch" generated by a weaker probe model after a specific reasoning prefix, the trajectory of a stronger model can be redirected toward a correct solution. This suggests that the strong model often possesses the underlying knowledge but requires a nudge to stay on track.
However, the researchers found that simply fine-tuning a strong model on these patches or repaired trajectories is unreliable. The useful signal appears to lie not in the intervention text itself, but in how that text reshapes the model's future reasoning distribution.
Introducing Woodpecker Distillation
To capture this signal, the researchers developed Woodpecker Distillation, a weak-to-strong training framework based on contrastive local interventions. The process involves several key steps:
- Contrasting successful and unsuccessful patches from a weak model at the same reasoning prefix.
- Constructing a corrective teacher distribution based on the future token predictions induced by these patches.
- Distilling this corrective signal into the stronger model.
Experiments conducted on mathematical reasoning benchmarks indicate that Woodpecker Distillation consistently improves the performance of strong models, outperforming standard direct imitation baselines.