Long-horizon planning is a critical challenge for AI agents, as initial assumptions are frequently invalidated by environmental changes or tool failures. New research introduces Plan-and-Patch, a framework utilizing diffusion language models (dLLMs) to generate and repair structured, program-like action plans.
Diffusion vs. Autoregressive Planning
Unlike traditional autoregressive (AR) models that often require regenerating an entire plan when an error occurs, Plan-and-Patch enables selective repair. By using parallel unmasking, the dLLM fills in affected regions while keeping the surrounding prefix and suffix steps fixed. This localized repair prevents unnecessary changes to functional parts of the plan.
On the Natural Plan benchmark, the diffusion-based DreamReasoner-8B achieved a plan repair success rate of 53.7%, nearly doubling the 27.0% success rate of the autoregressive Qwen3-8B. This suggests that diffusion models are inherently better suited for iterative refinement in agentic tasks.
Performance Gains and Latency
Beyond repair effectiveness, the framework offers significant efficiency improvements. Following task-specific training on benchmarks like ALFWorld and TextCraft, the study observed the following:
- Diffusion models reduced mean plan-generation latency by 39-46% relative to AR models.
- Observed success rates in initial plan generation remained comparable between the two architectures.
The results indicate that Plan-and-Patch provides a robust foundation for agents requiring both speed and the ability to adapt to unexpected outcomes without the overhead of full plan regeneration.