DIA: Denoising Intermediate Advantage
for Diffusion Policy Optimization

Arjun Sohal1, Yuchi Zhao1,2, Miroslav Bogdanovic1,3, Alán Aspuru-Guzik1,2,3,4
1University of Toronto, 2Vector Institute, 3Acceleration Consortium, 4NVIDIA

DIA across six tasks.

Abstract

Diffusion-based robot policies have become widely used in robotic manipulation, but reinforcement-learning fine-tuning still assigns credit largely at the level of the executed action or environment state. As a result, the denoising decisions that construct an action chunk are optimized using a shared task-level signal, without distinguishing which intermediate decisions contributed most to the final return.

We introduce DIA (Denoising Intermediate Advantage), a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising-level advantage for each step of the generative process. DIA combines this inner credit signal with the standard environment-level PPO advantage, providing state-dependent credit throughout the denoising chain.

Across Robomimic, FurnitureBench, Franka Kitchen, and D3IL, DIA consistently improves final performance over existing diffusion-policy fine-tuning methods. Beyond final reward, DIA reaches successful states more efficiently and can shift farther from the pretrained behavior distribution, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fail to reach.

Motivation

Existing policy-gradient methods for fine-tuning generative diffusion policies for robotics do not assign value to specific denoising steps of the generative process. We show that this can be learned stably, and that it leads to significant gains over existing methods.

Two-level MDP formulation: an inner MDP that generates an action by
                  iterative denoising, nested inside an outer MDP over environment states.

Two-level MDP formulation for fine-tuning a diffusion policy. At each outer state St the inner MDP generates the action at+1 by iterative denoising, and the inner reward is zero throughout. The resulting action drives the transition to St+1 and produces the outer reward Rt+1, which is the only feedback the policy receives.

Results

DIA consistently improves downstream performance across tasks

Across all four benchmarks DIA attains the highest final return of the methods compared.

Fine-tuning curves on Robomimic (top) and FurnitureBench (bottom),
                mean over five seeds.

Robomimic return (top) and FurnitureBench success rate and return (bottom), mean over five seeds.

Long-horizon assembly, same part layout under both policies.

The gains hold on the shorter assembly task as well.

DIA reaches successful states more efficiently

DIA arrives at successful states in fewer environment steps.

The same scene under both policies, with each side counting environment steps and its running return.

End-effector paths as they are executed: DIA commits to a more rewarding route to the goal.

DIA learns task sequences that are weakly represented or absent in the demonstrations

Where the full solution is rare in the demonstration data, DIA reaches successful sequences that the baselines do not.

Franka Kitchen, same reset: DIA completes the full target sequence where DPPO stops short.

BibTeX

@inproceedings{sohal2026dia,
  author    = {Arjun Sohal and Yuchi Zhao and Miroslav Bogdanovic and Al\'an Aspuru-Guzik},
  title     = {DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization},
  year      = {2026},
}