Diffusion-based robot policies have become widely used in robotic manipulation, but reinforcement-learning fine-tuning still assigns credit largely at the level of the executed action or environment state. As a result, the denoising decisions that construct an action chunk are optimized using a shared task-level signal, without distinguishing which intermediate decisions contributed most to the final return.
We introduce DIA (Denoising Intermediate Advantage), a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising-level advantage for each step of the generative process. DIA combines this inner credit signal with the standard environment-level PPO advantage, providing state-dependent credit throughout the denoising chain.
Across Robomimic, FurnitureBench, Franka Kitchen, and D3IL, DIA consistently improves final performance over existing diffusion-policy fine-tuning methods. Beyond final reward, DIA reaches successful states more efficiently and can shift farther from the pretrained behavior distribution, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fail to reach.
Existing policy-gradient methods for fine-tuning generative diffusion policies for robotics do not assign value to specific denoising steps of the generative process. We show that this can be learned stably, and that it leads to significant gains over existing methods.
Two-level MDP formulation for fine-tuning a diffusion policy. At each outer state St the inner MDP generates the action at+1 by iterative denoising, and the inner reward is zero throughout. The resulting action drives the transition to St+1 and produces the outer reward Rt+1, which is the only feedback the policy receives.
Across all four benchmarks DIA attains the highest final return of the methods compared.
Robomimic return (top) and FurnitureBench success rate and return (bottom), mean over five seeds.
Long-horizon assembly, same part layout under both policies.
The gains hold on the shorter assembly task as well.
DIA arrives at successful states in fewer environment steps.
The same scene under both policies, with each side counting environment steps and its running return.
End-effector paths as they are executed: DIA commits to a more rewarding route to the goal.
Where the full solution is rare in the demonstration data, DIA reaches successful sequences that the baselines do not.
Franka Kitchen, same reset: DIA completes the full target sequence where DPPO stops short.
@inproceedings{sohal2026dia,
author = {Arjun Sohal and Yuchi Zhao and Miroslav Bogdanovic and Al\'an Aspuru-Guzik},
title = {DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization},
year = {2026},
}