Artificial intelligence

Diffusion policy

Definition

A diffusion policy generates robot actions through a learned denoising process conditioned on observations. It commonly predicts an action sequence by progressively refining a noisy candidate rather than predicting one action with a single direct regression.

Updated

Denoising produces an action sequence

The Diffusion Policy research models visuomotor behavior as a conditional denoising diffusion process. During training, the model learns to remove noise from demonstrated actions using observations as context. During execution, repeated refinement turns an initial noisy sequence into a candidate robot motion.

The noisy sequence exists inside the model's computation. It is not a command to make the physical robot move randomly.

Multiple valid motions can remain distinct

A robot may be able to push an object around either side of an obstacle. A model trained to average incompatible demonstrations could propose an unsuitable middle path. Diffusion Policy is designed to represent multiple modes of an action distribution and select a coherent sequence, as illustrated in the authors' manipulation experiments.

It also combines action chunking with receding-horizon execution: the controller executes part of a predicted sequence, observes again, and generates another sequence.

Sampling and feedback impose practical limits

Iterative generation takes computation, so sampling settings and observation timing affect the control loop. The paper's demonstrations include pushing, mug flipping, and sauce manipulation. These results support those evaluated setups; diffusion decoding alone does not provide a collision guarantee or establish humanoid walking capability.

Sources