Artificial intelligence
State hallucination
Definition
State hallucination is a reported robot-policy failure in which the policy continues acting as though a required robot-object state has occurred even though it has not. The term has been used for vision-language-action policies that fail to re-ground their action sequence after an unrealized transition.
Also known as: Policy state hallucination
Updated
Act on a transition that did not happen
Suppose a gripper misses an object but the policy proceeds with the motion intended to place that object in a container. The later action is consistent with an imagined successful grasp rather than the observed physical state. Lee and colleagues call this recurring VLA failure pattern state hallucination.
The term is recent and is not yet a standardized robotics failure taxonomy. It should be tied to an operational test, such as whether a required contact or object transition occurred before the next action, rather than applied to any failed manipulation.
Different from an incorrect state estimate
A state estimator can output the wrong pose because of noise or occlusion. State hallucination, as used in the cited work, describes the policy behavior of continuing as if an unrealized state had been achieved. A bad estimate can cause that behavior, but the policy may also ignore visible contrary evidence or continue an open-loop action sequence.
It also differs from a world model producing an inaccurate future prediction. The observable failure is that the action policy does not adjust to the unrealized transition, regardless of whether it has an explicit predictive model.
Diagnosis needs closed-loop evidence
The cited authors report weaker attention to task-relevant image regions during these failures and identify associated sparse features in tested VLA models. They also report reduced hallucinated failures after selectively unlearning those features. These are findings for the studied architectures, simulations, real-world tasks, and intervention.
A readable action sequence alone cannot establish what internal assumption caused a failure. Contact may be hidden, the success condition may be ambiguous, or latency may make the policy act on an older frame. Evaluation should synchronize observations and actions, define the missing transition, and show whether fresh evidence reaches the policy before labeling the behavior state hallucination.
Sources
Related terms
Vision-language-action model
A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute.
State estimation
State estimation infers quantities describing a robot or its environment from measurements and a model. A robot state may include position, orientation, velocity, and other variables that are not all directly measured.
World model
A world model is an internal predictive model of an environment and how it changes. In robot learning, it can predict future states or observations under possible actions to support planning or policy training.
Action chunking
Action chunking is the prediction or organization of several future robot actions as one sequence. A policy can execute all or part of a chunk before using new observations to produce another sequence.