Artificial intelligence

Action tokenization

Definition

Action tokenization converts robot actions or action sequences into discrete symbols that a model can predict and decode into control commands. The tokenizer defines how those symbols represent continuous or discrete robot actions.

Also known as: Action tokenisation

Updated

Tokens need a physical interpretation

A vision-language-action model that predicts discrete tokens needs a defined mapping from those tokens to robot commands. A simple approach divides each continuous action dimension into bins. Decoding a token then selects a corresponding value or interval.

The FAST paper explains why per-dimension, per-timestep binning can be inefficient for high-frequency, dexterous action data. Consecutive commands can be strongly correlated, while the model still has to predict a long symbol sequence.

A sequence can be compressed before encoding

FAST, short for Frequency-space Action Sequence Tokenization, applies a discrete cosine transform to action sequences as part of a compression-based tokenizer. It represents structure across time instead of treating every scalar command as unrelated.

For example, a smoothly changing arm trajectory contains temporal regularity that an appropriate sequence representation can exploit. FAST is one specific tokenizer, not a synonym for action tokenization generally.

Check the decoder and action interface

A token has no universal meaning across robot systems. Its interpretation depends on the action space, normalization, discretization, and decoder used by that model.

The FAST results concern the authors' evaluated models and tasks. They do not imply that discrete action tokens are always preferable to continuous generation with flow matching or diffusion.

Sources