Artificial intelligence

Self-supervised learning

Definition

Self-supervised learning builds a training signal from the structure of the data itself rather than requiring a human label for every example. In robotics it can learn useful visual or temporal representations before a downstream task policy is trained.

Also known as: SSL

Updated

Bootstrap Your Own Latent, or BYOL, trains an image representation by predicting a target network's representation of another augmented view of the same image. The two related views provide the learning relationship without a separate object-class label for that image.

This is one self-supervised objective, not the definition of every self-supervised method. Other objectives can use relationships over time or between parts of an observation.

Reuse representations for robot learning

R3M pretrains visual representations on human video using a combination of time-contrastive learning, video-language alignment, and a sparsity objective. It then uses the representation as a frozen perception module for downstream robot-policy learning.

R3M's mixture includes language information, so it should not be described as learning solely from unlabeled image similarity. The example shows how several supervision sources can contribute to a reusable representation.

A representation is not a complete behavior

Learning that observations are related does not itself specify the robot's task or motor commands. A downstream imitation-learning or other policy-learning stage may still require task data.

Evaluate the downstream task separately. Strong representation-learning results on images or human video do not automatically demonstrate physical competence on a new robot body.

Sources