Artificial intelligence

Synthetic training data

Definition

Synthetic training data is data produced computationally for model training rather than collected directly as the corresponding real-world examples. In robotics, it often includes rendered sensor observations, simulated trajectories, and labels available from the simulator.

Also known as: Synthetic data

Updated

Generate examples and their labels

Tremblay and colleagues train object detection using synthetic images with randomized lighting, poses, and textures. Because the scene is generated, its construction can supply training labels without separately hand-annotating every real camera frame.

A robot-perception dataset could similarly render many object arrangements and provide their positions. Tobin and colleagues study simulated visual data for localization and demonstrate the learned detector in real grasping.

Synthetic experience can include actions

Data need not consist only of images. Dynamics-randomization research trains control policies from simulated robot interaction before testing them on a physical arm.

The simulator can provide observations, actions, and outcomes, but those outcomes reflect its model of physics.

More generated data does not remove the reality gap

Synthetic examples can increase variation and reduce some collection costs. Their usefulness still depends on whether they represent features and behavior relevant to the physical task.

Domain randomization deliberately varies simulated properties to support transfer. It does not make simulated evidence equivalent to hardware evaluation. Reports should state which data were generated, how real data were used, and which real tasks were tested before claiming a perception or control improvement.

Sources