Artificial intelligence

Open-vocabulary object detection

Definition

Open-vocabulary object detection locates objects in an image and associates them with text-defined categories beyond a fixed closed list used during detector training. It combines spatial detection with vision-language representations that can be queried using class names or phrases.

Also known as: OVD, Open vocabulary object detection

Updated

Text expands the detector's query set

A closed-vocabulary detector predicts boxes and labels from categories fixed by its training dataset. An open-vocabulary detector uses language-aligned visual features so that a user or robot can request categories through text, including categories not present in the detector's labelled box training set.

OWL-ViT transfers image-text representations to detection and accepts text queries for target categories. Grounding DINO combines a Transformer detector with grounded language pretraining and reports evaluation as an open-set detector. The papers use related terminology, and individual benchmarks define differently which categories count as unseen.

A robot can ask for task-relevant objects

A general-purpose robot may be instructed to find a particular tool, package, or household object. Open-vocabulary detection can propose image regions associated with that phrase without adding a dedicated output class for every possible instruction. The detections can guide active perception, tracking, or a manipulation pipeline.

The output is usually a two-dimensional region and a text match score. Grasping still needs depth, pose estimation, segmentation or geometry, reachability checks, and a controller. Language-conditioned detection is one perception component, not an action policy.

An open vocabulary is not unlimited recognition

Results depend on the image-text pretraining data, detector training, prompt wording, image quality, viewpoint, and decision threshold. Rare objects, fine-grained distinctions, occlusion, and unfamiliar environments can cause missed or incorrect detections. A model may also match visual context or text correlations instead of the physical attribute relevant to the task.

"Open" therefore describes how categories are queried and evaluated, not a guarantee that the detector understands every phrase. A robot should preserve uncertainty and verify critical detections through additional views, sensors, or task-specific checks.

Sources