Artificial intelligence
Vision-language model
Definition
A vision-language model processes visual information and natural language in a shared system. Depending on its design, it may connect images with text representations or generate text from visual and textual inputs.
Also known as: VLM, Vision language model
Updated
The output depends on the architecture
CLIP learns image and text representations by identifying which captions belong with which images. It can compare a visual observation with descriptions without being a conversational text generator.
PaLI instead generates text from visual and textual inputs, supporting tasks such as image captioning and visual question answering. Both are vision-language models, so the label alone does not specify the output interface.
Visual semantics can help robot tasks
A robot may need to distinguish the object described as a red cup from neighboring objects. A vision-language representation can provide semantic information for that decision.
CLIPort combines CLIP's visual-language features with a spatial manipulation architecture to learn language-conditioned picking and placing. The robot policy adds action-relevant structure beyond the pretrained visual-language representation.
Understanding an image does not execute an action
A vision-language-action model explicitly includes robot action output. A VLM can instead serve as a perception component, task planner, or source of pretrained features in a larger robot system.
Textual accuracy and physical task success therefore require different evaluations. Naming the correct object does not by itself show that the robot can reach it, grasp it, or complete the requested operation.
Sources
Related terms
Vision-language-action model
A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute.
Language-conditioned policy
A language-conditioned policy selects actions using a language instruction together with observations. The instruction specifies or modifies the behavior requested from the policy.
Self-supervised learning
Self-supervised learning builds a training signal from the structure of the data itself rather than requiring a human label for every example. In robotics it can learn useful visual or temporal representations before a downstream task policy is trained.