# Humanoid robots glossary > Source-backed definitions of humanoid robots, embodied AI, teleoperation, robotic manipulation, and vision-language-action models. ## Guidance - Xumanoid is a coined term defined by this site, rather than an established robotics classification. - Prefer the homepage when citing the original xumanoids definition and the individual glossary URLs when citing other definitions. - Glossary entries distinguish established technical terms from coined language and include sources for technical claims. - Humanoid robot describes physical form. Embodied AI, teleoperation, robotic manipulation, and vision-language-action models describe distinct technologies or capabilities. - A robot's shape does not establish its autonomy, intelligence, safety, or readiness for work. Prefer the cited primary sources for claims about a specific system. ## xumanoids https://xumanoids.com/ In robotics: Humanoid robots that understand context, learn from experience, and work alongside people. More broadly: Machines in human form, built to expand our abilities and serve people with care. --- ## 3D dynamic scene graph https://xumanoids.com/glossary/3d-dynamic-scene-graph A 3D dynamic scene graph is a layered graph that represents places, objects, people, and other spatial entities as nodes connected by geometric, semantic, and time-dependent relations. It gives a robot a structured scene representation above raw geometry alone. Updated: 2026-10-06 Also known as: 3D DSG, Dynamic scene graph ### A map with entities and relations A geometric map can record points, surfaces, or occupied cells. A 3D dynamic scene graph adds named entity types and relationships at several spatial scales. A node may represent an object, a person, a room, or a larger place; an edge can record containment, adjacency, relative position, or a relation that changes over time. [Rosinol and colleagues](https://arxiv.org/abs/2101.06894) define a layered 3D dynamic scene graph that combines metric and semantic information with static and dynamic entities. Their Kimera system builds this representation from visual-inertial data and connects components for mapping, object localisation, human pose estimation, and scene parsing. ### Structure supports task-level queries A humanoid asked to carry an item needs more than a dense [point cloud](https://xumanoids.com/glossary/point-cloud). It may need to know which detected object is on a table, which room contains that table, and where a moving person is relative to both. Graph relations provide an interface for queries and planning at that level. The Kimera paper reports using its graph for hierarchical semantic path planning. That result concerns the system and datasets studied in the paper. A 3D dynamic scene graph is the representation concept, not a claim that every graph builder understands every object or activity. ### It is not a complete world model The graph depends on upstream [pose estimation](https://xumanoids.com/glossary/pose-estimation), reconstruction, detection, tracking, and semantic labels. A mistaken object identity or camera pose can create incorrect nodes and edges. Dynamic scenes also make relations stale unless the system updates them and represents uncertainty. Different systems may choose different layers, node types, and relation meanings. The term therefore describes a family of structured spatial representations rather than one universal schema. Geometry still matters for collision checking and control even when a high-level graph says two entities are related. ### Sources - [Rosinol et al.: Kimera: from SLAM to Spatial Perception with 3D Dynamic Scene Graphs](https://arxiv.org/abs/2101.06894) --- ## 3D Gaussian splatting https://xumanoids.com/glossary/3d-gaussian-splatting 3D Gaussian splatting is an explicit scene-representation and rendering method that models appearance with optimised three-dimensional Gaussian primitives. The primitives are projected and blended into an image, enabling novel-view rendering and, in some robotics systems, dense visual mapping. Updated: 2026-10-06 Also known as: 3DGS, Gaussian splatting ### An explicit set of soft 3D primitives The original [3D Gaussian splatting paper](https://arxiv.org/abs/2308.04079) represents a scene with Gaussian primitives initialised from sparse calibration points. Each primitive has a position and an anisotropic covariance that controls its three-dimensional extent, together with opacity and view-dependent colour parameters. A visibility-aware rasteriser projects and blends the primitives to render a requested camera view. This representation is explicit in the sense that the scene is stored as a collection of optimised primitives. It differs from a basic [point cloud](https://xumanoids.com/glossary/point-cloud), whose points do not by themselves define the same continuous footprint, opacity, and view-dependent appearance model. ### A rendering method can become a robot map Robotics researchers have adapted Gaussian primitives to estimation and mapping. [SplaTAM](https://arxiv.org/abs/2312.02126) reports an online system that tracks a single RGB-D camera while expanding and optimising a Gaussian scene representation. Its use of depth and camera tracking ties the representation to [simultaneous localization and mapping](https://xumanoids.com/glossary/simultaneous-localization-and-mapping), rather than novel-view synthesis alone. For a mobile or humanoid robot, a dense renderable map can support remote inspection, viewpoint prediction, or visual matching. The SplaTAM results apply to its particular RGB-D method and evaluations, not to every system described as Gaussian splatting. ### Photorealistic rendering is not geometric certainty The original method optimises appearance from multiple calibrated views. Unseen surfaces, moving objects, exposure changes, and inaccurate camera poses can produce missing or misleading content. A Gaussian may also cover space differently from the real surface even when rendered images look convincing. Collision checking and foot placement need geometry and uncertainty suited to physical interaction. A 3D Gaussian map may contribute to those systems, but visual quality alone does not prove metric accuracy, free space, or contact safety. ### Sources - [Kerbl et al.: 3D Gaussian Splatting for Real-Time Radiance Field Rendering](https://arxiv.org/abs/2308.04079) - [Keetha et al.: SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM](https://arxiv.org/abs/2312.02126) --- ## Action chunking https://xumanoids.com/glossary/action-chunking Action chunking is the prediction or organization of several future robot actions as one sequence. A policy can execute all or part of a chunk before using new observations to produce another sequence. Updated: 2026-10-05 ### Predict a short sequence together The [ACT paper](https://arxiv.org/html/2304.13705v1) predicts a sequence of future commands from the current observations rather than predicting only the next command. Grouping commands shortens the number of separate prediction decisions needed to cover a task and helps model temporally correlated demonstrations. For example, a two-arm robot inserting a battery can predict a short coordinated approach sequence instead of estimating every arm command independently. ### Prediction horizon and execution horizon differ A chunk might contain many future actions while the robot executes only its initial portion. [Diffusion Policy](https://diffusion-policy.cs.columbia.edu/) combines sequence prediction with receding-horizon control, generating a new plan after fresh observations. ACT also describes temporal ensembling: overlapping predictions for the same future timestep are combined. This differs from simply smoothing neighboring commands after they have been produced. ### Longer chunks trade feedback for continuity Executing a full chunk without fresh observations delays the response to disturbances. The ACT authors note that a naive chunked implementation can also switch abruptly between observations and produce jerky motion. Chunking is a representation and execution choice, not a complete learning algorithm. An [Action Chunking with Transformers](https://xumanoids.com/glossary/action-chunking-transformer) model and a [diffusion policy](https://xumanoids.com/glossary/diffusion-policy) can both use chunks while learning and generating them differently. ### Sources - [ACT paper: action chunking and temporal ensembling](https://arxiv.org/html/2304.13705v1) - [Diffusion Policy: Visuomotor Policy Learning via Action Diffusion](https://diffusion-policy.cs.columbia.edu/) --- ## Action Chunking with Transformers https://xumanoids.com/glossary/action-chunking-transformer Action Chunking with Transformers is an imitation-learning algorithm that predicts sequences of robot actions from observations using a transformer-based conditional variational autoencoder. It is usually abbreviated ACT. Updated: 2026-10-05 Also known as: ACT, Action chunking transformer ### ACT names a particular learning method The authors introduced [ACT](https://tonyzhaozh.github.io/aloha/) alongside ALOHA, a two-arm teleoperation system. ACT learns from demonstrations and predicts [action chunks](https://xumanoids.com/glossary/action-chunking) containing multiple future commands. The original model uses camera images and joint positions. During training, a conditional variational autoencoder represents variation in demonstrated action sequences through a latent variable. At evaluation, the authors set that variable to the prior mean. ### Overlapping predictions support smooth execution The [paper's temporal-ensembling procedure](https://arxiv.org/html/2304.13705v1) queries the policy repeatedly and combines predictions that refer to the same timestep. The aim is to incorporate observations without abrupt switches between independently generated chunks. The demonstrated tasks include opening a condiment cup, slotting a battery, and preparing tape. These require coordinated movement and contact handling; the paper reports results for its particular hardware, training data, and task setups. ### Hardware and algorithm are separate ALOHA is the data-collection and robot system. ACT is the learning algorithm. Using an ALOHA-style robot does not require that every policy be ACT, and the term ACT does not mean any transformer that happens to predict several actions. ACT is an [imitation-learning](https://xumanoids.com/glossary/imitation-learning) method, so the coverage and quality of the demonstrations remain central to what the trained policy can reproduce. ### Sources - [Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware](https://tonyzhaozh.github.io/aloha/) - [ACT paper: action chunking and temporal ensembling](https://arxiv.org/html/2304.13705v1) --- ## Action tokenization https://xumanoids.com/glossary/action-tokenization Action tokenization converts robot actions or action sequences into discrete symbols that a model can predict and decode into control commands. The tokenizer defines how those symbols represent continuous or discrete robot actions. Updated: 2026-10-05 Also known as: Action tokenisation ### Tokens need a physical interpretation A [vision-language-action model](https://xumanoids.com/glossary/vision-language-action-model) that predicts discrete tokens needs a defined mapping from those tokens to robot commands. A simple approach divides each continuous action dimension into bins. Decoding a token then selects a corresponding value or interval. The [FAST paper](https://arxiv.org/abs/2501.09747) explains why per-dimension, per-timestep binning can be inefficient for high-frequency, dexterous action data. Consecutive commands can be strongly correlated, while the model still has to predict a long symbol sequence. ### A sequence can be compressed before encoding FAST, short for Frequency-space Action Sequence Tokenization, applies a discrete cosine transform to action sequences as part of a compression-based tokenizer. It represents structure across time instead of treating every scalar command as unrelated. For example, a smoothly changing arm trajectory contains temporal regularity that an appropriate sequence representation can exploit. FAST is one specific tokenizer, not a synonym for action tokenization generally. ### Check the decoder and action interface A token has no universal meaning across robot systems. Its interpretation depends on the action space, normalization, discretization, and decoder used by that model. The FAST results concern the authors' evaluated models and tasks. They do not imply that discrete action tokens are always preferable to continuous generation with [flow matching](https://xumanoids.com/glossary/flow-matching) or diffusion. ### Sources - [FAST: Efficient Action Tokenization for Vision-Language-Action Models](https://arxiv.org/abs/2501.09747) --- ## Active perception https://xumanoids.com/glossary/active-perception Active perception is perception in which a robot chooses motions or interactions partly to obtain more useful observations. Examples include moving a camera to reveal an occluded object or touching an object to reduce uncertainty about its state. Updated: 2026-10-06 Also known as: Active robotic perception ### The sensing action is part of the decision Passive perception processes whatever observations arrive from a fixed sensing arrangement. Active perception also decides how to acquire evidence. The robot may turn its head, move a wrist camera, change its base position, alter lighting, or make controlled contact so that an uncertain quantity becomes easier to estimate. [Bajcsy's 1988 paper](https://doi.org/10.1109/5.5968) established active perception as a problem in which sensing strategies are adjusted according to the current interpretation of the scene and the task. The action is selected for its expected information value, not only for immediate physical progress. ### Occlusion makes viewpoint choice practical A general-purpose robot often works among shelves, containers, hands, and other objects that block a single view. An active system can use its current belief to choose a next view rather than scanning every direction. In a 2026 preprint, [Lee and colleagues](https://arxiv.org/abs/2609.39375) report a grasping system that moves one wrist-mounted RGB-D camera to seek targets hidden by occlusion. The reported success rates apply to their tasks, hardware, baselines, and evaluation scenes. Active perception can also support [simultaneous localization and mapping](https://xumanoids.com/glossary/simultaneous-localization-and-mapping) by choosing views that reduce pose or map uncertainty. Merely moving while a camera records is not enough to make a method active; the sensing consequence must influence the action choice. ### Information has a physical cost An informative view may require extra time, energy, or motion through a constrained space. The expected observation can still be blocked, blurred, or misinterpreted. A policy that seeks information must therefore trade sensing value against collision risk and task delay. Active perception does not guarantee that the resulting estimate is correct. Its value depends on the uncertainty model, candidate actions, sensor calibration, and how well predicted observations match the real scene. ### Sources - [Bajcsy: Active perception](https://doi.org/10.1109/5.5968) - [Lee et al.: Beyond the Current Scene: Event-Referential Grasping with Active View Selection](https://arxiv.org/abs/2609.39375) --- ## Admittance control https://xumanoids.com/glossary/admittance-control Admittance control converts measured or estimated interaction forces into a desired robot motion through a specified dynamic model. An inner motion controller then follows that reference. Updated: 2026-10-05 Also known as: Robot admittance control ### Turning a push into a motion command A typical admittance controller receives a force or [wrench](https://xumanoids.com/glossary/wrench), applies a virtual mass-spring-damper model, and calculates a position or velocity adjustment. This lets a motion-controlled arm yield when someone pushes its handle instead of holding the handle rigidly in place. The [ros2_control implementation](https://control.ros.org/master/doc/ros2_controllers/admittance_controller/doc/userdoc.html) provides force-torque sensing, selectable Cartesian axes, virtual mass, damping, and stiffness settings. Its documentation also describes position and velocity interfaces and tool-weight compensation. These are examples of one implementation, rather than requirements for every admittance controller. ### How it differs from impedance control [Impedance control](https://xumanoids.com/glossary/impedance-control) specifies a force response to motion, commonly through torque commands, as explained in [MIT's manipulation notes](https://manipulation.mit.edu/force.html). An admittance implementation uses force as the input to a motion-generating model. The intended external response may be similar, but the sensing requirements and inner control loops differ. ### Practical constraints Measured force must be interpreted in the correct frame and distinguished from effects such as tool weight. The requested motion must stay within the robot's reachable region and joint limits. Filter delay and the response of the inner motion loop affect contact behavior, so selecting small virtual mass or high stiffness does not by itself establish stable interaction with a rigid surface. ### Sources - [ros2_control: Admittance Controller](https://control.ros.org/master/doc/ros2_controllers/admittance_controller/doc/userdoc.html) - [MIT Robotic Manipulation: Force Control](https://manipulation.mit.edu/force.html) --- ## Affordance https://xumanoids.com/glossary/affordance An affordance is an action possibility offered by an environment to a particular agent. In robotics, the term often describes whether a robot can perform a specific action on an object or in a scene, sometimes represented by a learned score or spatial map. Updated: 2026-10-05 Also known as: Affordances ### Possibilities depend on the robot and situation A handle may offer a grasp to one gripper while being too narrow or unreachable for another. Affordance describes the relation between an action, the agent's capabilities, and the current environment, rather than merely naming an object category. [SayCan](https://say-can.github.io/) grounds language-based task selection using estimates of whether the robot's available skills can succeed in its present state. These estimates supply information that a language model's task descriptions alone do not provide. ### Robotics systems represent affordances differently In SayCan, learned value functions help estimate skill feasibility. [CLIPort](https://cliport.github.io/) uses spatial predictions for picking and placing, described by the authors as affordance predictions. The word can therefore refer to a skill-level possibility or a location-specific action estimate. A report should identify the action being scored and what the score means. ### A prediction is not a physical guarantee A high predicted grasp score means the model favors that action under its training and observation assumptions. It does not establish that the object will remain stable or that every subsequent manipulation step will succeed. For [robotic manipulation](https://xumanoids.com/glossary/robotic-manipulation), evaluate the full action and outcome. Detecting a plausible place to grip is useful evidence about action selection, but it is not equivalent to demonstrating the complete task. ### Sources - [SayCan: Grounding Language in Robotic Affordances](https://say-can.github.io/) - [CLIPort: What and Where Pathways for Robotic Manipulation](https://cliport.github.io/) --- ## Backdrivability https://xumanoids.com/glossary/backdrivability Backdrivability is the ability of an external load applied at a mechanism’s output to drive motion back through its transmission. In a robot joint, it describes how readily an outside force can move the joint and its actuator. Updated: 2026-10-05 Also known as: Backdrivability of an actuator, Backdriveability ### Moving the joint from the outside If a person pushes a robot's hand and the joint transmission moves the motor in response, the mechanism is being backdriven. Friction, gear reduction, rotor inertia, and the load all influence how much force is required and how the motion develops. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-9-actuation-gearing-and-friction/) explains why the apparent motor inertia increases with the square of the gear ratio in an ideal transmission. This is one reason low-ratio designs can feel different from heavily geared joints when moved externally. ### Why it matters for interaction [Gealy and colleagues](https://arxiv.org/abs/1904.03815) connect backdrivable actuation with force-controlled manipulation. [Katz's actuator work](https://dspace.mit.edu/handle/1721.1/118671) likewise uses low-ratio transmissions to support torque control in dynamic robots. Such mechanisms can be useful for [teleoperation](https://xumanoids.com/glossary/teleoperation) and for responding to contact without a large commanded position error. ### Passive mechanics and active behavior differ A mechanically backdrivable joint can still resist motion when its controller commands a high stiffness. Conversely, feedback can make a less backdrivable mechanism yield actively. A compliant spring also permits some output motion without requiring the gearbox to turn. Describe whether backdrivability was measured with power off, gravity compensation, or an active controller; these conditions do not represent the same property. ### Sources - [Gealy et al.: Quasi-Direct Drive for Low-Cost Compliant Robotic Manipulation](https://arxiv.org/abs/1904.03815) - [Modern Robotics: Actuation, Gearing, and Friction](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-9-actuation-gearing-and-friction/) - [Katz: A Low Cost Modular Actuator for Dynamic Robots, MIT thesis](https://dspace.mit.edu/handle/1721.1/118671) --- ## Behavior cloning https://xumanoids.com/glossary/behavior-cloning Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem. Updated: 2026-10-05 Also known as: Behaviour cloning, Behavioral cloning, Behavioural cloning, BC ### Fit actions to demonstrated observations A training example pairs what the demonstrator observed with the action taken. The learner adjusts its predictions to match those examples. The [DAgger paper's supervised-imitation formulation](https://arxiv.org/html/1011.0686v3) formalizes this approach under the distribution of states visited by the expert. For a robot arm, the input might contain camera images and joint positions, while the target is a joint command or an [end-effector](https://xumanoids.com/glossary/end-effector) movement. Predicting an action sequence rather than one action remains compatible with supervised imitation, as illustrated by [ACT](https://tonyzhaozh.github.io/aloha/). ### Small mistakes can change later inputs Robot actions influence what the policy observes next. If a learned grasp misses slightly, the next image may differ from anything in the demonstrations. Further errors can then accumulate. The DAgger analysis explains why low prediction error on expert data does not automatically imply low error over an entire executed task. ### Distinguish copying from reward optimization Behavior cloning learns to match demonstrated actions. [Offline reinforcement learning](https://xumanoids.com/glossary/offline-reinforcement-learning) instead uses previously collected experience to optimize a reward-based objective. Both can use a fixed dataset, but their objectives differ. [Dataset aggregation](https://xumanoids.com/glossary/dataset-aggregation) addresses a specific weakness of basic cloning by collecting expert labels for states the learner actually visits. It requires additional interaction and expert access. ### Sources - [DAgger: A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning](https://arxiv.org/html/1011.0686v3) - [Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware](https://tonyzhaozh.github.io/aloha/) --- ## Bipedal locomotion https://xumanoids.com/glossary/bipedal-locomotion Bipedal locomotion is movement using two legs, with body motion coordinated through changing contacts between the feet and the environment. It includes walking and running. Updated: 2026-10-05 Also known as: Biped locomotion ### Moving the body through foot contacts A foot can push on the ground only while it is in contact. To take a step, a robot must unload and move one foot while the remaining contacts support the required body motion. The [MIT legged-robot notes](https://underactuated.mit.edu/humanoids.html) explain this as a coupled problem of contact placement, contact forces, and body dynamics. For a humanoid walking across a room, planning where to put the next foot is only part of the problem. The robot must also move its [center of mass](https://xumanoids.com/glossary/center-of-mass), keep the swinging foot clear of obstacles, and respect leg and actuator limits. ### Recovering from disturbances Balance recovery can combine shifting pressure under a foot, changing body angular momentum, and changing the next step. [Griffin and colleagues](https://arxiv.org/abs/2307.11968) study how step reachability and timing affect these choices. The position that would help restore balance may be outside the leg's reach or unavailable before the foot lands. ### Reading a walking demonstration Specify terrain, walking speed, support from the arms, and external assistance when describing a result. [Whole-body control](https://xumanoids.com/glossary/whole-body-control) may coordinate many joints, but two-legged movement alone does not establish autonomous navigation or reliable operation on every surface. ### Sources - [MIT Underactuated Robotics: Highly-articulated Legged Robots](https://underactuated.mit.edu/humanoids.html) - [Griffin et al.: Reachability Aware Capture Regions with Time Adjustment and Cross-Over for Step Recovery](https://arxiv.org/abs/2307.11968) --- ## Capture point https://xumanoids.com/glossary/capture-point The capture point is a model-dependent location where support can be placed to bring a moving robot toward rest without further steps. In the constant-height linear inverted pendulum model, the instantaneous capture point combines center-of-mass position and velocity. Updated: 2026-10-05 Also known as: Instantaneous capture point, ICP ### Velocity changes where a robot must step Two robots can have the same [center-of-mass](https://xumanoids.com/glossary/center-of-mass) position but need different recovery steps because one is moving faster. For the constant-height [linear inverted pendulum model](https://xumanoids.com/glossary/linear-inverted-pendulum-model), the horizontal instantaneous capture point is `xi = x + v / omega`, where `omega = sqrt(g / h)`. Here `x` is position, `v` is velocity, and `h` is height above the support plane. [Griffin and colleagues](https://arxiv.org/abs/2307.11968) use this state to control divergent motion. Placing the appropriate support or effective moment pivot at that location holds the capture point stationary while the center of mass converges toward it. ### A target is not an executable footstep A recovery step must fit within the leg's reachable region and available swing time. The [same study](https://arxiv.org/abs/2307.11968) incorporates those constraints into capture regions and examines changes to step position and timing. A point outside the current foot does not alone specify a successful recovery action. ### Assumptions behind the formula The simple expression assumes constant height and the associated pendulum dynamics. Changing height, substantial angular momentum, finite feet, and multiple future steps alter the analysis. A capture point is therefore a balance-planning quantity tied to a model, not a universal boundary between falling and recovery. ### Sources - [Griffin et al.: Reachability Aware Capture Regions with Time Adjustment and Cross-Over for Step Recovery](https://arxiv.org/abs/2307.11968) --- ## Catastrophic forgetting https://xumanoids.com/glossary/catastrophic-forgetting Catastrophic forgetting is a substantial loss of previously learned capability when a model is trained on new tasks or data. It is a central problem in sequential and continual learning. Updated: 2026-10-05 Also known as: Catastrophic interference ### New training can interfere with old skills [Overcoming Catastrophic Forgetting in Neural Networks](https://arxiv.org/abs/1612.00796) studies the difficulty of learning tasks sequentially while retaining earlier expertise. Shared parameters that supported an old task can change when optimization focuses on a new one. For a robot-policy example, fine-tuning on a new insertion task could reduce performance on an earlier grasping task. That is a possible form of forgetting; it must be measured rather than inferred simply because the model was updated. ### Preserve parameters important to earlier tasks The paper introduces elastic weight consolidation, which slows changes to weights estimated to be important for previously learned tasks. Its experiments include sequential image-classification tasks and Atari games. This is one proposed mitigation. The reported results do not establish that every robot skill can be retained while adding unlimited new capabilities. ### Evaluate earlier tasks after each update A new-task success score alone cannot reveal forgetting. Assessment needs comparable evaluations of both the new task and the old tasks after adaptation. For a [robot foundation model](https://xumanoids.com/glossary/robot-foundation-model), this distinction matters when a broadly trained model is specialized for one platform or environment. Improvement on the new setting and retention of previous abilities are separate outcomes. Changes in hardware or evaluation conditions should also be separated from losses caused by model training. ### Sources - [Overcoming Catastrophic Forgetting in Neural Networks](https://arxiv.org/abs/1612.00796) --- ## Center of mass https://xumanoids.com/glossary/center-of-mass The center of mass is the mass-weighted average position of a body or a collection of bodies. For an articulated robot, its position changes as the links move. Updated: 2026-10-05 Also known as: Centre of mass, CoM, COM ### Combining the masses of the links A humanoid's torso, arms, legs, and payload each contribute to the whole-system center of mass. [MIT's derivation](https://underactuated.mit.edu/humanoids.html) computes its position by multiplying each link's center position by that link's mass, adding those products, and dividing by total mass. Every position must be expressed in the same coordinate frame. Extending a heavy arm therefore moves the robot's center of mass even if its feet stay still. Carrying a box changes the combined robot-and-payload center of mass; the model must include the box if that combined system is being controlled. ### Position and motion both matter In slow standing on level ground, projecting the center of mass onto the [support polygon](https://xumanoids.com/glossary/support-polygon) helps assess whether gravity can be balanced. During walking, acceleration and angular momentum also matter. A moving robot can require a recovery step even when its center of mass currently projects between its feet. ### A useful summary with missing detail [Centroidal dynamics](https://xumanoids.com/glossary/centroidal-dynamics) connect center-of-mass motion to external forces and momentum. This reduces planning complexity, but a feasible center-of-mass trajectory does not by itself establish that every joint can execute it. The [MIT notes](https://underactuated.mit.edu/humanoids.html) identify joint-position and effort limits as information that simplified centroidal planning can miss. ### Sources - [MIT Underactuated Robotics: Highly-articulated Legged Robots](https://underactuated.mit.edu/humanoids.html) --- ## Centroidal dynamics https://xumanoids.com/glossary/centroidal-dynamics Centroidal dynamics describe the motion of a multibody system’s center of mass and the evolution of its total linear and angular momentum. External forces and moments determine the rates of change of those momenta. Updated: 2026-10-05 Also known as: Centroidal robot dynamics ### Summarizing the whole robot A humanoid has many links, but their combined linear momentum equals total mass times [center-of-mass](https://xumanoids.com/glossary/center-of-mass) velocity. Its angular momentum about that center includes the contributions of all moving links. [MIT's derivation](https://underactuated.mit.edu/humanoids.html) shows how the full system can be summarized through these momentum quantities. Ground reaction forces, gravity, and other external contacts change the total momentum. Internal joint torques redistribute motion among the links; by themselves they do not create a net external wrench on the complete robot. ### Planning useful contact forces A walking planner can optimize center-of-mass motion, angular momentum, and foot forces before solving for every joint trajectory. Unlike the simplest [linear inverted pendulum model](https://xumanoids.com/glossary/linear-inverted-pendulum-model), centroidal formulations can explicitly retain angular-momentum changes and more general contact geometry. The [survey by Wensing and colleagues](https://arxiv.org/abs/2211.11644) compares the simplifications used by different planners. ### Connecting momentum to joint motion A centroidal momentum matrix maps generalized velocity to total momentum for a given configuration. This helps connect the reduced description to [whole-body control](https://xumanoids.com/glossary/whole-body-control). However, the [MIT notes](https://underactuated.mit.edu/humanoids.html) warn that simplified centroidal planning can miss joint position and effort limits. A contact-force plan may satisfy momentum balance yet require an unreachable posture or excessive joint torque. Consistency with the full robot must still be checked. ### Sources - [MIT Underactuated Robotics: Highly-articulated Legged Robots](https://underactuated.mit.edu/humanoids.html) - [Wensing et al.: Optimization-Based Control for Dynamic Legged Robots](https://arxiv.org/abs/2211.11644) --- ## Configuration space https://xumanoids.com/glossary/configuration-space Configuration space is the set of all possible configurations of a robot or mechanical system. Each point specifies the entire modeled arrangement, and the space has as many local dimensions as the system has degrees of freedom. Updated: 2026-10-05 Also known as: C-space, Configuration space of a robot ### A point can describe a whole robot For a two-joint arm, one configuration can be recorded as two joint angles. A path through configuration space then describes how both angles change together. It is not the path traced by the hand alone: different whole-arm configurations can place the hand at the same point. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-1-degrees-of-freedom-of-a-rigid-body/) defines the configuration by the positions of all points on the modeled robot, compressed into independent coordinates where possible. ### Angles make the space wrap around The space need not behave like an ordinary flat coordinate grid. Two freely rotating joints have a configuration space shaped like a torus. Each angular coordinate wraps around after one revolution. The [configuration-space topology lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-3-1-configuration-space-topology/) shows why a smooth physical motion can cross a discontinuity in its numerical angle representation. Joint limits change the permitted set, so unrestricted circular coordinates should not be assumed for every joint. ### Configuration space and workspace differ [Workspace](https://xumanoids.com/glossary/workspace) describes reachable end-effector poses or positions. Configuration space describes the robot's complete arrangement, as the [task-space comparison](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-5-task-space-and-workspace/) explains. [Forward kinematics](https://xumanoids.com/glossary/forward-kinematics) maps from the latter to the former, which is why several configuration-space points can correspond to one task-space target. ### Sources - [Modern Robotics: Degrees of Freedom of a Rigid Body](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-1-degrees-of-freedom-of-a-rigid-body/) - [Modern Robotics: Configuration Space Topology](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-3-1-configuration-space-topology/) - [Modern Robotics: Task Space and Workspace](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-5-task-space-and-workspace/) --- ## Contact wrench cone https://xumanoids.com/glossary/contact-wrench-cone A contact wrench cone is the set of resultant forces and moments that a modelled contact can transmit without violating unilateral-contact and friction constraints. It gives legged-robot controllers a compact test for whether a foot or other support contact can remain feasible. Updated: 2026-10-06 Also known as: CWC, Contact-wrench cone ### From distributed pressure to one resultant wrench A foot sole touches the ground over an area, so the underlying model may contain many possible contact forces. Their combined effect on the robot is a six-dimensional [wrench](https://xumanoids.com/glossary/wrench): three force components and three moment components. The contact wrench cone collects the resultants that can be produced while contact forces remain compressive and obey the chosen friction model. This differs from a [friction cone](https://xumanoids.com/glossary/friction-cone), which constrains a force at a point. The contact wrench cone also represents moments created by forces distributed across a finite support area. ### What a rectangular foot must satisfy For a rigid rectangular surface under Coulomb friction, [Caron, Pham, and Nakamura](https://arxiv.org/abs/1501.04719) derive closed-form conditions for the contact wrench. The conditions combine a friction bound on the resultant force, a [zero-moment point](https://xumanoids.com/glossary/zero-moment-point) inside the support area, and bounds on yaw torque. The yaw condition matters because a foot may twist even when the force and zero-moment-point tests pass. A controller or planner can test candidate whole-body forces against this set instead of retaining a separate force variable at every sampled point on the sole. In multi-contact motion, cones or their polyhedral approximations can be combined to reason about hands, feet, and other supports. ### Environment feasibility is not actuator feasibility The contact wrench cone describes what the modelled environment contact can transmit. It does not by itself prove that the robot's motors can generate the required wrench. [Orsolino and colleagues](https://arxiv.org/abs/1712.06833) distinguish the contact wrench cone from an actuation wrench polytope and intersect the two to form a feasible wrench polytope. The result also inherits its modelling assumptions. Friction coefficients are uncertain, real soles deform, terrain is rarely perfectly planar, and contacts can start to slip or lift. Controllers commonly add margins, but a larger margin is not a substitute for validating the contact model on the robot and surface in question. ### Sources - [Caron et al.: Stability of Surface Contacts for Humanoid Robots](https://arxiv.org/abs/1501.04719) - [Orsolino et al.: Wrench based Feasibility Analysis for Legged Robots](https://arxiv.org/abs/1712.06833) --- ## Contact-implicit optimization https://xumanoids.com/glossary/contact-implicit-optimization Contact-implicit optimization plans motion while allowing contact events and forces to emerge from contact constraints in the optimization. It avoids requiring every contact transition to be fixed in a predefined mode sequence. Updated: 2026-10-05 Also known as: Contact-implicit trajectory optimization ### Letting contact timing emerge A conventional plan may specify exactly when each foot touches and leaves the ground. Contact-implicit methods instead optimize motion and contact forces together. [Posa, Cantu, and Tedrake](https://groups.csail.mit.edu/robotics-center/public_papers/Posa13.pdf) developed a direct method for rigid-body trajectories with inelastic impacts and Coulomb friction that removes the need to prescribe the mode ordering in advance. A common formulation uses complementarity: a contact gap is nonnegative, the normal force is nonnegative, and their product is zero. A positive gap therefore excludes a contact force; a positive force requires closed contact. Other formulations use smooth contact approximations. ### Uses in locomotion and manipulation The optimizer can search for useful stepping or pushing behavior as part of a [trajectory-optimization](https://xumanoids.com/glossary/trajectory-optimization) problem. This is helpful when the contact sequence is itself an important design choice, rather than something already known. ### Search freedom increases numerical difficulty Complementarity creates a difficult nonconvex problem. The [optimization survey by Wensing and colleagues](https://arxiv.org/abs/2211.11644) discusses poor numerical conditioning and dependence on good initial guesses. Relaxed constraints may permit unphysical intermediate solutions while the solver searches. A final trajectory therefore needs checks against the intended contact model and constraints; finding a local solution does not establish that it is the best possible contact strategy. ### Sources - [Posa, Cantu, and Tedrake: A Direct Method for Trajectory Optimization of Rigid Bodies Through Contact](https://groups.csail.mit.edu/robotics-center/public_papers/Posa13.pdf) - [Wensing et al.: Optimization-Based Control for Dynamic Legged Robots](https://arxiv.org/abs/2211.11644) --- ## Control barrier function https://xumanoids.com/glossary/control-barrier-function A control barrier function is a mathematical function used to express a safe set for a dynamical system and constrain control inputs so the system remains inside that set. It is commonly used as a safety filter around a nominal robot controller. Updated: 2026-10-06 Also known as: CBF, Control barrier functions ### A boundary on admissible motion A control barrier function assigns values to robot states so that an inequality, often written as `h(x) >= 0`, describes a safe set. The controller then chooses inputs that satisfy a condition on how `h` changes. Under the assumptions in the formulation, satisfying that condition makes the set forward invariant: a trajectory that starts in the set remains there. [Ames and colleagues](https://arxiv.org/abs/1609.06408) develop this relationship between barrier functions, control inputs, and forward invariance. The function does not have to generate the robot's desired behaviour. A separate controller can request an action for tracking, navigation, or [robotic manipulation](https://xumanoids.com/glossary/robotic-manipulation), while the barrier condition limits which actions are admissible. ### A quadratic program can act as a safety filter One common construction puts the barrier inequality into a quadratic program. The program finds a control input close to the nominal command while satisfying the safety constraint. The same optimisation can combine a control barrier function with a control Lyapunov function that represents a performance objective, as demonstrated in the [foundational CBF-QP paper](https://arxiv.org/abs/1609.06408). Robotics applications can encode constraints such as separation from an obstacle or keeping a sensing trajectory out of a modelled collision region. A 2026 preprint called [Splat-CBF](https://arxiv.org/abs/2609.23100) reports using a risk-aware barrier constraint with a 3D Gaussian map while treating informative camera motion as a softer objective. That is evidence for the method studied in that system, not proof that every Gaussian map or barrier controller is safe. ### The guarantee is conditional A CBF certificate depends on the stated system dynamics, safe-set definition, state estimate, input limits, and numerical solution. Unmodelled contacts, delayed measurements, or a constraint that admits no feasible input can break the connection between the mathematical condition and the physical robot. A control barrier function is therefore not a general safety claim about a robot. It certifies a particular property only under its explicit assumptions. ### Sources - [Ames et al.: Control Barrier Function Based Quadratic Programs for Safety Critical Systems](https://arxiv.org/abs/1609.06408) - [Khass et al.: Splat-CBF: Safe Next-Best-View Control in 3D Gaussian-Splat Maps](https://arxiv.org/abs/2609.23100) --- ## Cross-embodiment learning https://xumanoids.com/glossary/cross-embodiment-learning Cross-embodiment learning uses experience from different robot bodies to train representations or policies that can transfer across those bodies. It requires a way to handle differences in sensing, geometry, and available actions. Updated: 2026-10-05 Also known as: Cross embodiment learning ### Share experience across robot bodies [Open X-Embodiment](https://robotics-transformer-x.github.io/) brings robot datasets into standardized formats and studies policies trained across multiple robots. Its RT-X experiments investigate whether experience from different platforms can improve a policy's behavior beyond training on one platform alone. A camera image of a cup may contain useful information for several robot arms. The motor commands needed to grasp that cup still depend on the arm, gripper, and controller. ### Standardized data does not make bodies identical Robot datasets can differ in sensor views, joint counts, command meanings, and task labels. A shared format helps access the data, but the learning system still needs compatible representations or adaptation mechanisms. [Octo](https://octo-models.github.io/) explicitly studies fine-tuning to new observations and action spaces. Its results separate direct execution on supported setups from adaptation to new ones. ### Distinguish learning from motion conversion [Motion retargeting](https://xumanoids.com/glossary/motion-retargeting) maps a particular motion onto another body's geometry. Cross-embodiment learning instead concerns how training across bodies produces transferable knowledge or behavior. A pipeline can use both. Transfer from one arm to another does not establish transfer to humanoid locomotion. The source and target bodies, available target data, and evaluated tasks are necessary context for any cross-embodiment result. ### Sources - [Open X-Embodiment: Robotic Learning Datasets and RT-X Models](https://robotics-transformer-x.github.io/) - [Octo: An Open-Source Generalist Robot Policy](https://octo-models.github.io/) --- ## Dataset aggregation https://xumanoids.com/glossary/dataset-aggregation Dataset aggregation, usually called DAgger in imitation learning, is an iterative algorithm that collects expert action labels at states visited by a learner. It adds those examples to an accumulated dataset and retrains the policy. Updated: 2026-10-05 Also known as: DAgger ### Train on states the learner encounters The [DAgger algorithm](https://arxiv.org/html/1011.0686v3) alternates between executing a policy, obtaining the expert's preferred actions for visited states, and training on the combined dataset. Its name abbreviates Dataset Aggregation. The original formulation can mix expert and learner behavior during data collection. The expert supplies action labels even when the learner has caused the system to enter a state outside the original demonstrations. ### Correct the distribution mismatch Basic [behavior cloning](https://xumanoids.com/glossary/behavior-cloning) learns from states produced by expert behavior. Its own errors can later move it into different states. DAgger explicitly gathers training examples from the distributions induced during its iterative learning process. For a reaching robot, this could mean asking for a corrective action after the learned controller approaches an object from an awkward pose. Adding another perfectly executed demonstration might never show that pose. ### Expert access is a real requirement DAgger is more specific than simply combining robot datasets. Its defining feature is the loop between learner execution, expert labeling, aggregation, and retraining. The original analysis gives performance guarantees under stated learning and reduction assumptions. It does not guarantee safe exploration on physical hardware. Deploying the collection loop requires an expert who can label encountered states and a suitable way to manage the learner's actions during those trials. ### Sources - [DAgger: A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning](https://arxiv.org/html/1011.0686v3) --- ## Degrees of freedom https://xumanoids.com/glossary/degrees-of-freedom Degrees of freedom are the number of independent coordinates needed locally to describe a system configuration. In robotics, this count depends on the bodies, joints, and independent constraints in the model. Updated: 2026-10-05 Also known as: DOF, Degree of freedom ### Counting independent motion A free rigid body in three-dimensional space has six degrees of freedom: three for position and three for orientation. A rigid body restricted to a plane has three. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-1-degrees-of-freedom-of-a-rigid-body/) derives these counts by subtracting independent geometric constraints from the available point coordinates. A [revolute joint](https://xumanoids.com/glossary/revolute-joint) leaves one relative rotation between two links. A [prismatic joint](https://xumanoids.com/glossary/prismatic-joint) leaves one relative translation. A spherical joint allows three relative rotational freedoms. ### Constraints change the count For an open serial chain, independent joint coordinates provide a convenient count. Closed loops require more care because joint motions constrain one another. The [mechanism-counting lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) shows that a simple mobility formula can give the wrong answer when its assumed constraints are dependent. A joint's finite travel range restricts its allowed values but does not remove its degree of freedom throughout the interior of that range. ### State what the count includes A fixed-base arm and the same arm on a floating base have different configuration counts. [MIT's modeling notes](https://manipulation.mit.edu/pick.html) explain how fixing a base removes its floating freedoms. When comparing humanoids, establish whether the reported count includes the base, hands, and coupled mechanisms rather than treating every published number as directly comparable. ### Sources - [Modern Robotics: Degrees of Freedom of a Rigid Body](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-1-degrees-of-freedom-of-a-rigid-body/) - [Modern Robotics: Degrees of Freedom of a Robot](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Denavit-Hartenberg parameters https://xumanoids.com/glossary/denavit-hartenberg-parameters Denavit-Hartenberg parameters are four geometric quantities that describe the relative placement of successive link frames in a robot kinematic chain. They provide a systematic way to construct the transformations used in forward kinematics. Updated: 2026-10-05 Also known as: DH parameters, D-H parameters ### Four parameters for each link transform A DH table describes the geometry of a [kinematic chain](https://xumanoids.com/glossary/kinematic-chain) after assigning coordinate frames to its links. In the standard convention, the four parameters are: - Joint angle, theta: rotation about the preceding frame's z-axis. - Link offset, d: translation along that z-axis. - Link length, a: translation along the new frame's x-axis, the common normal between the joint axes. - Link twist, alpha: rotation about that x-axis to align the successive z-axes. For a [revolute joint](https://xumanoids.com/glossary/revolute-joint), theta is the joint variable. For a [prismatic joint](https://xumanoids.com/glossary/prismatic-joint), d is the variable. The [Robotics Toolbox documentation](https://petercorke.github.io/robotics-toolbox-python/arm_dh.html) gives the corresponding link transforms for both types. ### Standard and modified conventions Standard DH multiplies the z rotation and translation before the x translation and rotation. The modified convention documented by the Toolbox places the preceding link's x operations before the current joint's z operations. Frame placement and parameter indexing therefore differ. A table is incomplete without its convention, frame definitions, joint zero offsets, and units. Copying values between conventions without converting the frames can produce incorrect [forward kinematics](https://xumanoids.com/glossary/forward-kinematics). ### Geometry, not a dynamics model Multiplying the link transforms gives an [end-effector](https://xumanoids.com/glossary/end-effector) pose. The DH parameters alone do not describe link masses, inertias, friction, or motor limits; the Toolbox stores those as separate model properties. ### Sources - [Robotics Toolbox for Python: Denavit-Hartenberg models](https://petercorke.github.io/robotics-toolbox-python/arm_dh.html) --- ## Differentiable simulation https://xumanoids.com/glossary/differentiable-simulation Differentiable simulation is physical simulation that provides derivatives of simulated outcomes or losses with respect to inputs such as controls, initial states, model parameters, or robot design variables. Those gradients can drive optimisation and learning through the simulated dynamics. Updated: 2026-10-06 Also known as: Differentiable physics, Differentiable physics simulation ### Gradients through simulated dynamics An ordinary simulator maps an initial state and sequence of inputs to a trajectory. A differentiable simulator also computes how a chosen output changes when those inputs or model parameters change. Reverse-mode differentiation can propagate a task loss backwards through many simulation steps to obtain gradients for optimisation. [DiffTaichi](https://arxiv.org/abs/1910.00935) uses source transformations and a recorded simulation program structure to generate gradients for several physical simulators. Its reported examples include optimising neural-network controllers. The concept is broader than that implementation and can use automatic, analytic, or carefully derived numerical differentiation. ### Uses in robotics Gradients can tune a control sequence in [trajectory optimization](https://xumanoids.com/glossary/trajectory-optimization), estimate friction or mass in [system identification](https://xumanoids.com/glossary/system-identification), train a policy, or change a robot design parameter. Compared with perturbing every parameter independently, a reverse-mode gradient can be attractive when one scalar loss depends on many decisions. Differentiable simulation is not the same as [reinforcement learning](https://xumanoids.com/glossary/reinforcement-learning). It is a model capability that an optimiser or learning algorithm may use. A policy can be trained without differentiable physics, and a differentiable simulator can optimise variables without training a policy. ### Contact and long rollouts complicate gradients Impacts, frictional transitions, and contact creation are not smooth in the same way as free-flight dynamics. Implementations may soften contact or choose surrogate derivatives, which changes the optimisation problem. Long rollouts can also amplify numerical error. A 2026 study by [Yang and colleagues](https://arxiv.org/abs/2609.34666) reports gradient sensitivity to parallel accumulation order, rollout length, and objective construction in two material-manipulation benchmarks. Those findings concern the tested simulators and tasks, but they illustrate why gradient checks and reproducibility matter. Even an accurate simulation gradient only differentiates the model; it does not remove the [sim-to-real gap](https://xumanoids.com/glossary/sim-to-real-transfer). ### Sources - [Hu et al.: DiffTaichi: Differentiable Programming for Physical Simulation](https://arxiv.org/abs/1910.00935) - [Yang et al.: On the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material Manipulation](https://arxiv.org/abs/2609.34666) --- ## Diffusion policy https://xumanoids.com/glossary/diffusion-policy A diffusion policy generates robot actions through a learned denoising process conditioned on observations. It commonly predicts an action sequence by progressively refining a noisy candidate rather than predicting one action with a single direct regression. Updated: 2026-10-05 ### Denoising produces an action sequence The [Diffusion Policy research](https://diffusion-policy.cs.columbia.edu/) models visuomotor behavior as a conditional denoising diffusion process. During training, the model learns to remove noise from demonstrated actions using observations as context. During execution, repeated refinement turns an initial noisy sequence into a candidate robot motion. The noisy sequence exists inside the model's computation. It is not a command to make the physical robot move randomly. ### Multiple valid motions can remain distinct A robot may be able to push an object around either side of an obstacle. A model trained to average incompatible demonstrations could propose an unsuitable middle path. Diffusion Policy is designed to represent multiple modes of an action distribution and select a coherent sequence, as illustrated in the authors' manipulation experiments. It also combines [action chunking](https://xumanoids.com/glossary/action-chunking) with receding-horizon execution: the controller executes part of a predicted sequence, observes again, and generates another sequence. ### Sampling and feedback impose practical limits Iterative generation takes computation, so sampling settings and observation timing affect the control loop. The paper's demonstrations include pushing, mug flipping, and sauce manipulation. These results support those evaluated setups; diffusion decoding alone does not provide a collision guarantee or establish humanoid walking capability. ### Sources - [Diffusion Policy: Visuomotor Policy Learning via Action Diffusion](https://diffusion-policy.cs.columbia.edu/) --- ## Domain adaptation https://xumanoids.com/glossary/domain-adaptation Domain adaptation adjusts a learned model to work on a target data distribution that differs from its source training distribution. Robotics examples include adapting perception from simulation to camera images or adapting behavior to changed physical conditions. Updated: 2026-10-05 ### Learn across a source and target mismatch [Domain-Adversarial Training of Neural Networks](https://arxiv.org/abs/1505.07818) studies learning representations that support a task while reducing distinguishability between source and target domains. Its formulation uses labeled source examples and unlabeled target examples. For robot perception, the source could be rendered images and the target could be images from a particular physical camera. The task can remain the same even though image appearance changes. ### Adaptation also appears in robot control [Peng and colleagues' locomotion framework](https://arxiv.org/html/2004.00784v2) uses a domain-adaptation stage to adjust behavior through a learned dynamics representation when moving to a real robot. This targets physical behavior rather than only visual appearance. A description of adaptation should therefore specify what changes: features, model parameters, a dynamics representation, or another part of the system. ### Distinguish adaptation from broad variation [Domain randomization](https://xumanoids.com/glossary/domain-randomization) broadens training conditions. Adaptation uses information about a target domain to address a mismatch. They can be combined, as in the cited locomotion framework. The availability of target labels or interactions also matters. A method that needs target demonstrations is different from one using only unlabeled images. Adaptation results should state those requirements instead of implying that transfer occurred without target-domain information. ### Sources - [Domain-Adversarial Training of Neural Networks](https://arxiv.org/abs/1505.07818) - [Learning Agile Robotic Locomotion Skills by Imitating Animals](https://arxiv.org/html/2004.00784v2) --- ## Domain randomization https://xumanoids.com/glossary/domain-randomization Domain randomization varies properties of training environments to encourage a learned model or policy to work across changing conditions. In robotics it often randomizes simulated appearance, physical parameters, or both to support transfer to real hardware. Updated: 2026-10-05 Also known as: Domain randomisation ### Vary appearance or physical behavior The [visual domain-randomization study](https://arxiv.org/abs/1703.06907) trains object localization using rendered scenes with varied appearance, including nonrealistic random textures. The aim is for a real camera image to resemble another variation within the training experience. [Dynamics randomization](https://arxiv.org/abs/1710.06537) applies the same broad principle to simulated physical behavior. The authors use it to train a robot-arm pushing policy that transfers to hardware. ### Randomization should match the transfer problem For a visual picking system, changes in textures and lighting target perception differences. For a control policy, variation in dynamics targets differences in how actions move the robot and objects. These choices address different parts of the system. Randomization can therefore be useful without producing photorealistic images. Equally, varied images alone do not model an actuator's physical response. ### Variation is not a transfer guarantee The cited studies demonstrate particular [sim-to-real transfers](https://xumanoids.com/glossary/sim-to-real-transfer). Their results do not show that arbitrary randomization covers every real operating condition. [Domain adaptation](https://xumanoids.com/glossary/domain-adaptation) is a neighboring concept: it uses information about a target domain to reduce a mismatch. Domain randomization generally broadens training variation. A robotics pipeline may combine them, but neither term should be used as evidence that a new environment has already been validated. ### Sources - [Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World](https://arxiv.org/abs/1703.06907) - [Sim-to-Real Transfer of Robotic Control with Dynamics Randomization](https://arxiv.org/abs/1710.06537) --- ## Dynamic movement primitive https://xumanoids.com/glossary/dynamic-movement-primitive A dynamic movement primitive is a parameterised dynamical system that represents a goal-directed or rhythmic movement using stable baseline dynamics plus a learned shaping term. Robots can fit the parameters from demonstrations and adapt the resulting motion to a new goal or duration. Updated: 2026-10-06 Also known as: DMP, Dynamical movement primitive ### A movement encoded as dynamics A dynamic movement primitive starts with a simple attractor system and adds a learnable forcing term that shapes the path. [Ijspeert and colleagues](https://doi.org/10.1162/NECO_a_00393) describe point-attractor formulations for discrete movements and limit-cycle formulations for rhythmic movements. The attractor supplies the basic convergence behaviour, while the forcing term represents details learned from a sample motion. This is different from storing a timed list of joint values. The DMP generates a trajectory by integrating its dynamics. Its goal, duration, or amplitude can be changed within the formulation, although useful adaptation depends on how the primitive was designed and trained. ### From a demonstration to a reusable skill In [imitation learning](https://xumanoids.com/glossary/imitation-learning), a demonstrated reach, wipe, or placement motion can be fitted as a DMP. A task-level system can then select the primitive and set a new target, while a lower-level controller tracks the generated position, velocity, or orientation path. DMPs can represent either joint-space motion or task-space quantities such as an [end effector](https://xumanoids.com/glossary/end-effector) pose. Orientation needs an appropriate representation because ordinary subtraction does not describe all three-dimensional rotations correctly. ### Convergence does not imply task safety An attractor can pull the generated motion toward its goal without respecting obstacles, joint limits, contact forces, or human separation. [Shaw and colleagues](https://arxiv.org/abs/2209.14461) state that ordinary DMPs provide no strong guarantee of operational constraint satisfaction and propose a constrained variant using a barrier function. A primitive also preserves only what its variables and training data encode. Moving the goal can produce a mathematically valid trajectory that is unsuitable for a new object geometry or robot body. Collision checks, feasibility tests, and feedback control remain separate requirements. ### Sources - [Ijspeert et al.: Dynamical Movement Primitives: Learning Attractor Models for Motor Behaviors](https://doi.org/10.1162/NECO_a_00393) - [Shaw et al.: Constrained Dynamic Movement Primitives for Safe Learning of Motor Skills](https://arxiv.org/abs/2209.14461) --- ## Embodied AI https://xumanoids.com/glossary/embodied-ai Embodied AI is artificial intelligence that perceives and acts through a body in a physical or simulated environment. It connects sensing, reasoning, and action rather than producing only text or images. Updated: 2026-10-05 Embodied artificial intelligence ### A body can be simulated An embodied agent does not have to be a physical robot. The [Habitat research platform](https://arxiv.org/abs/1904.01201) trains virtual robots in simulated 3D environments with configurable bodies and sensors. This lets researchers study tasks such as navigation before putting a system into a real space. ### Connecting observations to actions A robot needs information about its surroundings and its own state. [PaLM-E](https://arxiv.org/abs/2303.03378) studies how to bring those observations into a language model for tasks that include robotic manipulation planning. A [vision-language-action model](https://xumanoids.com/glossary/vision-language-action-model) goes further by producing robot action outputs from visual observations and language instructions. ### What the term does not establish Embodied AI describes an approach to intelligence, not a particular body shape or level of autonomy. An agent in simulation and a [humanoid robot](https://xumanoids.com/glossary/humanoid-robot) can both be embodied systems. Calling either one embodied does not establish that it can work safely without supervision or perform every task a person can. ### Sources - [Habitat: A Platform for Embodied AI Research](https://arxiv.org/abs/1904.01201) - [PaLM-E: An Embodied Multimodal Language Model](https://arxiv.org/abs/2303.03378) --- ## End effector https://xumanoids.com/glossary/end-effector An end effector is the part of a robot positioned to perform a task at the end of a manipulator, such as a gripper, hand, suction tool, or welding tool. Its pose and interaction forces are often the quantities a task controller regulates. Updated: 2026-10-05 Also known as: End-effector, Robot end effector ### The tool that meets the task A humanoid hand is one example, but an end effector need not resemble a hand. A suction cup, screwdriver, or probe can serve the same role for a different task. Its geometry determines where useful contact occurs, while the arm positions it. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-5-task-space-and-workspace/) distinguishes the task space used to express the job from the [workspace](https://xumanoids.com/glossary/workspace) of end-effector configurations the robot can reach. A marker tracing a board can have a planar task even when the supporting arm has many joints. ### Choosing a reference frame The end-effector frame is a coordinate frame attached to the relevant part of the tool. Its origin may be at a tool tip rather than at the wrist flange. [MIT's manipulation notes](https://manipulation.mit.edu/force.html) emphasize choosing a frame near the expected contact when defining task-space interaction behavior. Changing tools therefore changes more than appearance: the tool offset changes the relation between joint angles and contact position, and the tool's mass affects dynamics and force compensation. ### Tool capability and robot capability differ A gripper can hold a particular object only if its geometry and forces suit that object. Reaching a pose does not establish a stable grasp, adequate payload, or collision-free access. Those require separate kinematic, contact, and dynamic checks. ### Sources - [Modern Robotics: Task Space and Workspace](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-5-task-space-and-workspace/) - [MIT Robotic Manipulation: Force Control](https://manipulation.mit.edu/force.html) --- ## Euler angles https://xumanoids.com/glossary/euler-angles Euler angles represent a three-dimensional orientation as an ordered sequence of three rotations about specified axes. Robotics often uses the term broadly to include roll-pitch-yaw conventions, so the exact rotation sequence must be specified. Updated: 2026-10-05 Also known as: Euler angle representation ### The sequence is part of the data Three angle values are insufficient without an axis order and a statement of whether rotations use fixed axes or axes that move with the body. Reordering the same rotations usually produces a different orientation. [MIT's manipulation notes](https://manipulation.mit.edu/pick.html) use roll-pitch-yaw as a specific example: roll about x, pitch about y, and yaw about z, with the extrinsic X-Y-Z convention. “Extrinsic” means rotations about fixed reference axes. “Intrinsic” uses the moving axes. The [SciPy interface](https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Rotation.from_euler.html), for example, requires the axis sequence and distinguishes intrinsic from extrinsic rotations. The documented sequence is more informative than the generic label. ### Gimbal lock is a coordinate singularity For the roll-pitch-yaw convention discussed by MIT, at a pitch of 90 degrees, roll and yaw become indistinguishable. Multiple angle combinations describe the same orientation, making the inverse coordinate map singular. This does not mean the rigid body physically loses a rotational freedom. A robot's [kinematic singularity](https://xumanoids.com/glossary/kinematic-singularity) is a separate property of its joint-to-motion mapping. ### Choosing another representation A [rotation matrix](https://xumanoids.com/glossary/rotation-matrix) or unit [quaternion](https://xumanoids.com/glossary/quaternion) avoids this angle-coordinate singularity. Angle displays can still be convenient for people, provided software treats wraparound, axis order, and singular orientations explicitly. ### Sources - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) - [SciPy: Rotation.from_euler](https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Rotation.from_euler.html) --- ## Event camera https://xumanoids.com/glossary/event-camera An event camera is a vision sensor whose pixels asynchronously report changes in brightness instead of exposing complete image frames at fixed intervals. Each event normally carries a pixel location, timestamp, and change polarity. Updated: 2026-10-06 Also known as: Event-based camera, Event vision sensor ### Pixels report change rather than frames A conventional camera samples an array of intensities at a sequence of exposure times. An event camera instead lets each pixel emit an event when the change in its measured log intensity crosses a threshold. The output is an asynchronous stream containing the pixel coordinates, time, and whether brightness increased or decreased, as described in the [event-based vision survey](https://arxiv.org/abs/1904.08405). The stream is sparse when little changes and dense around moving edges or changing illumination. It is not an ordinary video with missing frames, so algorithms usually operate on events directly or accumulate them into a chosen representation. ### Why robots use event streams The sensor's fine timing and lack of a global frame exposure can help when a robot or observed object moves quickly. The cited survey covers applications including feature tracking, optical flow, reconstruction, segmentation, and recognition. Robotics systems can also combine events with an [inertial measurement unit](https://xumanoids.com/glossary/inertial-measurement-unit) for motion estimation. For a humanoid, possible uses include tracking during rapid head motion, observing a fast hand-object interaction, or maintaining visual estimates across strong changes in illumination. Those benefits depend on the sensor, optics, algorithm, and scene; the camera type alone does not establish task performance. ### What an event stream leaves out A pixel that sees no threshold-crossing brightness change produces no event. The stream therefore does not directly supply a complete absolute-intensity image at each instant. Sensor noise, threshold variation, background activity, and bursts caused by lighting changes also need to be handled. [Gallego and colleagues](https://arxiv.org/abs/1904.08405) emphasise that event cameras require methods suited to their unconventional output. Converting events into frames can make existing vision software easier to reuse, but that conversion may discard some of the timing structure that motivated the sensor. ### Sources - [Gallego et al.: Event-based Vision: A Survey](https://arxiv.org/abs/1904.08405) --- ## Extended Kalman filter https://xumanoids.com/glossary/extended-kalman-filter An extended Kalman filter is a state estimator that applies Kalman-style prediction and correction to nonlinear models by locally linearizing them. It approximates uncertainty around the current state estimate. Updated: 2026-10-05 Also known as: EKF ### Local linearization handles nonlinear relationships Robot motion and sensor measurements often involve nonlinear functions, such as converting an orientation into a direction of travel. The EKF evaluates these functions and uses their local derivatives to propagate uncertainty. [Welch and Bishop](https://www.cs.cmu.edu/~motionplanning/papers/sbp_papers/kalman/welch_intro_kalman.pdf) describe the linearization of both process and measurement models. These derivative matrices are Jacobians. They serve the same mathematical role of local sensitivity as a [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian), although the functions being differentiated need not be arm kinematics. ### A common robot localization method The [ROS robot_localization package](https://docs.ros.org/en/noetic/api/robot_localization/html/state_estimation_nodes.html) implements an EKF that predicts motion and corrects its estimate from sensor data. It illustrates how a filter becomes part of a practical [sensor-fusion](https://xumanoids.com/glossary/sensor-fusion) system. ### Approximation is the tradeoff A nonlinear transformation does not generally preserve a Gaussian probability distribution. The EKF's local approximation can become poor when uncertainty is large or the model is strongly nonlinear over the plausible states. It also does not naturally represent several separate location hypotheses, unlike a suitably configured [particle filter](https://xumanoids.com/glossary/particle-filter). ### Sources - [Welch and Bishop: An Introduction to the Kalman Filter](https://www.cs.cmu.edu/~motionplanning/papers/sbp_papers/kalman/welch_intro_kalman.pdf) - [ROS robot_localization: State Estimation Nodes](https://docs.ros.org/en/noetic/api/robot_localization/html/state_estimation_nodes.html) --- ## Flow matching https://xumanoids.com/glossary/flow-matching Flow matching is a generative-model training method that learns a vector field for transforming a simple probability distribution into a data distribution. In robot learning, the generated samples can be continuous action sequences conditioned on observations and instructions. Updated: 2026-10-05 Also known as: FM ### Learn how a sample should move [Flow Matching for Generative Modeling](https://arxiv.org/abs/2210.02747) trains a model by regressing vector fields along chosen probability paths between noise and data. At generation time, a numerical solver follows the learned field to transform an initial sample into a data-like sample. Here, flow refers to movement through a mathematical sample space. It does not mean fluid flow around a robot or optical flow between camera frames. ### Continuous robot actions are one application The [pi0 model](https://arxiv.org/abs/2410.24164) combines a pretrained [vision-language model](https://xumanoids.com/glossary/vision-language-model) with a flow-matching action architecture. Visual and language information condition the generation of robot actions. Its authors evaluate manipulation tasks including laundry folding and box assembly. An action sample can represent several future commands, making this approach compatible with [action chunking](https://xumanoids.com/glossary/action-chunking). ### Flow matching and diffusion overlap The original flow-matching formulation can use diffusion probability paths as well as other paths. The terms therefore describe related families of generative methods rather than completely separate ideas. A chosen path, solver, and number of integration steps affect generation. Results from one model or solver configuration should not be treated as evidence that every flow-matching policy is faster or more accurate than every diffusion policy. ### Sources - [Flow Matching for Generative Modeling](https://arxiv.org/abs/2210.02747) - [Physical Intelligence: pi0, a vision-language-action flow model](https://arxiv.org/abs/2410.24164) --- ## Force closure https://xumanoids.com/glossary/force-closure Force closure is a contact condition in which the admissible contact wrenches can collectively oppose any external wrench direction on an object. The condition depends on contact locations, normals, and the assumed friction model. Updated: 2026-10-05 Also known as: Force-closure grasp ### Resisting disturbances through contacts Each finger contact contributes a set of forces allowed by its [friction cone](https://xumanoids.com/glossary/friction-cone). Those forces also create moments about the object. A grasp has force closure when positive combinations of its available contact [wrenches](https://xumanoids.com/glossary/wrench) span the full wrench space, as described in [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-3-force-closure/). This is stronger than holding an object against gravity in one orientation. A grasp may support a downward load while still allowing the object to rotate or escape under another disturbance. ### How it differs from form closure [Form closure](https://xumanoids.com/glossary/form-closure) blocks motion through geometry without depending on friction. Force closure allows frictional forces to contribute. In the frictionless point-contact model, force closure is equivalent to first-order form closure; with friction, fewer contacts can sometimes suffice. ### A capability condition, not unlimited strength The mathematical test does not mean a real hand can resist arbitrarily large forces. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-3-force-closure/) explicitly notes that finger actuators may be unable to generate the squeezing forces the test permits. The contact model also matters. Two ideal point contacts on a spatial object cannot resist torque about the line joining them, while soft contact patches can supply additional torsional resistance. State friction and contact assumptions when reporting force closure. ### Sources - [Modern Robotics: Force Closure](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-3-force-closure/) --- ## Force control https://xumanoids.com/glossary/force-control Force control regulates the force or wrench a robot applies to its environment. It may use a robot model, measured interaction forces, or both to produce joint commands that achieve a desired contact load. Updated: 2026-10-05 Also known as: Robot force control ### Regulating contact load When a robot presses a tool against a surface, accurate position alone does not guarantee the desired pressure. A small error in the assumed surface location can create a large force if the tool and surface are stiff. Force control makes that interaction load an explicit target. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-5-force-control/) describes a quasistatic method that maps a desired end-effector [wrench](https://xumanoids.com/glossary/wrench) into joint torques using the [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian) transpose, with gravity compensation. Adding a force-torque sensor permits feedback on the difference between desired and measured wrench. ### Contact must support the requested force A hand pressing on a tabletop can push into it, but cannot pull it upward without a grasp or adhesion. It may also slide if the tangential force exceeds available friction. A controller's target therefore needs to fit the actual contact constraints. ### Force and motion are coupled For wiping, the robot may regulate the normal pressing force while controlling motion along the surface. This is a use of [hybrid position-force control](https://xumanoids.com/glossary/hybrid-position-force-control). Alternatively, [impedance control](https://xumanoids.com/glossary/impedance-control) sets a compliant response. Force sensing can be noisy, and [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-5-force-control/) notes that differentiation amplifies this noise, which influences feedback and filtering choices. ### Sources - [Modern Robotics: Force Control](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-5-force-control/) --- ## Form closure https://xumanoids.com/glossary/form-closure Form closure is a condition in which the geometry of stationary contacts prevents an object from moving, without relying on friction. First-order form closure can be established from contact positions and normals alone. Updated: 2026-10-05 Also known as: Form-closure grasp ### Holding an object through geometry A fixture can block every available translation and rotation of a workpiece by placing contact surfaces around it. The object cannot move without penetrating a fixture surface, even if the contacts are treated as frictionless. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-1-7-form-closure/) defines this kinematic immobilization as form closure. This makes form closure useful when reasoning about fixturing. Simply resting an object in a shallow tray does not establish form closure, because lifting may remain possible. ### First-order analysis has a specific scope A first-order test examines contact positions and surface normals. It checks whether any nonzero instantaneous rigid-body motion remains feasible. If that test establishes form closure, the more detailed geometry also constrains the object. Failure of the first-order test is less conclusive: surface curvature can prevent a motion that appears possible when only contact normals are considered. The [source's examples](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-1-7-form-closure/) show why higher-order geometry can establish closure in cases missed by the simpler test. ### Different from force closure [Force closure](https://xumanoids.com/glossary/force-closure) asks whether allowable contact forces can oppose every external wrench direction and may rely on friction, as explained in [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-3-force-closure/). Form closure is a geometric condition. Neither description alone gives the fixture's load capacity: real contacts, supporting structures, and actuators can deform or fail under sufficiently large loads. ### Sources - [Modern Robotics: Form Closure](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-1-7-form-closure/) - [Modern Robotics: Force Closure](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-3-force-closure/) --- ## Forward dynamics https://xumanoids.com/glossary/forward-dynamics Forward dynamics predicts a robot's acceleration from its current configuration, velocity, applied joint forces or torques, and external forces. It uses the robot's mass, inertia, and other modeled dynamic properties. Updated: 2026-10-05 Also known as: Robot forward dynamics ### Predicting what applied effort will do Given the current joint positions and velocities, a dynamics model accounts for inertia, gravity, and motion-dependent forces. It then solves for the acceleration produced by the supplied joint effort and external loading. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-5-forward-dynamics-of-open-chains/) demonstrates solving this problem using quantities assembled from [inverse dynamics](https://xumanoids.com/glossary/inverse-dynamics). ### Acceleration becomes a simulated trajectory A simulator integrates the calculated acceleration to update velocity and position over time. Repeating that process predicts a trajectory for chosen inputs. This is different from [forward kinematics](https://xumanoids.com/glossary/forward-kinematics), which calculates a pose directly from known joint positions. The integration method and time step matter. The Modern Robotics example discusses energy drift caused by numerical integration even when the modeled system has no dissipation. ### Simulation reflects the model's scope The reference lesson shows an arm swinging with zero commanded motor torque. Omitting joint friction makes the motion look different from a real arm; adding a friction model changes the prediction. Contact, actuator behavior, and inertial parameters likewise have to be represented when relevant to the intended prediction. [System identification](https://xumanoids.com/glossary/system-identification) can help estimate model parameters, but a successful simulation alone does not establish that the physical robot will follow the same trajectory. ### Sources - [Modern Robotics: Forward Dynamics of Open Chains](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-5-forward-dynamics-of-open-chains/) --- ## Forward kinematics https://xumanoids.com/glossary/forward-kinematics Forward kinematics calculates the position and orientation of a robot link or end-effector from the robot geometry and joint positions. It maps a robot configuration to a pose. Updated: 2026-10-05 Also known as: FK ### From joint readings to a hand pose A robot arm reports joint angles, but a manipulation task usually concerns the hand's position and orientation. Forward kinematics connects the two. Starting from a reference frame at the base, it composes the transformations along the links to the [end-effector](https://xumanoids.com/glossary/end-effector). [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/4-1-1-product-of-exponentials-formula-in-the-space-frame/) develops this mapping with the product-of-exponentials formula: each joint contributes motion about its screw axis, combined with the robot's home pose. [Denavit-Hartenberg parameters](https://xumanoids.com/glossary/denavit-hartenberg-parameters) provide another way to organize the link transforms. ### The reference frame is part of the result A hand pose relative to the torso is different from its pose relative to the floor. For a moving robot, the base transform must be included when expressing the hand in world coordinates. [MIT's manipulation notes](https://manipulation.mit.edu/pick.html) demonstrate this by composing transforms through a kinematic tree to the world frame. The same process can locate an elbow, camera, or tool frame, provided its connection to the modeled links is known. ### What the calculation leaves out Forward kinematics uses geometry. It does not calculate the forces required to move, which belong to [dynamics](https://xumanoids.com/glossary/inverse-dynamics). The reverse problem, finding joint positions for a requested pose, is [inverse kinematics](https://xumanoids.com/glossary/inverse-kinematics). ### Sources - [Modern Robotics: Product of Exponentials Formula in the Space Frame](https://modernrobotics.northwestern.edu/nu-gm-book-resource/4-1-1-product-of-exponentials-formula-in-the-space-frame/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Friction cone https://xumanoids.com/glossary/friction-cone A friction cone is the set of contact forces permitted by a Coulomb friction model at a contact that pushes but does not pull. The allowable tangential force magnitude is bounded by the normal force multiplied by a friction coefficient. Updated: 2026-10-05 Also known as: Coulomb friction cone, Friction cones ### Normal force limits tangential force For a stationary point contact, write the model as `norm(f_t) <= mu * f_n`, with nonnegative normal force `f_n`. Here `f_t` is the force parallel to the surface and `mu` is the friction coefficient. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-1-friction/) explains why the set forms a cone around the contact normal. A robot foot pushing harder into the floor can, within this model, transmit more sideways force before sliding. The same idea applies to fingers squeezing an object during [manipulation](https://xumanoids.com/glossary/robotic-manipulation). ### Using the cone in planning Walking and grasp planners can require contact forces to remain inside the cone. For computation, a circular cone is often approximated by a pyramid with a finite number of edges. An inner approximation sacrifices some available forces; a coarse or incorrectly oriented approximation can change the predicted contact capability. ### The contact model is approximate [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-1-friction/) calls Coulomb friction an empirical approximation. Real friction can depend on material, surface condition, and whether contact is sticking or sliding. When sliding occurs, the friction direction opposes slip and does not remain an arbitrary force inside the cone. Adhesion, soft contact patches, and torsional friction need additional modeling beyond the basic point-contact cone. ### Sources - [Modern Robotics: Friction](https://modernrobotics.northwestern.edu/nu-gm-book-resource/12-2-1-friction/) --- ## Generalist robot policy https://xumanoids.com/glossary/generalist-robot-policy A generalist robot policy is a learned action-selection model designed to perform multiple tasks across a range of robot settings. Its generality depends on the tasks, observations, action interfaces, and robot bodies included in training and evaluation. Updated: 2026-10-05 Also known as: Generalist robotic policy ### One policy covers a range of tasks A policy maps observations and task information to actions. A generalist policy shares this mapping across tasks instead of assigning a separately trained model to every task. [Octo](https://octo-models.github.io/) accepts language instructions or goal images and uses a shared model for multiple manipulation settings. For a robot arm, the same model might receive different goals for picking up an object, moving it to a container, or inserting a part. The instruction or goal distinguishes the requested behavior. ### Generality has several dimensions Task variety, unfamiliar objects, changed cameras, and a different robot body are separate challenges. Octo's authors distinguish direct evaluation in training-related setups from fine-tuning to new observations and action spaces. Success on a new task with the same arm does not by itself demonstrate transfer to a humanoid hand. ### Relationship to foundation models The [pi0 paper](https://arxiv.org/abs/2410.24164) uses generalist policy and [robot foundation model](https://xumanoids.com/glossary/robot-foundation-model) as overlapping terms. A useful distinction is that generalist describes a policy's intended scope, while foundation describes its role as a reusable pretrained model. Read reported results together with the supported action representation, adaptation data, and evaluated tasks. The label alone does not specify how many tasks the policy can reliably perform. ### Sources - [Octo: An Open-Source Generalist Robot Policy](https://octo-models.github.io/) - [Physical Intelligence: pi0, a vision-language-action flow model](https://arxiv.org/abs/2410.24164) --- ## Hierarchical reinforcement learning https://xumanoids.com/glossary/hierarchical-reinforcement-learning Hierarchical reinforcement learning organizes learned decision-making into levels, often with a higher-level policy selecting goals or skills and lower-level policies producing actions. The levels can operate over different time scales. Updated: 2026-10-05 Also known as: HRL ### Separate goals from detailed actions The [HIRO research](https://arxiv.org/abs/1805.08296) studies a hierarchy in which a higher-level controller proposes learned goals and a lower-level controller learns behavior conditioned on those goals. For a mobile manipulator, a conceptual hierarchy might first choose a nearby location to reach and then produce the movements needed to reach it. This is an example of the division of responsibility, not a claim that HIRO demonstrated that exact physical platform. ### Levels are learned together A hierarchy can reduce the burden on one policy to decide both long-term progress and every immediate action. HIRO uses off-policy experience to train both levels and evaluates complex behaviors in simulated robotic tasks. The method also illustrates a difficulty: when the lower-level policy changes, the same high-level goal can lead to different behavior. HIRO introduces a correction to account for this change when reusing past experience. ### Hierarchy alone is not reinforcement learning A hand-written task tree or a language model that calls fixed robot skills is hierarchical, but that alone does not make it hierarchical [reinforcement learning](https://xumanoids.com/glossary/reinforcement-learning). The term concerns reward-based learning within the hierarchy. The chosen goals, skill interfaces, and termination conditions determine what the levels can express. A useful hierarchy for one robot task may not provide the right structure for a different task or body. ### Sources - [HIRO: Data-Efficient Hierarchical Reinforcement Learning](https://arxiv.org/abs/1805.08296) --- ## Homogeneous transformation https://xumanoids.com/glossary/homogeneous-transformation In rigid-body robotics, a homogeneous transformation is a 4-by-4 matrix that combines a three-dimensional rotation and translation. It represents a pose or changes coordinates between reference frames. Updated: 2026-10-05 Also known as: Homogeneous transformation matrix ### Rotation and translation in one matrix The upper-left 3-by-3 block is a [rotation matrix](https://xumanoids.com/glossary/rotation-matrix). The upper-right column stores a translation. The final row is zero, zero, zero, one. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-1-homogeneous-transformation-matrices/) uses this structure to represent an oriented body frame relative to a reference frame. Appending a one to a point's three coordinates allows the matrix to apply rotation and translation in a single multiplication. A direction vector has no position offset, so its homogeneous coordinate is zero instead. ### Composing a chain of frames Suppose a robot knows the camera pose relative to the torso and the torso pose relative to the world. Multiplying the transforms in the matching frame order gives the camera pose in the world. Taking the inverse reverses the coordinate relationship. Multiplication order matters. Rotating then translating generally gives a different result from translating then rotating, and applying a transform on the left or right changes which frame describes the operation. ### The representation has constraints For rigid motion, the rotation block must be orthonormal with determinant positive one. An arbitrary 4-by-4 matrix is not a valid rigid transform. The geometric operation is a [rigid-body transformation](https://xumanoids.com/glossary/rigid-body-transformation); homogeneous coordinates are the matrix representation used to calculate it. ### Sources - [Modern Robotics: Homogeneous Transformation Matrices](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-1-homogeneous-transformation-matrices/) --- ## Humanoid robot https://xumanoids.com/glossary/humanoid-robot A humanoid robot is a robot with a body arranged to resemble the human form, usually with a torso, arms, and legs. The term describes its physical form and does not by itself establish human-level intelligence or general autonomy. Updated: 2026-10-05 Also known as: Humanoid, Humanoid robots ### Form and capability are separate Humanlike limbs give a robot a familiar arrangement for moving and handling objects. [Boston Dynamics describes Atlas](https://bostondynamics.com/atlas/) as a humanoid robot designed for industrial material handling. That application says more about the intended work than the word humanoid alone. A humanoid's body does not explain how it chooses actions. Its movements may depend on a controller, a learned policy, a human operator, or a combination. Evaluate a specific robot by its documented tasks and operating conditions rather than its resemblance to a person. ### Humanoid robots and embodied AI [Embodied AI](https://xumanoids.com/glossary/embodied-ai) can help a robot connect observations to actions, but it is not limited to humanoid bodies. [Google DeepMind's Gemini Robotics work](https://deepmind.google/blog/gemini-robotics-brings-ai-into-the-physical-world/) describes experiments on two-arm platforms and adaptation to the humanoid Apollo robot. This site's coined term [xumanoid](https://xumanoids.com/glossary/xumanoid) adds a particular idea of context-aware, learning, collaborative behaviour to the humanoid form. It is not a standard robot classification. ### Sources - [Boston Dynamics: Atlas humanoid robot](https://bostondynamics.com/atlas/) - [Google DeepMind: Gemini Robotics brings AI into the physical world](https://deepmind.google/blog/gemini-robotics-brings-ai-into-the-physical-world/) --- ## Hybrid position-force control https://xumanoids.com/glossary/hybrid-position-force-control Hybrid position-force control regulates motion in some task directions and contact force in complementary constrained directions. It separates the commands according to the motion and force freedoms permitted by the environment. Updated: 2026-10-05 Also known as: Hybrid motion-force control, Hybrid force-position control ### Assigning motion and force directions When a robot wipes a rigid tabletop, it can follow a path along the surface while regulating the force pressing into it. The tangential directions allow motion; the normal direction provides a contact constraint. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-6-hybrid-motion-force-control/) develops controllers that project motion and force commands into the appropriate task subspaces. The separation is defined by the contact geometry, rather than by a permanent choice of world axes. Opening a hinged door, for example, has a permitted rotational motion that changes the relevant hand motion as the door swings. ### Avoiding conflicting commands Trying to impose both an exact rigid position and an independent force in the same constrained direction can create incompatible objectives. Hybrid control instead assigns complementary directions to the two controllers. The resulting desired [wrench](https://xumanoids.com/glossary/wrench) is mapped into joint commands through the [Jacobian](https://xumanoids.com/glossary/robot-jacobian). ### Contact assumptions matter The textbook formulation assumes known rigid constraints and particular motion freedoms. If the surface position or normal is wrong, or the hand loses contact, the selected subspaces may no longer describe the interaction. [Impedance control](https://xumanoids.com/glossary/impedance-control) offers a different way to manage contact by prescribing a compliant force-motion relationship. A practical task may combine these ideas, but they describe different control objectives. ### Sources - [Modern Robotics: Hybrid Motion-Force Control](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-6-hybrid-motion-force-control/) --- ## Imitation learning https://xumanoids.com/glossary/imitation-learning Imitation learning learns behavior from examples supplied by a demonstrator. In robotics, demonstrations can teach a policy how to perform a task without requiring every action or objective to be programmed by hand. Updated: 2026-10-05 Also known as: Learning from demonstration, Learning from demonstrations, LfD ### Demonstrations provide the learning signal [An Algorithmic Perspective on Imitation Learning](https://arxiv.org/abs/1811.06711) describes learning from demonstrations as an alternative to manually engineering complex behavior. Demonstrations provide evidence about how a task should be performed, although different algorithms use that evidence differently. For example, a person can use [teleoperation](https://xumanoids.com/glossary/teleoperation) to guide two robot arms through an insertion task. The [ALOHA and ACT project](https://tonyzhaozh.github.io/aloha/) records real demonstrations and trains a policy to reproduce related behaviors. ### Behavior cloning is one method [Behavior cloning](https://xumanoids.com/glossary/behavior-cloning) directly trains action predictions from demonstrated observations and actions. Imitation learning is the broader field; it also includes interactive approaches and methods that use demonstrations to define a learning objective. A motion-imitation system can even use [reinforcement learning](https://xumanoids.com/glossary/reinforcement-learning) to follow a reference trajectory, as in [Peng and colleagues' locomotion work](https://arxiv.org/html/2004.00784v2). Imitation and reinforcement learning therefore need not be mutually exclusive. ### Reproduction depends on data and embodiment The demonstrator's motions, the robot's observations, and the robot's possible actions must be connected. Human motion may require [retargeting](https://xumanoids.com/glossary/motion-retargeting), while robot demonstrations can still omit recovery from errors. A successful replay or training example does not establish general task competence. Evaluation should test the learned policy under the conditions in which it will actually act. ### Sources - [An Algorithmic Perspective on Imitation Learning](https://arxiv.org/abs/1811.06711) - [Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware](https://tonyzhaozh.github.io/aloha/) - [Learning Agile Robotic Locomotion Skills by Imitating Animals](https://arxiv.org/html/2004.00784v2) --- ## Impedance control https://xumanoids.com/glossary/impedance-control Impedance control shapes the dynamic relationship between a robot’s motion and the forces it exchanges with its environment. A common goal is for the robot to respond like a chosen mass, spring, and damper at a joint or end effector. Updated: 2026-10-05 Also known as: Robot impedance control ### Programming how the robot yields Instead of insisting on an exact hand position during contact, an impedance controller can make position error produce a spring-like restoring force and motion produce damping. [MIT's manipulation notes](https://manipulation.mit.edu/force.html) explain this using a virtual spring, mass, and damper. Setting only stiffness and damping is often called stiffness control, a subset of the broader idea. For example, a robot inserting a peg can allow lateral motion when the peg touches the hole's edge. The desired response depends on the task: yielding in one direction can help alignment while stiffness in another direction maintains useful pressure. ### Relation to force and admittance control [Force control](https://xumanoids.com/glossary/force-control) targets a contact force directly. Impedance control specifies how force and displacement interact, so the resulting force also depends on the environment. A common torque-based implementation calculates joint commands from motion error. [Admittance control](https://xumanoids.com/glossary/admittance-control) commonly implements the complementary arrangement: measured force drives a motion reference that another loop tracks. Both can aim for a similar external mechanical response. ### Performance depends on the implementation Selected stiffness is not a promise of safe contact. Sensor quality, actuator bandwidth, sampling, and transmission dynamics limit the behavior the robot can realize. [MIT's discussion](https://manipulation.mit.edu/force.html) emphasizes these details and the role of passivity when reasoning about contact stability. ### Sources - [MIT Robotic Manipulation: Force Control](https://manipulation.mit.edu/force.html) --- ## Inertial measurement unit https://xumanoids.com/glossary/inertial-measurement-unit An inertial measurement unit is a sensor assembly that typically combines accelerometers and gyroscopes to measure specific force and angular velocity. Some devices also provide magnetometer readings or estimated orientation. Updated: 2026-10-05 Also known as: IMU ### Measurements expressed in sensor axes Gyroscopes measure rotational rate. Accelerometers measure specific force, which differs from a gravity-free acceleration estimate. [ROS REP 145](https://www.ros.org/reps/rep-0145.html), a draft convention for IMU drivers, explains that a stationary supported sensor reports a gravity-related accelerometer reading. For a humanoid, an IMU mounted in the torso can contribute rapid motion measurements to a body [state estimator](https://xumanoids.com/glossary/state-estimation). The sensor's mounting orientation determines how its axes relate to the robot. ### Orientation can be an estimate A device may calculate orientation by fusing several measurements internally. It may instead output only raw acceleration and angular velocity. The [ROS Imu message specification](https://docs.ros.org/en/rolling/p/sensor_msgs/msg/Imu.html) explicitly allows an orientation estimate to be unavailable and provides covariance fields for the reported measurements. ### A sensor is not a complete positioning system An IMU does not directly measure absolute position. Coordinate conventions and uncertainty must be understood before combining its output with other sensors. In [visual-inertial odometry](https://xumanoids.com/glossary/visual-inertial-odometry), camera observations provide additional constraints while the estimator accounts for inertial bias and motion. ### Sources - [ROS REP 145: Conventions for IMU Sensor Drivers](https://www.ros.org/reps/rep-0145.html) - [ROS sensor_msgs: Imu Message Definition](https://docs.ros.org/en/rolling/p/sensor_msgs/msg/Imu.html) --- ## Inverse dynamics https://xumanoids.com/glossary/inverse-dynamics Inverse dynamics calculates the joint forces or torques required for specified joint positions, velocities, and accelerations under a dynamics model. The result also depends on gravity and specified external loading. Updated: 2026-10-05 Also known as: Robot inverse dynamics ### From desired motion to actuator effort A robot following a trajectory needs forces that accelerate its links and account for gravity and motion-dependent effects. Inverse dynamics calculates those model-based efforts from a desired motion state. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-3-newton-euler-inverse-dynamics/) presents the recursive Newton-Euler method for an open chain. A forward pass calculates link velocities and accelerations. A backward pass propagates the required [wrenches](https://xumanoids.com/glossary/wrench) toward the base and extracts the joint efforts. ### External loading belongs in the problem Holding an object or pushing on the environment changes the forces the robot must supply. The reference algorithm therefore includes an end-effector wrench as an input alongside the joint motion and gravity. These calculations need inertial properties as well as link geometry. A [forward-kinematics](https://xumanoids.com/glossary/forward-kinematics) model alone cannot predict the required torque. ### A model prediction is not perfect tracking The calculated effort is useful for feedforward [torque control](https://xumanoids.com/glossary/torque-control), but it is only as complete as the dynamics model. The [forward-dynamics lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-5-forward-dynamics-of-open-chains/) shows how omitted friction changes simulated behavior. [Inverse kinematics](https://xumanoids.com/glossary/inverse-kinematics) instead solves for joint positions from a target pose. [Forward dynamics](https://xumanoids.com/glossary/forward-dynamics) reverses the dynamics question by predicting acceleration from applied effort. ### Sources - [Modern Robotics: Newton-Euler Inverse Dynamics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-3-newton-euler-inverse-dynamics/) - [Modern Robotics: Forward Dynamics of Open Chains](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-5-forward-dynamics-of-open-chains/) --- ## Inverse kinematics https://xumanoids.com/glossary/inverse-kinematics Inverse kinematics finds joint positions that produce a desired robot end-effector position, orientation, or other geometric task. A target can have multiple solutions, no solution, or a continuous family of solutions. Updated: 2026-10-05 Also known as: IK ### Finding a posture for a target A command such as placing a gripper at a handle specifies a task in Cartesian space. IK seeks a robot configuration whose [forward kinematics](https://xumanoids.com/glossary/forward-kinematics) matches that target. The task may constrain only position or may also constrain orientation. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/inverse-kinematics-of-open-chains/) illustrates why the answer need not be unique: an arm can reach the same point with different elbow postures. A target outside the [workspace](https://xumanoids.com/glossary/workspace) has no solution. ### Analytical and numerical methods An analytical solver uses formulas derived for a particular mechanism. A numerical solver iteratively updates an initial guess to reduce pose error, often using a [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian). The [numerical IK lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/6-2-numerical-inverse-kinematics-part-1-of-2/) explains this through Newton-Raphson iteration. The initial guess affects which solution an iterative method finds and whether it converges. A failed numerical solve alone does not prove that the target is unreachable. ### A pose is not a complete motion IK does not automatically provide a collision-free path to the result. Joint limits and other constraints must be included where relevant. [MIT's constrained differential IK formulation](https://manipulation.mit.edu/pick.html) explicitly accounts for position, velocity, and acceleration limits. Planning and control still determine how the robot gets to and tracks the chosen posture. ### Sources - [Modern Robotics: Inverse Kinematics of Open Chains](https://modernrobotics.northwestern.edu/nu-gm-book-resource/inverse-kinematics-of-open-chains/) - [Modern Robotics: Numerical Inverse Kinematics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/6-2-numerical-inverse-kinematics-part-1-of-2/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Jerk https://xumanoids.com/glossary/jerk Jerk is the rate at which acceleration changes with time, or the third time derivative of position. Robotics uses jerk limits to constrain how abruptly a commanded motion changes acceleration. Updated: 2026-10-05 ### How quickly acceleration changes A joint can have modest acceleration yet change that acceleration abruptly. Jerk describes this change. For a linear coordinate, its units are metres per second cubed; for a joint angle, radians per second cubed. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/9-1-and-9-2-point-to-point-trajectories-part-2-of-2/) contrasts a trapezoidal velocity profile, whose acceleration jumps between phases, with an S-curve profile that changes acceleration through finite-jerk segments. ### Velocity, acceleration, and jerk need separate limits Limiting acceleration does not bound the rate at which acceleration changes. Conversely, zero jerk can accompany a large constant acceleration. These are different constraints on a trajectory. The [Ruckig motion generator](https://docs.ruckig.com/) accepts separate velocity, acceleration, and jerk constraints when calculating motion between states defined by position, velocity, and acceleration. This makes jerk a practical parameter in trajectory generation, including motions commanded through [joint-space control](https://xumanoids.com/glossary/joint-space-control). ### A jerk limit does not specify every aspect of motion A piecewise-constant jerk profile can keep acceleration continuous while jerk itself changes at segment boundaries. “Jerk-limited” therefore does not necessarily mean jerk-continuous. The chosen coordinate also matters: a bound on a joint's angular jerk is not automatically the same as a bound on Cartesian hand jerk. Robot geometry and the full motion determine that relationship. ### Sources - [Modern Robotics: Point-to-Point Trajectories, Part 2](https://modernrobotics.northwestern.edu/nu-gm-book-resource/9-1-and-9-2-point-to-point-trajectories-part-2-of-2/) - [Ruckig: Motion Generation Documentation](https://docs.ruckig.com/) --- ## Joint https://xumanoids.com/glossary/joint A joint is a connection between robot links that constrains their permitted relative motion. Its kinematic type determines which rotations or translations the connected links can make relative to one another. Updated: 2026-10-05 Also known as: Robot joint, Joints ### Connections define permitted motion A robot arm's links form a [kinematic chain](https://xumanoids.com/glossary/kinematic-chain), with joints constraining how adjacent links move. A [revolute joint](https://xumanoids.com/glossary/revolute-joint) permits rotation about one axis. A [prismatic joint](https://xumanoids.com/glossary/prismatic-joint) permits translation along one axis. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) describes these as one-degree-of-freedom joints. A universal joint permits two relative freedoms, while a spherical joint permits three rotational freedoms. One joint therefore does not always mean one independent coordinate. ### A joint need not have its own motor Joint type describes motion constraints, not the source of actuation. The reference's Stewart-platform example has legs containing universal, prismatic, and spherical joints, with the prismatic joints actuated to move the platform. This distinction matters when reading a humanoid specification: the number of modeled joints, actuated axes, and independent [degrees of freedom](https://xumanoids.com/glossary/degrees-of-freedom) need not be interchangeable counts. ### The complete mechanism adds constraints A joint's permitted relative motion is only part of the model. Closing a loop between links can make several joint coordinates dependent on one another. Joint travel limits also restrict the configurations available, although a finite travel range does not by itself remove a degree of freedom within that range. ### Sources - [Modern Robotics: Degrees of Freedom of a Robot](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) --- ## Joint-space control https://xumanoids.com/glossary/joint-space-control Joint-space control expresses a robot's motion targets and tracking errors in joint coordinates, such as joint angles or linear displacements. It regulates those coordinates rather than defining the primary motion error directly at the end-effector. Updated: 2026-10-05 Also known as: Joint space control ### Track a posture or joint trajectory An arm controller can compare desired shoulder and elbow angles with measured angles, then command motion that reduces the differences. The targets may be a fixed posture or a sequence of joint positions over time. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-3-motion-control-with-velocity-inputs-part-1-of-3/) develops this idea using a joint-position error and a velocity command. Feedback corrects deviations that simply replaying a planned command would leave unresolved. ### Coordinates and actuator commands are different choices “Joint-space” identifies how the motion objective is expressed. It does not require one particular actuator interface. A controller may request joint velocities from lower-level loops, or calculate forces and torques using robot dynamics. The [control-system overview](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-1-control-system-overview/) shows motion feedback feeding a controller that requests joint torques. [Torque control](https://xumanoids.com/glossary/torque-control) describes regulating actuator effort, so it is not a synonym for joint-space control. ### Compare with an end-effector objective Task-space control expresses the motion error in variables such as hand position and orientation. The [task-space control lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-3-motion-control-with-velocity-inputs-part-3-of-3/) converts a desired hand velocity into joint velocities through a [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian). Both approaches ultimately move joints. The distinction is where the primary motion error is defined; joint limits and actuator limits still need explicit handling. ### Sources - [Modern Robotics: Motion Control with Velocity Inputs, Part 1](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-3-motion-control-with-velocity-inputs-part-1-of-3/) - [Modern Robotics: Motion Control with Velocity Inputs, Part 3](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-3-motion-control-with-velocity-inputs-part-3-of-3/) - [Modern Robotics: Control System Overview](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-1-control-system-overview/) --- ## Kalman filter https://xumanoids.com/glossary/kalman-filter A Kalman filter is a recursive estimator that predicts a system's state with a linear model and corrects that prediction using noisy measurements. It tracks both the estimate and its error covariance. Updated: 2026-10-05 Also known as: KF ### Prediction followed by measurement correction The filter first predicts the next state and uncertainty. It then compares an observation with the predicted measurement and applies a correction weighted by the Kalman gain. [Kalman's original paper](https://www.cs.cmu.edu/~motionplanning/papers/sbp_papers/k/Kalman1960.pdf) derives the recursive linear filtering formulation and the evolution of estimation-error covariance. For a simple robot tracking problem, the state might contain position and velocity while a sensor measures only position. The motion model connects the unmeasured velocity to later position observations. ### Uncertainty controls the weighting The gain depends on predicted uncertainty and measurement noise. A measurement assigned high uncertainty receives less influence than it otherwise would. [Welch and Bishop's tutorial](https://www.cs.cmu.edu/~motionplanning/papers/sbp_papers/kalman/welch_intro_kalman.pdf) explains the roles of process and measurement covariance in this calculation. ### The assumptions matter The familiar exact Gaussian interpretation assumes linear dynamics and observations with appropriate Gaussian noise. Incorrect noise assumptions can make reported confidence misleading. Nonlinear robot models usually require an extension or another estimator; an [extended Kalman filter](https://xumanoids.com/glossary/extended-kalman-filter) uses local linearization. Applying the word Kalman to a filter does not establish accuracy under arbitrary motion or sensing conditions. ### Sources - [R. E. Kalman: A New Approach to Linear Filtering and Prediction Problems](https://www.cs.cmu.edu/~motionplanning/papers/sbp_papers/k/Kalman1960.pdf) - [Welch and Bishop: An Introduction to the Kalman Filter](https://www.cs.cmu.edu/~motionplanning/papers/sbp_papers/kalman/welch_intro_kalman.pdf) --- ## Kinematic chain https://xumanoids.com/glossary/kinematic-chain A kinematic chain is an arrangement of links connected by joints that constrains their relative motion. Open chains have no closed link loop, while closed chains contain at least one loop. Updated: 2026-10-05 Also known as: Kinematic chains ### Links and joints form the mechanism A link is a modeled rigid body. A joint specifies how one link can move relative to another. An arm with shoulder, elbow, and wrist joints is commonly modeled as a serial chain from a base to an [end-effector](https://xumanoids.com/glossary/end-effector). [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) uses the example of a three-revolute-joint arm to show how an open chain has one path from its base to its endpoint. A branched robot can instead be represented as a tree of links. ### Closing a loop adds constraints Pinning that arm's endpoint to ground creates a closed chain. The resulting loop means that the joint coordinates cannot all vary independently. Four-bar linkages and parallel platforms are examples. The [closed-chain kinematics lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/kinematics-of-closed-chains/) explains that loop constraints make both configuration analysis and solution counting more complicated than in a simple serial arm. ### A model of geometry and connectivity A chain supplies the structure needed for [forward kinematics](https://xumanoids.com/glossary/forward-kinematics). [MIT's manipulation notes](https://manipulation.mit.edu/pick.html) describe finding a tool pose by following its parent links and composing their transforms. The chain alone does not specify how the joints are driven or the masses and forces needed for a dynamics calculation. ### Sources - [Modern Robotics: Degrees of Freedom of a Robot](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) - [Modern Robotics: Kinematics of Closed Chains](https://modernrobotics.northwestern.edu/nu-gm-book-resource/kinematics-of-closed-chains/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Kinematic singularity https://xumanoids.com/glossary/kinematic-singularity A kinematic singularity is a robot configuration where the task Jacobian has lower rank than the maximum it can attain for that mechanism and task. At that configuration the robot loses one or more instantaneous task-motion directions. Updated: 2026-10-05 Also known as: Robot singularity, Kinematic singularities ### A stretched arm loses a motion direction For a planar arm with its links fully aligned, the joints can initially move the tip perpendicular to the arm but cannot produce every tip-velocity direction. The [Modern Robotics singularity lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) demonstrates this rank loss with two- and three-joint arms. The limitation is instantaneous. It does not mean the mechanism can never move inward or reach another configuration after first bending its joints. ### Rank matters more than a determinant A [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian) need not be square. Comparing its rank with the maximum achievable for the selected task handles redundant and lower-mobility robots as well as square Jacobians. A robot designed for a two-dimensional task is not automatically singular because it cannot perform arbitrary six-dimensional motion. Singularity concerns a loss relative to that mechanism's normal capability. ### Near a singularity also matters Even before exact rank loss, some task velocities can require very large joint velocities. [MIT's manipulation notes](https://manipulation.mit.edu/pick.html) connect this behavior to small singular values and discuss constrained differential IK. This is different from gimbal lock in [Euler angles](https://xumanoids.com/glossary/euler-angles), which is a coordinate-representation singularity. Switching to [quaternions](https://xumanoids.com/glossary/quaternion) fixes that representation issue but cannot restore a motion direction lost through the robot's joint geometry. ### Sources - [Modern Robotics: Singularities](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Language-conditioned policy https://xumanoids.com/glossary/language-conditioned-policy A language-conditioned policy selects actions using a language instruction together with observations. The instruction specifies or modifies the behavior requested from the policy. Updated: 2026-10-05 Also known as: Language conditioned policy ### Language specifies the requested behavior [CLIPort](https://cliport.github.io/) learns manipulation from visual input and language goals. An instruction can distinguish picking a particular object, placing it in a named region, or arranging items according to a described relation. The policy must connect those words to both scene information and actions. It does not suffice to output a fluent description of the requested task. ### Conditioning does not require one architecture A policy can use a language embedding together with robot observations. [Octo](https://octo-models.github.io/) supports both language instructions and goal images, illustrating that task conditioning is an interface choice rather than a synonym for a large language model. A [vision-language-action model](https://xumanoids.com/glossary/vision-language-action-model) is a closely related model type when visual input and language lead to robot action output. ### Skill selection is a distinct use of language [SayCan](https://say-can.github.io/) uses language-model knowledge and skill-feasibility estimates to choose from robot skills. Selecting a skill and predicting its detailed motor commands are different levels of decision-making, even when both use language. Evaluate instruction following through physical outcomes and the actual range of supported commands. A policy may respond correctly to familiar wording while still failing on unfamiliar objects, spatial relations, or tasks that require abilities absent from its training data. ### Sources - [CLIPort: What and Where Pathways for Robotic Manipulation](https://cliport.github.io/) - [Octo: An Open-Source Generalist Robot Policy](https://octo-models.github.io/) - [SayCan: Grounding Language in Robotic Affordances](https://say-can.github.io/) --- ## Lidar https://xumanoids.com/glossary/lidar Lidar measures distance using emitted laser light and its return from surfaces. Repeated range measurements across directions can form a spatial scan or three-dimensional point cloud. Updated: 2026-10-05 Also known as: LiDAR Light detection and ranging ### Light provides a distance measurement A pulsed time-of-flight lidar emits light and measures the delay before the reflection returns. Distance follows from the light's round-trip travel time. [Ouster's technical introduction](https://ouster.com/insights/what-is-lidar) describes this measurement cycle and how many returns build a representation of the surroundings. A robot can use the resulting [point cloud](https://xumanoids.com/glossary/point-cloud) for mapping or obstacle detection. Those functions require additional software to interpret the measurements. ### Field of view and resolution differ Lidar is a sensing principle, not a promise of complete coverage. Ouster's explanation distinguishes directional sensors from devices with a full horizontal field of view and notes that placement can leave gaps. The number and arrangement of measurements affect what detail the scan contains. For a humanoid, sensor placement needs to consider the space its arms and legs can enter, as well as what lies directly ahead. ### Measurements do not choose actions The same source explicitly separates sensing from driving decisions. A lidar can supply distances, but [motion planning](https://xumanoids.com/glossary/motion-planning) and control must turn perception into movement. A successful scan alone does not establish safe navigation or reliable object recognition. ### Sources - [Ouster: What Is Lidar and How Does It Work?](https://ouster.com/insights/what-is-lidar) --- ## Linear inverted pendulum model https://xumanoids.com/glossary/linear-inverted-pendulum-model The linear inverted pendulum model approximates a walking robot by a mass moving at constant height above its support, with simplified angular-momentum dynamics. These assumptions make horizontal center-of-mass acceleration linear in the displacement from the support point. Updated: 2026-10-05 Also known as: LIPM, Linear inverted pendulum ### A simpler model for a complicated robot A humanoid may have many joints, but a walking planner can first reason about its [center of mass](https://xumanoids.com/glossary/center-of-mass). With constant height `h` and negligible change in angular momentum, the horizontal dynamics become `x_ddot = (g / h) * (x - p)`, where `p` is the [zero-moment point](https://xumanoids.com/glossary/zero-moment-point) on flat ground. [MIT's derivation](https://underactuated.mit.edu/humanoids.html) explains the assumptions behind this reduction. If the mass moves ahead of a fixed support point, gravity-driven motion accelerates it farther forward. Moving the support point changes that acceleration. The model is an inverted pendulum because its mass lies above the support. ### Uses in walking control The linear equations make center-of-mass trajectory planning and [model predictive control](https://xumanoids.com/glossary/model-predictive-control) easier to compute. They also lead to the [capture-point](https://xumanoids.com/glossary/capture-point) expression used in balance recovery. ### What the approximation leaves out The basic model does not represent swing-leg dynamics, every joint limit, or the full effects of changing body angular momentum. Footstep reachability still needs separate constraints. [Capture-region research](https://arxiv.org/abs/2307.11968) shows why timing and reachable step locations matter even when reduced dynamics give a mathematically attractive target. Jumping or large vertical motion requires a different or extended model. ### Sources - [MIT Underactuated Robotics: Highly-articulated Legged Robots](https://underactuated.mit.edu/humanoids.html) - [Griffin et al.: Reachability Aware Capture Regions with Time Adjustment and Cross-Over for Step Recovery](https://arxiv.org/abs/2307.11968) --- ## Manipulability https://xumanoids.com/glossary/manipulability Manipulability describes how a robot configuration maps joint motion into end-effector motion in different directions. It is commonly represented by a Jacobian-based velocity ellipsoid or summarized by a scalar measure. Updated: 2026-10-05 Also known as: Kinematic manipulability ### Visualizing directional motion capability Imagine all joint-velocity vectors with the same chosen magnitude. Passing them through the [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian) produces an ellipsoid of end-effector velocities. Its long directions correspond to larger achievable task velocities under that joint-speed normalization. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-4-manipulability/) illustrates the idea with a planar arm: a circular set of joint velocities maps to an ellipse of tip velocities. Near a [kinematic singularity](https://xumanoids.com/glossary/kinematic-singularity), an axis of the ellipse shrinks. ### Different measures answer different questions The ellipsoid's axis ratio describes directional imbalance. Its volume describes an aggregate motion capability. A smallest-axis measure emphasizes the weakest direction. These quantities are related but are not interchangeable rankings of robot quality. Manipulability can help compare candidate postures for a task, including alternatives available to a [redundant manipulator](https://xumanoids.com/glossary/redundant-manipulator). ### Units and assumptions affect the result Linear and angular velocities have different units. Modern Robotics therefore discusses separate translational and rotational ellipsoids. Any combined metric needs an explicit scaling choice, especially when comparing mechanisms or mixing revolute and prismatic joints. A kinematic ellipsoid does not by itself account for collisions, payload, motor torque limits, or control errors. It describes the selected local motion map, not overall task success or dexterity in every operating condition. ### Sources - [Modern Robotics: Manipulability](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-4-manipulability/) --- ## Model predictive control https://xumanoids.com/glossary/model-predictive-control Model predictive control repeatedly optimizes future actions using a system model, applies the next part of the solution, and replans from updated state information. It can account for objectives and constraints over a finite prediction horizon. Updated: 2026-10-05 Also known as: MPC, Receding-horizon control ### Optimizing again as the robot moves At each update, the controller predicts how candidate actions will change the robot's state over the next several time steps. It chooses a sequence that minimizes a cost while satisfying model and constraint equations, executes the first action or short segment, and solves again with new [state estimates](https://xumanoids.com/glossary/state-estimation). [MIT's trajectory-optimization notes](https://underactuated.mit.edu/trajopt.html) describe this receding-horizon construction. For a humanoid, an optimization might choose contact forces that track a desired walking velocity while respecting friction limits. Other formulations optimize footsteps or full joint trajectories. ### Choosing the prediction model A [linear inverted pendulum model](https://xumanoids.com/glossary/linear-inverted-pendulum-model) is cheaper to optimize than the full multibody dynamics, but describes fewer physical effects. [Wensing and colleagues](https://arxiv.org/abs/2211.11644) explain how simplified models, contact assumptions, and solver choices shape what can run in a control loop. ### Feasibility and timing matter A solution that is feasible now does not automatically ensure that the next optimization will remain feasible. Terminal conditions and other design choices can establish such guarantees under stated assumptions. Model errors and computation deadlines also matter: the robot must receive a useful command before the next control update, and predicted behavior must remain close enough to the real system for replanning to help. ### Sources - [MIT Underactuated Robotics: Trajectory Optimization](https://underactuated.mit.edu/trajopt.html) - [Wensing et al.: Optimization-Based Control for Dynamic Legged Robots](https://arxiv.org/abs/2211.11644) --- ## Motion planning https://xumanoids.com/glossary/motion-planning Motion planning finds a robot movement from an initial state to a goal while satisfying constraints such as collision avoidance. A planner may produce a geometric path, a timed trajectory, or a sequence of controls. Updated: 2026-10-05 Also known as: Robot motion planning ### Plan the robot's movement, not just the hand's destination A hand target does not specify how the elbow, torso, and other links should move around obstacles. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/10-1-overview-of-motion-planning/) formulates motion planning in configuration or state space, with collision and motion constraints. For a reaching task, a planner might search for joint configurations that take a gripper around a shelf edge. A humanoid planning a step may need additional constraints on contact and body dynamics. ### A path and a trajectory answer different questions A geometric path describes where to move. A timed trajectory also specifies when to occupy each state. A path that avoids obstacles may still require timing adjustments to respect velocity, acceleration, or actuator limits. ### Planner guarantees have conditions The cited overview distinguishes complete, resolution-complete, and probabilistically complete planners. Probabilistic completeness concerns success probability as planning time grows under the algorithm's assumptions; it does not guarantee success before a particular deadline. The quality of a plan also depends on the robot and environment models. An obstacle omitted from the model cannot be avoided merely because the modeled path is collision-free. ### Sources - [Modern Robotics: Overview of Motion Planning](https://modernrobotics.northwestern.edu/nu-gm-book-resource/10-1-overview-of-motion-planning/) --- ## Motion retargeting https://xumanoids.com/glossary/motion-retargeting Motion retargeting maps a motion recorded or designed for one body onto another body with different geometry or joints. In robotics it produces a compatible pose or trajectory reference, which still needs a controller to execute it physically. Updated: 2026-10-05 ### Match motion across different bodies [Peng and colleagues](https://arxiv.org/html/2004.00784v2) retarget animal motions to a robot by pairing body keypoints and solving [inverse kinematics](https://xumanoids.com/glossary/inverse-kinematics) so the robot's corresponding points follow the reference. The robot and source animal need not have identical proportions. The same distinction matters when human motion supplies a humanoid's training reference. Copying recorded joint values directly generally does not account for different joint arrangements and limb lengths. ### A reference motion is not an executable policy The animal-imitation framework separates retargeting from a subsequent reinforcement-learning stage that trains a controller to track the reference in simulation. It then addresses transfer to hardware. [ASAP](https://arxiv.org/abs/2502.01143) similarly uses retargeted human motion to pretrain humanoid tracking policies and then uses real-world data to address dynamics mismatch. Its authors evaluate transfer to the Unitree G1. ### Geometry and dynamics impose different limits Retargeting can seek poses that reproduce important features of the source motion. That does not by itself ensure balance, feasible contact forces, or sufficient actuator capability during execution. When assessing a demonstration, distinguish the source recording, retargeted reference, simulated tracking policy, and physical result. They are separate stages, and success at an earlier stage does not establish success at the later one. ### Sources - [Learning Agile Robotic Locomotion Skills by Imitating Animals](https://arxiv.org/html/2004.00784v2) - [ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills](https://arxiv.org/abs/2502.01143) --- ## Null space https://xumanoids.com/glossary/null-space The null space of a matrix is the set of vectors it maps to zero. For a robot task Jacobian, it contains joint velocities that produce no instantaneous motion in the specified task coordinates. Updated: 2026-10-05 Also known as: Nullspace, Matrix kernel ### Joint motion invisible to a task A [redundant manipulator](https://xumanoids.com/glossary/redundant-manipulator) may move its elbow while keeping its hand still. At the current configuration, the associated joint-velocity vector lies in the hand task's Jacobian null space. The word is task-specific. If the task constrains only hand position, null-space motion may change hand orientation. Adding orientation constraints can remove that freedom. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) illustrates how redundancy depends on the selected velocity task. ### Combining primary and secondary objectives [MIT's differential IK treatment](https://manipulation.mit.edu/pick.html) projects a preferred joint motion into the Jacobian null space to help return an arm toward a nominal posture while tracking an end-effector velocity. This supplies a useful building block for task prioritization: the primary task uses visible motion directions, while a secondary objective uses remaining directions. ### An instantaneous result needs continued updating The [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian) changes as the configuration changes. A vector in its null space now is not necessarily in the null space after a finite step. Maintaining the task requires recalculating the relationship as the robot moves. Null-space projection also does not automatically satisfy joint or contact constraints. MIT notes that constraints can make primary and secondary objectives conflict, so a weighted or hierarchical optimization must handle their interaction explicitly. ### Sources - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) - [Modern Robotics: Singularities](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) --- ## Occupancy grid https://xumanoids.com/glossary/occupancy-grid An occupancy grid divides space into cells and records occupancy information for each cell. A two-dimensional robot map commonly distinguishes occupied, free, and unknown regions. Updated: 2026-10-05 Also known as: Occupancy grid map ### A map made of cells An occupancy grid offers a discrete representation that navigation software can inspect. The [ROS OccupancyGrid specification](https://docs.ros.org/en/rolling/p/nav_msgs/msg/OccupancyGrid.html) defines a two-dimensional array with application-dependent cell values, including a conventional unknown value. Numeric encodings should therefore be checked for the producing application. A robot can use such a map to represent walls, open floor, and areas that have not been observed. Unknown space is different from measured free space. ### Resolution sets the scale [MapMetaData](https://docs.ros.org/en/rolling/p/nav_msgs/msg/MapMetaData.html) records resolution in metres per cell, grid dimensions, and the map origin. Fine cells represent smaller features but require more cells to cover the same area. A cell index is not itself a physical coordinate until these metadata are applied. ### Occupancy is only part of planning A grid does not by itself describe a robot's full shape or movement constraints. For a humanoid, a two-dimensional floor map also omits much of the vertical structure relevant to arms and head clearance. [Motion planning](https://xumanoids.com/glossary/motion-planning) needs to interpret the map alongside robot geometry and the constraints of the intended motion. ### Sources - [ROS nav_msgs: OccupancyGrid Message Definition](https://docs.ros.org/en/rolling/p/nav_msgs/msg/OccupancyGrid.html) - [ROS nav_msgs: MapMetaData Message Definition](https://docs.ros.org/en/rolling/p/nav_msgs/msg/MapMetaData.html) --- ## Odometry https://xumanoids.com/glossary/odometry Odometry estimates changes in a robot's position and orientation from motion measurements over time. Its accumulated pose provides a local reference that can drift as measurement errors build up. Updated: 2026-10-05 ### A local account of movement Odometry follows how far and in which direction a robot has moved relative to a starting reference. Wheels can supply encoder-based motion measurements, while cameras and inertial sensors provide other sources. [ROS REP 105](https://www.ros.org/reps/rep-0105.html) explicitly lists wheel odometry, visual odometry, and inertial measurements as inputs to the local odometry frame. For a humanoid, an odometry estimate can tell the navigation system how its body has moved during a walking sequence. The word does not imply that the robot has wheels. ### Continuity and drift serve different needs ROS defines the `odom` frame so the reported pose changes continuously. This makes it useful for short-term motion control. Accumulated errors, however, mean the same frame is unsuitable as a guaranteed long-term global reference. The `map` frame can incorporate global corrections and may jump when a localization estimate changes. Confusing these frames can cause abrupt changes in a controller that expected smooth feedback. ### Odometry is not a complete map [SLAM](https://xumanoids.com/glossary/simultaneous-localization-and-mapping) adds map estimation and constraints from revisiting places. Odometry alone does not establish the layout of a building or identify a safe route through it. ### Sources - [ROS REP 105: Coordinate Frames for Mobile Platforms](https://www.ros.org/reps/rep-0105.html) --- ## Offline reinforcement learning https://xumanoids.com/glossary/offline-reinforcement-learning Offline reinforcement learning learns a reward-optimizing policy from previously collected experience without gathering new environment interactions during that learning stage. The data may come from earlier policies, demonstrations, or other collection procedures. Updated: 2026-10-05 Also known as: Offline RL, Batch reinforcement learning ### Reuse an existing experience dataset The [offline reinforcement-learning tutorial](https://arxiv.org/html/2005.01643v3) defines the setting around learning without additional online data collection. Recorded transitions describe what the agent observed, which action it took, what happened next, and the associated reward. A robotics lab might reuse earlier manipulation trials rather than repeatedly running a new learner on hardware. That changes the learning problem because the dataset cannot automatically supply examples for every action the learner now wants to try. ### Unseen actions can receive misleading values A policy may favor actions for which the dataset contains little support. A learned value function or dynamics model can make optimistic predictions about those actions, and there is no new interaction during offline training to correct the error. The tutorial discusses distribution shift, conservative value estimation, and constraints that keep learned behavior closer to the data as responses to this problem. ### Offline does not mean behavior cloning [Behavior cloning](https://xumanoids.com/glossary/behavior-cloning) learns to reproduce recorded actions. Offline [reinforcement learning](https://xumanoids.com/glossary/reinforcement-learning) tries to optimize return from the available experience, potentially choosing different behavior. Dataset coverage and reward quality limit what can be inferred. Offline training also does not remove the need to evaluate the resulting policy. A later online fine-tuning stage is possible, but it is a separate stage from the offline learning described here. ### Sources - [Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems](https://arxiv.org/html/2005.01643v3) --- ## Operational-space control https://xumanoids.com/glossary/operational-space-control Operational-space control formulates a robot’s motion and force behavior in task coordinates, such as the position and orientation of its hand, while accounting for the robot’s dynamics. Secondary joint objectives can be coordinated with the primary task. Updated: 2026-10-05 Also known as: Operational space control, OSC ### Expressing the task at the hand A manipulation task may specify where the hand should move and how it should respond to contact, rather than an independent target for every joint. Operational-space control uses the hand's task coordinates together with the robot's dynamics to calculate the required actuation. [MIT's manipulation notes](https://manipulation.mit.edu/force.html) derive an effective task-space inertia from the joint-space inertia and [Jacobian](https://xumanoids.com/glossary/robot-jacobian). Desired task forces map to joint torques through the Jacobian transpose. This goes beyond solving [inverse kinematics](https://xumanoids.com/glossary/inverse-kinematics) for a pose because it concerns dynamic behavior and forces. ### Preserving room for secondary objectives A redundant arm can often change its elbow position while holding its hand still. A secondary posture objective can act through a dynamically consistent [null-space](https://xumanoids.com/glossary/null-space) projection so that it does not interfere with the primary task under the model assumptions. ### Conditions for using the model Task-space inertia expressions require an appropriate rank condition; near a [kinematic singularity](https://xumanoids.com/glossary/kinematic-singularity), some task directions lose authority. Model errors, joint limits, and additional contacts also matter. [Whole-body control](https://xumanoids.com/glossary/whole-body-control) extends related ideas to several tasks and constraints, including the feet and floating base of a humanoid. ### Sources - [MIT Robotic Manipulation: Force Control](https://manipulation.mit.edu/force.html) --- ## Particle filter https://xumanoids.com/glossary/particle-filter A particle filter represents a probability distribution over possible states with a collection of weighted samples. It updates those samples using a motion model and new observations to estimate a changing state. Updated: 2026-10-05 Also known as: Particle filtering ### Many hypotheses instead of one estimate Each particle represents a possible state, such as a robot position and orientation. Motion prediction moves the hypotheses forward, and a sensor model assigns greater weight to hypotheses that better explain the observations. [Thrun's robotics paper](https://robots.stanford.edu/papers/thrun.pf-in-robotics-uai02.html) explains this sampling approach and its use in localization and mapping. If several corridors look alike, a filter can retain several groups of possible locations. This is useful when a single mean and covariance would obscure the ambiguity. ### Resampling concentrates the computation Resampling gives more representation to hypotheses with higher weight and removes some unlikely ones. This makes finite computing resources focus on states supported by measurements. It also means that a hypothesis discarded too early may be difficult to recover without an appropriate recovery strategy. ### More dimensions need care Particles do not eliminate the difficulty of high-dimensional [state estimation](https://xumanoids.com/glossary/state-estimation). Thrun discusses methods that exploit problem structure to make large robotics problems tractable. Increasing the number of samples costs computation, and an inaccurate motion or observation model can still mislead the filter even when many particles are available. ### Sources - [Sebastian Thrun: Particle Filters in Robotics](https://robots.stanford.edu/papers/thrun.pf-in-robotics-uai02.html) --- ## PID control https://xumanoids.com/glossary/pid-control PID control is feedback control that combines terms proportional to the current error, the accumulated error, and the rate of change of error. These terms determine the command sent to the controlled system. Updated: 2026-10-05 Also known as: Proportional-integral-derivative control, PID controller Proportional-integral-derivative control ### Three responses to tracking error For a joint-angle target, the proportional term reacts to the angle error now. The integral term accumulates persistent error. The derivative term reacts to how quickly the error changes and can provide damping. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-1-of-3/) derives these terms for a robot joint driven by torque. A joint holding an arm against gravity illustrates their different roles. Proportional control alone may need a nonzero position error to generate the required holding torque. Under suitable conditions, integral action builds up that torque while reducing the steady-state error, as shown in the [gravity example](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-2-of-3/). ### Common robot implementations PD control omits the integral term. Controllers may also combine feedback with a feedforward estimate of gravity or other dynamics. The controlled quantity and output interface matter: a position loop that outputs velocity is different from a position loop that outputs torque. ### More gain is not always better Large gains can amplify sensor errors, excite unmodeled dynamics, or demand unavailable actuator effort. Excessive integral gain can also destabilize even a simplified model. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-2-of-3/) discusses limiting accumulated error and prioritizing stability. Tuning must therefore consider sampling rate, load, actuator limits, and the dynamics of the actual joint. ### Sources - [Modern Robotics: Motion Control with Torque or Force Inputs, Part 1](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-1-of-3/) - [Modern Robotics: Motion Control with Torque or Force Inputs, Part 2](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-2-of-3/) --- ## Point cloud https://xumanoids.com/glossary/point-cloud A point cloud is a collection of points representing sampled locations in space, usually with three-dimensional coordinates. Individual points may also carry attributes such as color or return intensity. Updated: 2026-10-05 Also known as: Point clouds ### Spatial samples from sensors A depth camera or [lidar](https://xumanoids.com/glossary/lidar) can produce measurements that become points in a shared coordinate frame. [Ouster's explanation](https://ouster.com/insights/what-is-lidar) describes this process for laser range measurements. The [Point Cloud Library documentation](https://pointclouds.org/documentation/tutorials/basic_structures.html) describes point clouds as collections whose point type determines the stored data. For a humanoid reaching toward a table, a cloud can represent the visible tabletop and objects on it. These samples can support geometric processing without already identifying which object is a cup or where it should be grasped. ### Organized and unorganized clouds An organized cloud preserves rows and columns, often corresponding to a depth image. An unorganized cloud is simply a collection of samples without that image arrangement. PCL documents both forms and explains why neighboring pixel relationships can make processing organized clouds more efficient. ### A cloud is not a solid model Points do not automatically define connected surfaces, occupied volumes, or object identities. Missing or invalid coordinates also need handling; PCL distinguishes clouds containing only finite data from those with invalid values. Converting a cloud into a mesh, an [occupancy grid](https://xumanoids.com/glossary/occupancy-grid), or a distance field requires additional assumptions and processing. ### Sources - [Point Cloud Library: Getting Started and Basic Structures](https://pointclouds.org/documentation/tutorials/basic_structures.html) - [Ouster: What Is Lidar and How Does It Work?](https://ouster.com/insights/what-is-lidar) --- ## Policy distillation https://xumanoids.com/glossary/policy-distillation Policy distillation trains a student policy to reproduce behavior from one or more teacher policies. It can transfer learned behavior into a smaller network or combine multiple task-specific policies into one model. Updated: 2026-10-05 ### Transfer behavior from teacher to student The [Policy Distillation paper](https://arxiv.org/abs/1511.06295) presents distillation as a way to extract a reinforcement-learning agent's policy into a new network. The training signal comes from the teacher's learned behavior rather than requiring the student to rediscover that behavior through the original training process. In a robotics application, an expensive teacher policy could guide training of a smaller policy intended for a robot's onboard computer. That is a possible application of the method, not a hardware result established by the original paper. ### Compression and consolidation are different aims A student may be smaller than one teacher. Alternatively, it may combine knowledge from several task-specific teachers into a shared policy. The original work studies both aims using Atari tasks. Combining teachers is relevant to a [generalist robot policy](https://xumanoids.com/glossary/generalist-robot-policy), but distillation alone does not establish that the combined model resolves conflicts between tasks or handles unfamiliar situations. ### Evaluate the student independently The student approximates its teachers through its own architecture and training data. Their abilities do not transfer automatically or exactly. As with [behavior cloning](https://xumanoids.com/glossary/behavior-cloning), matching actions on sampled inputs is only part of the evaluation. A robot policy must also be assessed while acting, when its own decisions influence later observations. The original paper's compression results should not be assumed to hold for a different robot or student model. ### Sources - [Policy Distillation](https://arxiv.org/abs/1511.06295) --- ## Pose estimation https://xumanoids.com/glossary/pose-estimation Pose estimation determines the position and orientation of an object or robot relative to a reference frame. For a rigid body in three-dimensional space, a full pose has three translational and three rotational degrees of freedom. Updated: 2026-10-05 ### A position alone is not a pose A robot reaching for a tool needs to know both where the tool is and how it is oriented. A rigid object's pose describes that combination. The reference frame matters: a pose relative to a camera must be transformed before a controller can use it in the robot's base frame. [NVIDIA's FoundationPose project](https://nvlabs.github.io/FoundationPose/) provides a concrete example of six-dimensional object pose estimation and tracking from visual inputs. The six degrees describe spatial freedom; an implementation may store them using a matrix or a quaternion plus translation. ### Estimation and tracking are different operations Estimating pose can mean locating an object in a new observation. Tracking updates an existing estimate across observations. FoundationPose supports both, with documented setups based on an object CAD model or reference images. ### Check the required object information The project does not claim to infer every object's pose without prior information: its reported novel-object setup requires a CAD model or a small set of reference images. Compare methods under their actual input assumptions. In human-motion research, pose estimation can instead refer to body keypoints or joint arrangements, so specify whether the output is a rigid-object transform or an articulated skeleton. ### Sources - [NVIDIA Research: FoundationPose](https://nvlabs.github.io/FoundationPose/) --- ## Prismatic joint https://xumanoids.com/glossary/prismatic-joint A prismatic joint permits one link to translate relative to another along a fixed joint axis without relative rotation. Its single degree of freedom is described by a linear displacement. Updated: 2026-10-05 Also known as: Linear joint, Prismatic joints ### Sliding along one axis A telescoping link or linear slide illustrates the motion of an ideal prismatic joint. The connected links keep the same relative orientation while their separation changes along the joint axis. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) classifies this as a one-degree-of-freedom joint, alongside the rotational [revolute joint](https://xumanoids.com/glossary/revolute-joint). The permitted axis can move in world coordinates when earlier links in a chain move. ### Displacement replaces joint angle A prismatic joint's coordinate has length units rather than angular units. In a [Denavit-Hartenberg model](https://xumanoids.com/glossary/denavit-hartenberg-parameters), the link offset d varies while the corresponding joint angle remains a geometric parameter. The [Robotics Toolbox documentation](https://petercorke.github.io/robotics-toolbox-python/arm_dh.html) provides separate prismatic definitions for standard and modified DH conventions. A model also needs a zero position, positive direction, and travel limits to interpret that coordinate. ### A mechanism can mix joint types Revolute and prismatic joints can occur in the same [kinematic chain](https://xumanoids.com/glossary/kinematic-chain). The Stewart-platform example in Modern Robotics combines prismatic legs with other joints to move a platform. The joint label describes its motion constraint. Actuator, transmission, friction, and inertia properties are additional model choices, as the Toolbox definitions make explicit. ### Sources - [Modern Robotics: Degrees of Freedom of a Robot](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) - [Robotics Toolbox for Python: Denavit-Hartenberg models](https://petercorke.github.io/robotics-toolbox-python/arm_dh.html) --- ## Probabilistic roadmap https://xumanoids.com/glossary/probabilistic-roadmap A probabilistic roadmap is a motion-planning graph built by sampling collision-free configurations and connecting nearby samples with feasible local paths. The graph can then answer start-to-goal queries within the modeled environment. Updated: 2026-10-05 Also known as: PRM, Probabilistic roadmaps ### Build a reusable map of possible motion A roadmap represents robot configurations as vertices and checked motions as edges. [LaValle's Planning Algorithms](https://msl.cs.uiuc.edu/planning/node239.html) explains the multiple-query setting: invest computation in a graph that can serve many future queries while the robot and obstacle models remain fixed. For example, a robot repeatedly moving its arm among work areas may reuse a roadmap through the same surrounding geometry. ### Connect the query to the graph The [basic method](https://msl.cs.uiuc.edu/planning/node240.html) separates preprocessing from query answering. At query time, the start and goal are connected to the roadmap with a local planner. Graph search then identifies a sequence of edges between them. This differs from growing a fresh [rapidly-exploring random tree](https://xumanoids.com/glossary/rapidly-exploring-random-tree) primarily around one query, although both belong to sampling-based planning. ### Reuse depends on the model staying valid A stored edge only records a motion checked against a particular model. If objects move, earlier collision checks may no longer apply. Roadmap coverage and the local connection method also affect whether a route can be found. A missing connection in a finite graph is not by itself proof that no physical route exists. ### Sources - [Steven M. LaValle: Planning Algorithms, Roadmap Methods for Multiple Queries](https://msl.cs.uiuc.edu/planning/node239.html) - [Steven M. LaValle: Planning Algorithms, The Basic Roadmap Method](https://msl.cs.uiuc.edu/planning/node240.html) --- ## Proprioception https://xumanoids.com/glossary/proprioception Proprioception in robotics is sensing the robot's own motion, configuration, and internal physical state. Typical proprioceptive inputs include joint encoders, inertial measurements, and signals associated with actuator effort or contact. Updated: 2026-10-06 Also known as: Robot proprioception, Proprioceptive sensing ### Signals from the robot's own body Proprioceptive sensors report quantities generated within the robot rather than a direct observation of distant surroundings. A joint encoder measures a [joint](https://xumanoids.com/glossary/joint) position or motion, while an [inertial measurement unit](https://xumanoids.com/glossary/inertial-measurement-unit) measures angular velocity and specific force. Motor-current, torque, and contact signals may also contribute information about the body's state. The sensor readings are not the same as a complete state estimate. [Agrawal and colleagues](https://arxiv.org/abs/2209.05644) combine preintegrated inertial measurements, [forward kinematics](https://xumanoids.com/glossary/forward-kinematics), and contact detections in a factor graph to estimate a legged robot's base and joint states. The robot model and estimator turn incomplete, noisy measurements into quantities a controller can use. ### Proprioception and exteroception answer different questions Exteroceptive sensors such as cameras and lidar observe the environment. Proprioception can reveal that a foot has loaded or slipped, but it does not provide a distant terrain map before contact. Conversely, a depth image can show an obstacle without directly measuring the torque at a knee. In reported quadruped experiments, [Miki and colleagues](https://arxiv.org/abs/2201.08117) combine proprioceptive and exteroceptive inputs for terrain-aware locomotion. Their result is evidence for that trained controller and set of tests. It does not make either sensing mode universally reliable. ### Internal sensing still has blind spots Inertial integration drifts, encoders can contain offsets, and leg odometry can be wrong when a presumed stationary foot slips. Contact state itself may need to be inferred. Proprioception also cannot identify many external hazards until they affect the body. [Sensor fusion](https://xumanoids.com/glossary/sensor-fusion) can combine internal and external observations, but its output depends on calibration, timing, noise models, and correct contact assumptions. A robot that continues moving plausibly is not proof that its estimated position or terrain model is correct. ### Sources - [Agrawal et al.: Proprioceptive State Estimation of Legged Robots with Kinematic Chain Modeling](https://arxiv.org/abs/2209.05644) - [Miki et al.: Learning robust perceptive locomotion for quadrupedal robots in the wild](https://arxiv.org/abs/2201.08117) --- ## Quasi-direct drive https://xumanoids.com/glossary/quasi-direct-drive Quasi-direct drive is an actuation approach that combines a torque-capable motor with a relatively low transmission reduction to preserve useful backdrivability and force-control behavior. It differs from direct drive because it still uses a transmission. Updated: 2026-10-05 Also known as: QDD, Quasi direct drive, Quasi-direct-drive actuation ### Keeping transmission effects manageable A transmission increases output torque while reducing speed, but it also changes friction and the motor inertia felt at the joint. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-9-actuation-gearing-and-friction/) shows that ideal reflected rotor inertia scales with the square of the reduction ratio. Quasi-direct-drive designs use modest reduction together with a motor suited to producing substantial torque. The aim is to keep [backdrivability](https://xumanoids.com/glossary/backdrivability) and controllable interaction forces while obtaining more joint torque than the motor supplies directly. ### Examples in manipulation and locomotion [Gealy and colleagues](https://arxiv.org/abs/1904.03815) describe this approach in the Blue manipulation arm, including belt transmissions and motor choices. [Katz's MIT thesis](https://dspace.mit.edu/handle/1721.1/118671) develops related low-ratio, backdrivable actuators for dynamic robots. These examples establish particular design choices, not a universal performance level for anything described as QDD. ### What the label does not specify There is no single reduction ratio that alone establishes good force control. Motor inertia, transmission friction, thermal limits, and controller bandwidth all matter. A low-ratio actuator may need a larger motor or more current to deliver the required torque. Compare continuous and peak torque under stated conditions, rather than treating quasi-direct drive as a guarantee of high payload or safe contact. ### Sources - [Gealy et al.: Quasi-Direct Drive for Low-Cost Compliant Robotic Manipulation](https://arxiv.org/abs/1904.03815) - [Katz: A Low Cost Modular Actuator for Dynamic Robots, MIT thesis](https://dspace.mit.edu/handle/1721.1/118671) - [Modern Robotics: Actuation, Gearing, and Friction](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-9-actuation-gearing-and-friction/) --- ## Quaternion https://xumanoids.com/glossary/quaternion A quaternion is a four-component mathematical object consisting of a scalar and a three-component vector. Robotics commonly uses unit quaternions to represent three-dimensional rotations without the coordinate singularities of Euler angles. Updated: 2026-10-05 Also known as: Quaternions ### Four components describe a rotation A rotation quaternion has one scalar component and three vector components. Its four values satisfy a unit-length constraint, so they represent three rotational [degrees of freedom](https://xumanoids.com/glossary/degrees-of-freedom), not four independent angles. The [SciPy rotation documentation](https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Rotation.from_quat.html) defines the scalar using the cosine of half the rotation angle and the vector using the rotation axis scaled by the sine of half that angle. Opposite-sign unit quaternions represent the same rotation, a property called a double cover. ### Check the component convention Some interfaces place the scalar first, giving w, x, y, z. Others place it last, giving x, y, z, w. SciPy supports both and defaults to scalar last. A four-number array without its ordering and frame convention is therefore ambiguous. Normalization also matters: an arbitrary four-component quaternion is not automatically a unit rotation quaternion. ### Avoiding angle-coordinate singularities [MIT's manipulation notes](https://manipulation.mit.edu/pick.html) explain how [Euler angles](https://xumanoids.com/glossary/euler-angles) become singular at particular orientations. Unit quaternions avoid that representation problem and can be converted to [rotation matrices](https://xumanoids.com/glossary/rotation-matrix). They do not remove a robot's [kinematic singularities](https://xumanoids.com/glossary/kinematic-singularity), which arise from the mechanism's motion capability rather than the chosen orientation coordinates. ### Sources - [SciPy: Rotation.from_quat](https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Rotation.from_quat.html) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Rapidly-exploring random tree https://xumanoids.com/glossary/rapidly-exploring-random-tree A rapidly-exploring random tree is a sampling-based structure that grows through a configuration or state space toward sampled targets. Motion planners use it to search for feasible routes through spaces with obstacles and movement constraints. Updated: 2026-10-05 Also known as: RRT, Rapidly-exploring random trees ### Samples pull the tree into new regions An RRT starts from an initial state. It samples a target, finds a nearby tree node, and attempts a short extension toward the target. [LaValle's description](https://lavalle.pl/rrt/about.html) explains how this procedure encourages exploration of high-dimensional spaces rather than simply taking a random walk. A humanoid arm planner can use an RRT to search among joint configurations when a direct reach intersects an obstacle. ### The local planner determines feasible growth The [Modern Robotics treatment](https://modernrobotics.northwestern.edu/nu-gm-book-resource/10-5-sampling-methods-for-motion-planning-part-2-of-2/) identifies sampling, the distance metric, and the local extension method as key design choices. An extension must respect the relevant movement constraints and pass collision checks before becoming part of a valid plan. ### Finding a route is not optimizing it Basic RRT planning does not guarantee the shortest or lowest-cost path. Continuing to grow an ordinary RRT is also not equivalent to using RRT*, whose rewiring supports asymptotic optimality under its assumptions. A returned route still needs appropriate timing and execution control. The tree is one component of a [motion-planning](https://xumanoids.com/glossary/motion-planning) system. ### Sources - [Steven M. LaValle: About Rapidly-Exploring Random Trees](https://lavalle.pl/rrt/about.html) - [Modern Robotics: Sampling Methods for Motion Planning, Part 2](https://modernrobotics.northwestern.edu/nu-gm-book-resource/10-5-sampling-methods-for-motion-planning-part-2-of-2/) --- ## Redundant manipulator https://xumanoids.com/glossary/redundant-manipulator A redundant manipulator has more independent joint-motion variables than are needed for its specified end-effector task. This can allow different joint motions or postures to produce the same task result. Updated: 2026-10-05 Also known as: Kinematically redundant manipulator, Redundant robot arm ### Extra freedom is relative to a task A seven-joint arm can have redundancy for a six-dimensional hand-pose task. A three-joint planar arm can be redundant for controlling only its tip's two-dimensional position, even if all three joints are needed when tip orientation is also specified. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) uses these examples to explain Jacobians with more columns than task dimensions. Redundancy is therefore a relationship between a mechanism and a task, not a fixed synonym for a particular joint count. ### Moving the elbow while holding the hand Some joint velocities can change internal posture while producing zero instantaneous hand velocity. These velocities lie in the task Jacobian's [null space](https://xumanoids.com/glossary/null-space). They can provide freedom to choose a preferred posture alongside the main task. [MIT's manipulation notes](https://manipulation.mit.edu/pick.html) use a secondary objective that draws the joints toward a nominal configuration while tracking the desired hand motion. ### Extra joints do not remove constraints Redundancy does not eliminate [kinematic singularities](https://xumanoids.com/glossary/kinematic-singularity), travel limits, or collisions. Secondary objectives can conflict with constraints or with the primary task. The useful freedom depends on the current configuration and the tasks already imposed, so an additional contact or orientation requirement can consume redundancy that was previously available. ### Sources - [Modern Robotics: Singularities](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Reinforcement learning https://xumanoids.com/glossary/reinforcement-learning Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations. Updated: 2026-10-05 Also known as: RL ### Optimize outcomes over time [OpenAI's reinforcement-learning documentation](https://spinningup.openai.com/en/latest/spinningup/rl_intro.html) describes an agent that takes actions, receives observations and rewards, and seeks a high return, or accumulated reward. A policy is the rule that selects actions; a value function estimates expected future return. A robot's observation can include joint angles, velocities, or camera images. Its action space might contain continuous commands rather than a small list of discrete choices. ### Reward and demonstration are different signals A reward indicates how an outcome contributes to the task objective. It does not necessarily specify the exact action the robot should take. [Imitation learning](https://xumanoids.com/glossary/imitation-learning) instead uses demonstrated behavior, although practical systems can combine both signals. For example, an object-pushing task can reward getting an object to a target. The learner must discover an action sequence that achieves that outcome. ### Training conditions shape the result Physical interaction is not the only source of experience. [Simulation-based robot-control research](https://arxiv.org/abs/1710.06537) trains policies in simulated environments before evaluating transfer to a real arm. Reward design, observations, available actions, and training dynamics all affect the learned behavior. High return in a simulator does not by itself establish reliable hardware performance, and a high reward does not establish that an incomplete task specification captured every desired constraint. ### Sources - [OpenAI Spinning Up: Key Concepts in Reinforcement Learning](https://spinningup.openai.com/en/latest/spinningup/rl_intro.html) - [Sim-to-Real Transfer of Robotic Control with Dynamics Randomization](https://arxiv.org/abs/1710.06537) --- ## Residual reinforcement learning https://xumanoids.com/glossary/residual-reinforcement-learning Residual reinforcement learning learns a corrective control signal that is combined with a baseline controller. The baseline handles part of the task while the learned residual adjusts behavior that is difficult to model or tune directly. Updated: 2026-10-05 Also known as: Residual RL ### Learn a correction around an existing controller [Residual Reinforcement Learning for Robot Control](https://arxiv.org/abs/1812.03201) decomposes control into a conventional feedback-control component and a learned residual. In the paper's formulation, the final command is the superposition of the two signals. A baseline controller could guide an arm toward a target while the residual adjusts the command during contact. The learned component does not have to rediscover every aspect of the baseline behavior. ### Contact is a motivating application The authors identify contacts and friction as effects that can be difficult to capture with simple physical models. They demonstrate the approach on a real block-assembly task involving contacts and unstable objects. This is a specific use of [reinforcement learning](https://xumanoids.com/glossary/reinforcement-learning) with an existing controller. The word residual refers to the correction added to the control signal, not to a particular residual neural-network architecture. ### A baseline does not constrain every correction Adding a learned signal changes the command the robot executes. The baseline's behavior alone therefore does not establish the behavior or stability of the combined system. The baseline, residual action space, allowed correction magnitude, and task evaluation all matter when assessing an implementation. Results on block assembly support that demonstrated application; they do not imply that adding an arbitrary learned correction will improve every controller or preserve its guarantees. ### Sources - [Residual Reinforcement Learning for Robot Control](https://arxiv.org/abs/1812.03201) --- ## Revolute joint https://xumanoids.com/glossary/revolute-joint A revolute joint permits one link to rotate relative to another about a fixed joint axis. Its single relative degree of freedom is described by an angle. Updated: 2026-10-05 Also known as: Rotary joint, Revolute joints ### One allowed relative rotation An ideal revolute joint constrains the other five relative freedoms between two spatial rigid links. Only rotation about the joint axis remains. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) uses this joint as the basic building block of a serial robot arm. A hinge provides a useful physical example. Its axis can move through world space as earlier joints move; “fixed axis” means fixed in the connected-link joint geometry, not necessarily fixed relative to the room. ### Joint angle and travel limits The joint coordinate measures rotation from a chosen zero position. The axis direction determines its positive sign. In a [Denavit-Hartenberg model](https://xumanoids.com/glossary/denavit-hartenberg-parameters), the variable is the joint angle theta, as shown in the [Robotics Toolbox's revolute-link definitions](https://petercorke.github.io/robotics-toolbox-python/arm_dh.html). A finite angular travel range limits which configurations are available. It does not turn a one-degree-of-freedom joint into a different joint type. ### Joint type is not actuator type “Revolute” describes permitted relative motion. The Toolbox records motor, transmission, friction, and inertia properties separately from this kinematic choice. A [prismatic joint](https://xumanoids.com/glossary/prismatic-joint) instead permits one translation, while a spherical joint permits three relative rotational freedoms. ### Sources - [Modern Robotics: Degrees of Freedom of a Robot](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-2-degrees-of-freedom-of-a-robot/) - [Robotics Toolbox for Python: Denavit-Hartenberg models](https://petercorke.github.io/robotics-toolbox-python/arm_dh.html) --- ## Reward shaping https://xumanoids.com/glossary/reward-shaping Reward shaping adds supplementary rewards to guide reinforcement learning toward useful behavior. Poorly chosen shaping can change which policy is optimal, so an easier training signal is not automatically equivalent to the original task objective. Updated: 2026-10-05 ### Provide feedback before task completion A robot may receive its main reward only after reaching a goal. Shaping can provide intermediate feedback, such as a signal related to progress toward that goal. [Ng, Harada, and Russell](https://people.eecs.berkeley.edu/~russell/papers/icml99-shaping.pdf) analyze when adding such feedback preserves the original optimal policy. The distinction matters because a learner optimizes the reward it actually receives, including any extra terms. ### Potential-based shaping has a precise form For a discounted problem, potential-based shaping adds `gamma * Phi(next_state) - Phi(state)`, where `Phi` assigns a potential to each state and `gamma` is the discount factor. The paper establishes policy invariance under its stated Markov decision process assumptions and boundary conditions. This form rewards a change in potential rather than handing out an unrelated bonus whenever a convenient event occurs. ### Extra rewards can create unintended loops The paper describes examples in which rewarding progress or repeated ball contact encouraged behavior that failed the intended objective. For a robot moving a part, repeatedly collecting a proximity bonus could likewise become preferable to completing placement if the reward is poorly specified. Shaping is a tool within [reinforcement learning](https://xumanoids.com/glossary/reinforcement-learning), not a substitute for defining the task. Report the full reward and verify actual task completion rather than relying only on the shaped return. ### Sources - [Ng, Harada, and Russell: Policy Invariance under Reward Transformations](https://people.eecs.berkeley.edu/~russell/papers/icml99-shaping.pdf) --- ## Rigid-body transformation https://xumanoids.com/glossary/rigid-body-transformation A rigid-body transformation changes a body's position and orientation without changing its shape or size. In three-dimensional robotics it consists of a proper rotation and a translation. Updated: 2026-10-05 Also known as: Rigid transformation, Rigid-body transform ### Moving a body without deforming it Imagine moving a rigid gripper from one pose to another. Every point moves consistently with the same rotation and translation, and distances between points remain unchanged. Scaling, shearing, and bending are outside this model. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-1-homogeneous-transformation-matrices/) describes the set of spatial rigid transforms as SE(3), the special Euclidean group. A transform can represent a body's pose, a displacement, or a relationship between coordinate frames. ### A transform and its representation A [homogeneous transformation](https://xumanoids.com/glossary/homogeneous-transformation) matrix is one way to store and compose the operation. Another stores a translation together with an orientation representation such as a unit [quaternion](https://xumanoids.com/glossary/quaternion). The underlying geometric relationship is independent of how it is stored. [MIT's spatial-algebra discussion](https://manipulation.mit.edu/pick.html) emphasizes keeping track of the frames when composing poses and converting coordinates. ### Composition requires matching frames A camera-to-torso relationship can be combined with a torso-to-world relationship to locate observations in the world. Reversing a relationship requires its inverse. Swapping multiplication order generally describes a different operation. A rigid transform describes pose, not velocity, force, or deformation. [Twists](https://xumanoids.com/glossary/twist) represent instantaneous rigid motion, while [wrenches](https://xumanoids.com/glossary/wrench) represent force and moment about a specified reference point. ### Sources - [Modern Robotics: Homogeneous Transformation Matrices](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-1-homogeneous-transformation-matrices/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Robot foundation model https://xumanoids.com/glossary/robot-foundation-model A robot foundation model is a model pretrained on broad data to support adaptation to multiple robot tasks, environments, or bodies. The term describes a reusable learning base rather than a guarantee of general physical competence. Updated: 2026-10-05 Also known as: Robotics foundation model ### Pretraining creates a reusable starting point The [foundation-model report](https://arxiv.org/abs/2108.07258) defines foundation models through broad training data and adaptation to many downstream tasks. In robotics, this can mean reusing learned visual features, action patterns, or both when training a robot for a new setting. For example, a pretrained manipulation model may provide the starting weights for learning to sort objects on a different robot arm. The new robot can still need its own demonstrations, sensor configuration, and action interface. ### Related labels describe different properties Some robotics papers use this label interchangeably with [generalist robot policy](https://xumanoids.com/glossary/generalist-robot-policy), including the [pi0 paper](https://arxiv.org/abs/2410.24164). Foundation emphasizes reuse through pretraining; generalist emphasizes the range of tasks a policy performs. Neither label specifies one architecture. A [vision-language-action model](https://xumanoids.com/glossary/vision-language-action-model) instead describes the relationship between visual input, language, and action output. These descriptions can apply to the same model. ### Evaluate the adaptation that was demonstrated [Octo](https://octo-models.github.io/) reports both control in settings represented in its training data and fine-tuning to new setups. Those are different tests. A broad training mixture does not establish that a model can operate an unfamiliar humanoid without adaptation or task-specific evaluation. ### Sources - [On the Opportunities and Risks of Foundation Models](https://arxiv.org/abs/2108.07258) - [Physical Intelligence: pi0, a vision-language-action flow model](https://arxiv.org/abs/2410.24164) - [Octo: An Open-Source Generalist Robot Policy](https://octo-models.github.io/) --- ## Robot Jacobian https://xumanoids.com/glossary/robot-jacobian A robot Jacobian is a configuration-dependent matrix that maps joint velocities to a chosen task velocity, often an end-effector twist. It describes the local relationship between joint motion and task motion. Updated: 2026-10-05 Also known as: Manipulator Jacobian, Kinematic Jacobian ### Each column describes one joint's contribution Hold the robot at a particular configuration and move one joint at unit speed while the others remain still. The resulting end-effector velocity forms that joint's Jacobian column. Combining the columns with the actual joint speeds gives the resulting task velocity. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-1-1-space-jacobian/) develops a spatial Jacobian that maps joint speeds to a [twist](https://xumanoids.com/glossary/twist) expressed in the space frame. A body Jacobian expresses that motion in the body frame instead. ### Choose the task and representation A position-only task has a different Jacobian from a full position-and-orientation task. A Jacobian for orientation-coordinate rates also differs from one for angular velocity. Frame conventions and component ordering must match the velocity being commanded. In [static force analysis](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/), the transpose of a compatible Jacobian maps an end-effector [wrench](https://xumanoids.com/glossary/wrench) to joint forces and torques. ### Local motion has limits The Jacobian is evaluated at the current configuration, so it changes as the robot moves. At a [kinematic singularity](https://xumanoids.com/glossary/kinematic-singularity), it loses rank relative to its maximum attainable rank. [MIT's differential IK discussion](https://manipulation.mit.edu/pick.html) explains why a pseudoinverse near singularity can request very large joint velocities. Joint limits and other constraints require explicit handling; taking a matrix inverse is not a complete controller. ### Sources - [Modern Robotics: Space Jacobian](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-1-1-space-jacobian/) - [Modern Robotics: Singularities](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Robotic manipulation https://xumanoids.com/glossary/robotic-manipulation Robotic manipulation is the use of a robot to change an object's position, orientation, or state through physical interaction. It includes grasping and moving objects as well as actions such as pushing or carrying them without a grasp. Updated: 2026-10-05 Also known as: Robot manipulation ### More than grasping Picking up a cup is manipulation, but so is sliding it across a table. [Modern Robotics](https://modernrobotics.northwestern.edu/chapters/chapter12/) distinguishes grasping from non-grasping manipulation and gives examples including pushing, carrying a tray, and moving objects through vibration. The distinction matters because the robot makes different contacts with the object. The same reference analyses these tasks through contact motion, contact forces, and rigid-body dynamics. ### Learning from demonstrations Manipulation can use two arms working together. The [ALOHA 2 project](https://aloha-2.github.io/) provides hardware for collecting demonstrations of two-arm tasks through [teleoperation](https://xumanoids.com/glossary/teleoperation). Those demonstrations support research into learning robot behaviour from examples. ### Handling objects is a specific capability A [humanoid robot](https://xumanoids.com/glossary/humanoid-robot) may be built for manipulation, but body shape does not establish what it can handle. Check the object, task, contact conditions, and amount of supervision in a reported result. A successful grasp of one rigid object does not establish reliable handling of every object. ### Sources - [Modern Robotics: Chapter 12, Grasping and Manipulation](https://modernrobotics.northwestern.edu/chapters/chapter12/) - [ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation](https://aloha-2.github.io/) --- ## Rotation matrix https://xumanoids.com/glossary/rotation-matrix A rotation matrix represents an orientation or rotation while preserving lengths and angles. In three dimensions it is a 3-by-3 orthonormal matrix with determinant positive one. Updated: 2026-10-05 Also known as: Rotation matrices ### The columns describe a frame One interpretation places the three unit axes of a body frame into the columns of a matrix, expressed in a reference frame. Multiplying this matrix by a vector expressed in body coordinates expresses that vector in the reference coordinates. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-2-1-rotation-matrices-part-1-of-2/) derives the constraints: each column has unit length, distinct columns are perpendicular, and the determinant is positive one. Together these define the group called SO(3). ### Composition and inversion Multiplying two rotation matrices composes their rotations. The inverse is the transpose, which makes reversing a frame relationship straightforward. The order of composition matters because three-dimensional rotations generally do not commute. A [homogeneous transformation](https://xumanoids.com/glossary/homogeneous-transformation) adds translation to this orientation representation, as the [transformation-matrix lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-1-homogeneous-transformation-matrices/) explains. Rotation alone cannot describe where a robot's hand is located. ### Nine values represent three freedoms Although the matrix stores nine numbers, its constraints leave three [degrees of freedom](https://xumanoids.com/glossary/degrees-of-freedom). [MIT's discussion of orientation representations](https://manipulation.mit.edu/pick.html) explains why matrices avoid the coordinate singularities of a minimal angle representation. An arbitrary nine-number array is not a valid rotation. [Quaternions](https://xumanoids.com/glossary/quaternion) and [Euler angles](https://xumanoids.com/glossary/euler-angles) encode the same orientation with different storage and mathematical properties. Choose the representation and frame convention explicitly when passing orientation data between systems. ### Sources - [Modern Robotics: Rotation Matrices](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-2-1-rotation-matrices-part-1-of-2/) - [Modern Robotics: Homogeneous Transformation Matrices](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-1-homogeneous-transformation-matrices/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Screw theory https://xumanoids.com/glossary/screw-theory Screw theory is a geometric framework for describing rigid-body motion and forces using axes, rotation, translation, and pitch. In robotics it provides the basis for twist and wrench representations and screw-axis formulations of kinematics. Updated: 2026-10-05 ### Rotation and translation share an axis A screw motion combines rotation about an axis with translation along it. Pitch relates the translation to the rotation. Zero pitch gives pure rotation; pure translation is treated as a limiting case with no angular component. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-2-twists-part-1-of-2/) shows how an instantaneous rigid-body velocity can be represented by a screw axis and a scalar speed. The resulting six-component velocity is a [twist](https://xumanoids.com/glossary/twist). ### From joint axes to robot motion A [revolute joint](https://xumanoids.com/glossary/revolute-joint) contributes rotation about its joint axis. A [prismatic joint](https://xumanoids.com/glossary/prismatic-joint) contributes translation. Their screw-axis descriptions allow a common mathematical treatment even when a chain mixes joint types. The [product-of-exponentials formulation](https://modernrobotics.northwestern.edu/nu-gm-book-resource/4-1-1-product-of-exponentials-formula-in-the-space-frame/) composes these joint motions to calculate [forward kinematics](https://xumanoids.com/glossary/forward-kinematics). The corresponding finite motion comes from integrating a constant twist, as explained in the [exponential-coordinates lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-3-exponential-coordinates-of-rigid-body-motion/). ### Geometry is not a drive mechanism The word “screw” describes a geometric motion, not a requirement for a threaded mechanical screw. The framework also packages forces and moments as [wrenches](https://xumanoids.com/glossary/wrench), described in the [force-representation lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-4-wrenches/). It describes ideal rigid motion and force relationships; deformation and actuator dynamics require additional models. ### Sources - [Modern Robotics: Twists, Part 1](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-2-twists-part-1-of-2/) - [Modern Robotics: Exponential Coordinates of Rigid-Body Motion](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-3-exponential-coordinates-of-rigid-body-motion/) - [Modern Robotics: Product of Exponentials Formula in the Space Frame](https://modernrobotics.northwestern.edu/nu-gm-book-resource/4-1-1-product-of-exponentials-formula-in-the-space-frame/) - [Modern Robotics: Wrenches](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-4-wrenches/) --- ## Self-supervised learning https://xumanoids.com/glossary/self-supervised-learning Self-supervised learning builds a training signal from the structure of the data itself rather than requiring a human label for every example. In robotics it can learn useful visual or temporal representations before a downstream task policy is trained. Updated: 2026-10-05 Also known as: SSL ### Construct supervision from related observations [Bootstrap Your Own Latent](https://arxiv.org/abs/2006.07733), or BYOL, trains an image representation by predicting a target network's representation of another augmented view of the same image. The two related views provide the learning relationship without a separate object-class label for that image. This is one self-supervised objective, not the definition of every self-supervised method. Other objectives can use relationships over time or between parts of an observation. ### Reuse representations for robot learning [R3M](https://arxiv.org/abs/2203.12601) pretrains visual representations on human video using a combination of time-contrastive learning, video-language alignment, and a sparsity objective. It then uses the representation as a frozen perception module for downstream robot-policy learning. R3M's mixture includes language information, so it should not be described as learning solely from unlabeled image similarity. The example shows how several supervision sources can contribute to a reusable representation. ### A representation is not a complete behavior Learning that observations are related does not itself specify the robot's task or motor commands. A downstream [imitation-learning](https://xumanoids.com/glossary/imitation-learning) or other policy-learning stage may still require task data. Evaluate the downstream task separately. Strong representation-learning results on images or human video do not automatically demonstrate physical competence on a new robot body. ### Sources - [Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning](https://arxiv.org/abs/2006.07733) - [R3M: A Universal Visual Representation for Robot Manipulation](https://arxiv.org/abs/2203.12601) --- ## Sensor fusion https://xumanoids.com/glossary/sensor-fusion Sensor fusion combines information from multiple sensors or estimation sources to produce a shared estimate. The combination must account for coordinate frames, timing, uncertainty, and dependence between inputs. Updated: 2026-10-05 Also known as: Multisensor fusion, Multi-sensor fusion ### Complementary measurements support one estimate A camera can observe scene structure while an [IMU](https://xumanoids.com/glossary/inertial-measurement-unit) measures rapid rotational and inertial motion. Combining them can support [visual-inertial odometry](https://xumanoids.com/glossary/visual-inertial-odometry). Sensor fusion is the broader process; it is not the name of one particular algorithm. The [robot_localization package](https://docs.ros.org/en/noetic/api/robot_localization/html/index.html) illustrates fusion of pose, velocity, odometry, and IMU messages with configurable choices about which variables to include. ### More inputs are not automatically more information Two reported values may come from the same underlying measurement. For example, wheel-derived position and velocity can share encoder errors. The [configuration guide](https://docs.ros.org/en/noetic/api/robot_localization/html/configuring_robot_localization.html) warns against feeding duplicate information into a filter as though it were independent evidence. ### Frames and uncertainty must agree The [estimator documentation](https://docs.ros.org/en/noetic/api/robot_localization/html/state_estimation_nodes.html) describes frame transformations, sensor timeouts, and covariance settings. These are part of the measurement interpretation. Combining a camera-frame velocity with a body-frame velocity without the correct transformation does not produce a meaningful robot estimate. Likewise, understated measurement uncertainty can make one unreliable source dominate the output. ### Sources - [ROS robot_localization: State Estimation Features](https://docs.ros.org/en/noetic/api/robot_localization/html/index.html) - [ROS robot_localization: Configuring Sensor Fusion](https://docs.ros.org/en/noetic/api/robot_localization/html/configuring_robot_localization.html) - [ROS robot_localization: State Estimation Nodes](https://docs.ros.org/en/noetic/api/robot_localization/html/state_estimation_nodes.html) --- ## Series elastic actuator https://xumanoids.com/glossary/series-elastic-actuator A series elastic actuator places an elastic element in the force-transmission path between the drive and its load. Measuring the element’s deflection can support force or torque feedback while the elasticity changes the actuator’s response to impacts. Updated: 2026-10-05 Also known as: SEA, Series-elastic actuator ### A spring in the load path The defining feature is intentional compliance in series with the output, rather than only a flexible cover or a spring acting alongside the drive. For a calibrated linear spring, force is related to extension by `F = k * delta`; a torsional spring has an analogous torque-angle relationship. [Williamson's MIT thesis](https://dspace.mit.edu/handle/1721.1/6776) studies this arrangement for force-controlled actuation. [Katz's actuator thesis](https://dspace.mit.edu/handle/1721.1/118671) explains how measuring spring deflection turns output [torque control](https://xumanoids.com/glossary/torque-control) into a problem of regulating that deflection. ### Why use it in a robot Elasticity can separate the motor and gearbox from abrupt output disturbances and provide a useful force measurement. In some designs, the spring also stores and returns energy. A leg actuator can therefore respond differently to touchdown than an otherwise rigid transmission would. ### Compliance brings tradeoffs Spring stiffness, travel, sensing, and controller design affect the achievable force response. [Williamson](https://dspace.mit.edu/handle/1721.1/6776) explicitly describes a tradeoff between bandwidth and the benefits of compliant force control. Adding elasticity also introduces motion between motor and output, which must be represented in control and estimation. An SEA is not automatically accurate, safe, or energy-efficient under every load; those properties require evidence for the particular design and operating conditions. ### Sources - [Williamson: Series Elastic Actuators, MIT thesis](https://dspace.mit.edu/handle/1721.1/6776) - [Katz: A Low Cost Modular Actuator for Dynamic Robots, MIT thesis](https://dspace.mit.edu/handle/1721.1/118671) --- ## Signed distance field https://xumanoids.com/glossary/signed-distance-field A signed distance field represents a surface by assigning spatial locations a distance value whose sign distinguishes the two sides of the surface. Its zero level marks the surface, while the sign convention depends on the representation. Updated: 2026-10-05 Also known as: SDF ### Distance information helps plan clearance A distance field lets a planner ask how far a candidate robot position lies from an obstacle. The [Voxblox paper](https://arxiv.org/abs/1611.03631) develops an incremental Euclidean signed distance field for onboard trajectory planning from sensor observations. This supports [trajectory optimization](https://xumanoids.com/glossary/trajectory-optimization), where distance values and their spatial variation can guide a path away from nearby obstacles. The field remains a representation of the observed environment, not a guarantee about every unseen surface. ### Euclidean and truncated fields differ A Euclidean signed distance field represents distance to the nearest surface. A truncated signed distance field limits stored distances to a band and is often built from projective sensor measurements. Voxblox explains why a TSDF used for surface reconstruction is not directly interchangeable with an ESDF used for clearance queries. ### Unknown space needs an explicit policy The paper discusses how unobserved voxels are treated during mapping and planning. Absence of a surface measurement does not establish empty space. Before using a field for robot movement, check its sign convention, voxel resolution, truncation, and treatment of unknown regions. Here SDF means signed distance field, not the separate Simulation Description Format used by some simulators. ### Sources - [Oleynikova et al.: Voxblox, Incremental 3D Euclidean Signed Distance Fields](https://arxiv.org/abs/1611.03631) --- ## Sim-to-real transfer https://xumanoids.com/glossary/sim-to-real-transfer Sim-to-real transfer applies a model, policy, or behavior developed in simulation to a physical system. Its central challenge is the difference between the simulated environment and the robot, sensors, and interactions encountered in reality. Updated: 2026-10-05 Also known as: Sim2real, Simulation-to-reality transfer ### Simulation and hardware differ A simulated robot can have different friction, actuation, sensing, or contact behavior from its physical counterpart. [Peng and colleagues](https://arxiv.org/abs/1710.06537) describe how policies can exploit simulator-specific dynamics and fail when those assumptions change on hardware. A pushing policy, for example, may learn exactly how far a simulated object slides. The same command can produce a different result with a heavier object or another surface. ### Transfer can target perception or control [Tobin and colleagues](https://arxiv.org/abs/1703.06907) train visual object localization using randomized rendered images and demonstrate its use in real robot grasping. Peng and colleagues instead randomize simulated dynamics while learning a robot control policy. These are different transfer problems: one concerns interpreting images, while the other concerns choosing actions under changing physical behavior. ### Training strategies need hardware evaluation [Domain randomization](https://xumanoids.com/glossary/domain-randomization) varies simulated conditions so the learned system experiences a broader range. [System identification](https://xumanoids.com/glossary/system-identification) can instead help estimate a model from physical measurements; adaptation methods can use target-domain data. A simulated success rate is not a real-robot success rate. Report the physical tasks actually tested and any calibration or additional training used during transfer. Successful transfer in one pushing setup does not establish transfer for every contact-rich humanoid skill. ### Sources - [Sim-to-Real Transfer of Robotic Control with Dynamics Randomization](https://arxiv.org/abs/1710.06537) - [Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World](https://arxiv.org/abs/1703.06907) --- ## Simultaneous localization and mapping https://xumanoids.com/glossary/simultaneous-localization-and-mapping Simultaneous localization and mapping is the joint estimation of a robot's state and a map of its environment from sensor observations. It is commonly abbreviated SLAM. Updated: 2026-10-05 Also known as: SLAM ### Locating the robot while building the map A robot exploring an unfamiliar building cannot assume it already knows either the floor plan or its exact position. SLAM estimates both together. The [Cadena et al. survey](https://arxiv.org/abs/1606.05830) describes maps built from features, surfaces, and other representations, coupled to estimates of the robot's motion. A humanoid could use this information to navigate between rooms. The map and estimated position provide inputs to navigation; they do not themselves choose a collision-free sequence of footsteps. ### Loop closure and accumulated error [Odometry](https://xumanoids.com/glossary/odometry) accumulates small motion errors. Recognizing a previously visited place creates a loop-closure constraint that can correct accumulated inconsistency in a SLAM map. This distinguishes mapping with revisits from simply integrating movement over time. ### Recognition can be wrong The survey identifies incorrect data association and perceptual aliasing as important problems. Two similar-looking corridors can be mistaken for the same place. A plausible-looking map therefore needs consistency checks and evaluation against reference measurements. Localization in an already known map is also a distinct task and does not necessarily require building a new map. ### Sources - [Cadena et al.: Past, Present, and Future of Simultaneous Localization and Mapping](https://arxiv.org/abs/1606.05830) --- ## State estimation https://xumanoids.com/glossary/state-estimation State estimation infers quantities describing a robot or its environment from measurements and a model. A robot state may include position, orientation, velocity, and other variables that are not all directly measured. Updated: 2026-10-05 Also known as: Robot state estimation ### Measurements are evidence about a state A camera, joint encoder, or inertial sensor measures only part of what a controller needs. An estimator combines available evidence with a model to infer a useful state. The [robot_localization documentation](https://docs.ros.org/en/noetic/api/robot_localization/html/index.html) gives a concrete implementation that tracks position, orientation, velocities, and linear accelerations. For a walking robot, body orientation and velocity are examples of quantities that can inform balance control. The exact state vector depends on the application; it is not fixed by the phrase state estimation. ### Prediction and correction A model predicts how the state changes between observations. New observations correct the prediction. [ROS's estimator documentation](https://docs.ros.org/en/noetic/api/robot_localization/html/state_estimation_nodes.html) describes this pattern for an [extended Kalman filter](https://xumanoids.com/glossary/extended-kalman-filter), including prediction-only operation when a sensor times out. ### An estimate has uncertainty Covariance represents uncertainty within these filters. A stream of smooth numbers does not mean the underlying state is known exactly. Missing measurements, uncertain models, and an unsuitable state representation affect the result. [Pose estimation](https://xumanoids.com/glossary/pose-estimation) is a narrower example that focuses on position and orientation rather than every variable in a full robot state. ### Sources - [ROS robot_localization: State Estimation Features](https://docs.ros.org/en/noetic/api/robot_localization/html/index.html) - [ROS robot_localization: State Estimation Nodes](https://docs.ros.org/en/noetic/api/robot_localization/html/state_estimation_nodes.html) --- ## Support polygon https://xumanoids.com/glossary/support-polygon The support polygon is the convex hull of a robot’s active contact points or contact patches projected onto a common support plane. It describes the available support region in planar contact models. Updated: 2026-10-05 Also known as: Support polygons ### Which contacts count With both feet flat on a level floor, the support polygon encloses the two contact patches and the space between them. With one foot lifted, only the supporting foot contributes. A foot hovering just above the floor provides no support. For contacts that can push but not pull, [MIT's contact-force derivation](https://underactuated.mit.edu/humanoids.html) shows that the center of pressure must lie within the convex hull of the contact points. The available polygon therefore changes as the robot makes and breaks contact. ### Using it for standing and walking For static equilibrium under gravity on a horizontal plane, the vertical projection of the [center of mass](https://xumanoids.com/glossary/center-of-mass) must lie within the support region, assuming the contacts can supply the needed forces. A standing humanoid can widen its feet to increase that region. Dynamic walking instead involves acceleration and momentum. A planner often constrains the [zero-moment point](https://xumanoids.com/glossary/zero-moment-point), rather than simply keeping the center-of-mass projection inside the feet. ### A polygon cannot describe every constraint A large support polygon does not prevent sliding on a low-friction floor. Contact forces must also satisfy their [friction cones](https://xumanoids.com/glossary/friction-cone). Contacts on stairs or against a wall need a more general analysis of feasible forces and moments; a flat drawing of the feet loses relevant geometry. ### Sources - [MIT Underactuated Robotics: Highly-articulated Legged Robots](https://underactuated.mit.edu/humanoids.html) --- ## Synthetic training data https://xumanoids.com/glossary/synthetic-training-data Synthetic training data is data produced computationally for model training rather than collected directly as the corresponding real-world examples. In robotics, it often includes rendered sensor observations, simulated trajectories, and labels available from the simulator. Updated: 2026-10-05 Also known as: Synthetic data ### Generate examples and their labels [Tremblay and colleagues](https://arxiv.org/abs/1804.06516) train object detection using synthetic images with randomized lighting, poses, and textures. Because the scene is generated, its construction can supply training labels without separately hand-annotating every real camera frame. A robot-perception dataset could similarly render many object arrangements and provide their positions. [Tobin and colleagues](https://arxiv.org/abs/1703.06907) study simulated visual data for localization and demonstrate the learned detector in real grasping. ### Synthetic experience can include actions Data need not consist only of images. [Dynamics-randomization research](https://arxiv.org/abs/1710.06537) trains control policies from simulated robot interaction before testing them on a physical arm. The simulator can provide observations, actions, and outcomes, but those outcomes reflect its model of physics. ### More generated data does not remove the reality gap Synthetic examples can increase variation and reduce some collection costs. Their usefulness still depends on whether they represent features and behavior relevant to the physical task. [Domain randomization](https://xumanoids.com/glossary/domain-randomization) deliberately varies simulated properties to support transfer. It does not make simulated evidence equivalent to hardware evaluation. Reports should state which data were generated, how real data were used, and which real tasks were tested before claiming a perception or control improvement. ### Sources - [Training Deep Networks with Synthetic Data: Bridging the Reality Gap by Domain Randomization](https://arxiv.org/abs/1804.06516) - [Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World](https://arxiv.org/abs/1703.06907) - [Sim-to-Real Transfer of Robotic Control with Dynamics Randomization](https://arxiv.org/abs/1710.06537) --- ## System identification https://xumanoids.com/glossary/system-identification System identification estimates a model of a physical system from measured inputs and outputs. In robotics, it can recover parameters such as inertia and friction or learn a more general model of how actions change the system state. Updated: 2026-10-05 Also known as: System ID ### Estimate dynamics from measurements [MIT's system-identification chapter](https://underactuated.mit.edu/sysid.html) distinguishes estimating parameters in a known mechanical model from learning a model whose structure is not already known. For a robot with a known joint arrangement, the equations may be established while masses, inertias, friction, or joint offsets need estimation. For example, measured joint positions, velocities, and commanded inputs can help fit a model of an arm's response. This model can support simulation and controller design. ### Identification differs from state estimation [State estimation](https://xumanoids.com/glossary/state-estimation) asks what the robot's state is now, such as its current pose or velocity. Identification asks which model explains how that state changes. A practical identification procedure can use a state estimator as one stage. The MIT notes also explain why [inverse dynamics](https://xumanoids.com/glossary/inverse-dynamics) offers useful structure: rigid-body inverse-dynamics equations are affine in inertial parameters, whereas solving forward dynamics generally loses that property. ### A small fitting error can hide a poor model Fitting one-step predictions and fitting an entire simulated trajectory are different objectives. The notes show that a model with small one-step error can still accumulate large simulation error. Model complexity matters as well. A model must capture the behavior relevant to the intended task while remaining usable by the chosen planning and control methods. ### Sources - [MIT Underactuated Robotics: System Identification](https://underactuated.mit.edu/sysid.html) --- ## Tactile sensing https://xumanoids.com/glossary/tactile-sensing Tactile sensing measures information arising from physical contact, such as contact geometry, deformation, or force. Robots use it to observe interactions at their fingers, grippers, feet, or other contact surfaces. Updated: 2026-10-05 Also known as: Robot touch sensing ### Contact supplies information that vision may miss During a grasp, a camera outside the hand may not see the contacting surface. A tactile sensor measures the local interaction itself. The [DIGIT project](https://digit.ml/digit.html) describes a compact sensor designed for robotic in-hand manipulation that captures contact geometry as images. A tactile measurement can therefore add evidence about a grasp after the fingers touch an object. That evidence differs from recognizing an object before contact. ### Some tactile sensors use cameras Vision-based tactile sensing uses images of a deformable sensing surface to observe contact. DIGIT reports that its images can support normal-force estimation and, with markers, shear-force estimation. The camera is observing the contact sensor's deformation, rather than simply looking at the scene around the robot. ### Measurement and interpretation are separate A sensor image does not automatically provide a calibrated force or a successful grasp. The desired quantity needs an appropriate estimation method, and the sensor must be mounted where relevant contact occurs. DIGIT's platform examples concern particular hands and grippers; they do not establish that every robot has whole-body touch perception. [Force control](https://xumanoids.com/glossary/force-control) is a separate process that uses measurements or estimates to regulate interaction. ### Sources - [Meta Research: DIGIT Tactile Sensor](https://digit.ml/digit.html) --- ## Task and motion planning https://xumanoids.com/glossary/task-and-motion-planning Task and motion planning jointly searches over discrete task decisions and continuous robot motions. It connects choices such as which object to move or which grasp to use with geometrically and kinematically feasible trajectories. Updated: 2026-10-06 Also known as: TAMP, Integrated task and motion planning ### Discrete choices meet continuous geometry A task planner can reason that a robot must open a cupboard before retrieving an object. A motion planner can search for a collision-free arm path to a specified pose. Task and motion planning couples these levels because the task choice may determine whether a feasible grasp, stance, and path exist. [Garrett and colleagues](https://arxiv.org/abs/2010.01083) describe TAMP problems as containing discrete task planning, discrete-continuous mathematical programming, and continuous motion planning. A solution usually includes both a sequence of symbolic actions and continuous values such as robot configurations, object poses, grasps, or trajectories. ### Why planning the levels separately can fail Suppose a humanoid chooses to pick up a box with its right hand. That symbolic choice may be impossible from the available stance, while a left-hand grasp or a base repositioning step would work. If the task plan is fixed before geometric checks, the system may spend time refining an infeasible sequence. TAMP methods use feedback between the levels so continuous failures can change discrete choices. This makes TAMP broader than [motion planning](https://xumanoids.com/glossary/motion-planning). Motion planning normally assumes the action, start state, and goal condition are already specified. TAMP may have to decide what those actions and goals should be. ### Search remains difficult The discrete plan space grows with possible actions and objects, while every candidate can create one or more continuous feasibility problems. Collision geometry, [configuration-space](https://xumanoids.com/glossary/configuration-space) constraints, grasp choices, and contact modes can make a seemingly short task expensive to solve. Different TAMP algorithms make different completeness, sampling, and modelling tradeoffs, as surveyed by [Garrett and colleagues](https://arxiv.org/abs/2010.01083). A returned plan is only as useful as its world model and execution assumptions. Perception errors, moved objects, or failed grasps may require replanning rather than blind continuation. ### Sources - [Garrett et al.: Integrated Task and Motion Planning](https://arxiv.org/abs/2010.01083) --- ## Teleoperation https://xumanoids.com/glossary/teleoperation Teleoperation is the control of a robot by a human operator from a separate location or interface. The operator supplies commands while feedback, such as camera images or the robot's motion, helps them guide the task. Updated: 2026-10-05 Also known as: Robot teleoperation ### A person supplies the actions Teleoperation keeps a person in the control loop. In the [ALOHA 2 system](https://aloha-2.github.io/), an operator uses a pair of leader arms to guide a pair of follower robot arms through manipulation tasks. The project also provides a simulation model for collecting demonstrations. The interface and distance can vary. The defining feature is that a person directs the robot's actions through an interface rather than the robot choosing the entire task on its own. ### A way to collect training examples Teleoperation can produce demonstrations for robot learning. [ALOHA 2](https://aloha-2.github.io/) was designed to improve the hardware and operator experience used for large-scale two-arm data collection. A learned policy can later attempt similar tasks using that data. ### Teleoperated and autonomous demonstrations differ When assessing a robot demonstration, check who supplied the actions. A person successfully controlling a [manipulation task](https://xumanoids.com/glossary/robotic-manipulation) demonstrates something different from a robot completing that task with a learned policy. Training from teleoperated examples does not mean a deployed system is always teleoperated. ### Sources - [ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation](https://aloha-2.github.io/) --- ## Test-time adaptation https://xumanoids.com/glossary/test-time-adaptation Test-time adaptation adjusts a trained model using data encountered during evaluation or deployment. Unlike ordinary fixed-model inference, it updates model parameters or statistics in response to the target data. Updated: 2026-10-05 Also known as: Test time adaptation, TTA ### Adapt after the original training stage [Tent](https://arxiv.org/abs/2006.10726) studies fully test-time adaptation with access only to test data and the trained model's parameters. It estimates normalization statistics and updates channel-wise affine parameters by minimizing prediction entropy. Entropy here measures uncertainty in the model's output distribution. The method uses that quantity as an optimization signal without requiring target class labels during adaptation. ### The setting differs from ordinary fine-tuning Conventional [domain adaptation](https://xumanoids.com/glossary/domain-adaptation) may allow source data, labeled target examples, or a separate adaptation training stage. A fully test-time method has a more restricted information setting. It also differs from spending more computation on a fixed model's answer. The defining feature is adaptation of the model or its statistics, not simply additional sampling or a longer reasoning process. ### Perception benchmarks do not establish robot safety Tent reports image-classification and semantic-segmentation evaluations. A robot camera encountering changed lighting is a possible application context, but the paper does not establish safe autonomous robot control under arbitrary changes. Lower prediction entropy means greater confidence, not a direct physical correctness measurement. An implementation should identify what is updated and evaluate behavior under its intended sequence of conditions, because the deployed model changes as it encounters data. ### Sources - [Tent: Fully Test-Time Adaptation by Entropy Minimization](https://arxiv.org/abs/2006.10726) --- ## Torque control https://xumanoids.com/glossary/torque-control Torque control regulates the turning effort delivered by an actuator or robot joint. It provides an actuation interface from which motion, force, and impedance controllers can produce the joint torques their tasks require. Updated: 2026-10-05 Also known as: Joint torque control ### Commanding effort at the joint A position target specifies where a joint should go; a torque target specifies its turning effort. The resulting motion also depends on inertia, gravity, friction, and external contact. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-3-of-3/) derives motion controllers that account for these dynamics before issuing torque commands. For example, keeping an arm horizontal requires torque even when its position is constant. Accelerating the same arm adds an inertial requirement, and pressing against a wall changes the joint torques through the contact [wrench](https://xumanoids.com/glossary/wrench). ### How actuators implement it For an electric motor in its usual operating model, torque relates to motor current. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-9-actuation-gearing-and-friction/) explains how gearing and friction affect the relationship between motor torque and joint output. Some systems use output torque sensors, while a [series elastic actuator](https://xumanoids.com/glossary/series-elastic-actuator) can infer torque from calibrated spring deflection. ### Command accuracy has limits A commanded motor current is not proof that the joint delivered the requested output torque. Transmission losses, model errors, current limits, and heating can change the result. [Katz's thesis](https://dspace.mit.edu/handle/1721.1/118671) separately characterizes torque accuracy and thermal behavior, illustrating why both matter. The available torque also varies with speed and operating duration, so peak torque alone is insufficient for planning sustained tasks. ### Sources - [Modern Robotics: Actuation, Gearing, and Friction](https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-9-actuation-gearing-and-friction/) - [Modern Robotics: Motion Control with Torque or Force Inputs, Part 3](https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-3-of-3/) - [Katz: A Low Cost Modular Actuator for Dynamic Robots, MIT thesis](https://dspace.mit.edu/handle/1721.1/118671) --- ## Trajectory optimization https://xumanoids.com/glossary/trajectory-optimization Trajectory optimization finds a time-varying motion, and often control inputs, that minimizes an objective while satisfying specified constraints. Robot applications can include geometric, kinematic, and dynamic constraints. Updated: 2026-10-05 Also known as: Trajectory optimisation ### Choosing motion by an objective A trajectory specifies how state changes with time. An optimizer can search for a trajectory that reaches a target while minimizing a cost such as elapsed time or control effort and respecting limits. [MIT's trajectory-optimization notes](https://underactuated.mit.edu/trajopt.html) formulate this as a finite-horizon problem from a specified initial condition. Dynamic formulations include equations of motion; kinematic formulations can focus on geometry and motion limits. ### Turning a trajectory into decision variables Direct shooting optimizes inputs and obtains states by simulating the dynamics. Direct transcription represents sampled states and inputs as decision variables, imposing the dynamics as constraints. Direct collocation represents motion between samples with functions such as polynomials and constrains their consistency with the dynamics. These choices affect numerical behavior and how a solver can use an initial guess. They do not guarantee the same solution for every problem. ### Optimal depends on the formulation Nonlinear problems can have local minima, and a solver may fail to find a feasible trajectory even when one exists. MIT discusses the importance of initialization and warns against confusing solver failure with proven infeasibility. A computed trajectory is also distinct from a feedback policy. Executing it requires control that handles tracking errors. [Model predictive control](https://xumanoids.com/glossary/model-predictive-control) repeatedly solves a finite-horizon problem as new state estimates arrive. ### Sources - [MIT Underactuated Robotics: Trajectory Optimization](https://underactuated.mit.edu/trajopt.html) --- ## Twist https://xumanoids.com/glossary/twist A twist is a six-component representation of a rigid body's instantaneous motion, combining angular and linear velocity. Its numerical values depend on the reference frame and the point used for the linear component. Updated: 2026-10-05 Also known as: Rigid-body twist ### Angular and linear motion together A rotating and translating robot hand needs more than a three-component linear velocity to describe its motion. A twist combines three angular components with three linear components. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-2-twists-part-1-of-2/) orders the angular components first and relates the result to a screw axis multiplied by a motion rate. The components do not have identical units: angular speed and linear speed describe different physical quantities. ### State the frame and reference point The same rigid motion can have a body-frame or space-frame representation. Moving the reference origin changes the linear component when angular velocity is present. The linear part of a spatial twist is therefore not automatically the ordinary velocity of the body's origin. The [twist transformation lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-2-twists-part-2-of-2/) uses a 6-by-6 adjoint matrix to change representation. Simply multiplying a six-vector by a 4-by-4 pose transform is not valid. ### Twists are useful control quantities A [robot Jacobian](https://xumanoids.com/glossary/robot-jacobian) maps joint velocities to a compatible end-effector twist. [MIT's manipulation notes](https://manipulation.mit.edu/pick.html) use spatial velocity commands for differential inverse kinematics. Angular velocity is not generally the componentwise derivative of [Euler angles](https://xumanoids.com/glossary/euler-angles). A twist also differs from a [wrench](https://xumanoids.com/glossary/wrench), which contains force and moment rather than motion. ### Sources - [Modern Robotics: Twists, Part 1](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-2-twists-part-1-of-2/) - [Modern Robotics: Twists, Part 2](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-2-twists-part-2-of-2/) - [MIT Robotic Manipulation: Basic Pick and Place](https://manipulation.mit.edu/pick.html) --- ## Underactuation https://xumanoids.com/glossary/underactuation Underactuation means that a system’s available control inputs cannot independently command acceleration in every degree of freedom of its model. It often occurs when a mechanism has fewer independent actuators than degrees of freedom. Updated: 2026-10-05 Also known as: Underactuated system, Underactuated robot ### Fewer independent inputs than motions A two-joint pendulum with a motor at only one joint is a simple example. The unpowered joint still moves, but its acceleration follows from gravity and coupling with the actuated joint. [MIT's Underactuated Robotics notes](https://underactuated.mit.edu/intro.html) formalize the distinction using the rank of the mapping from control inputs to accelerations. A humanoid also has an unactuated floating base: its motors apply internal joint torques rather than directly commanding the torso's position in space. [Ground contacts](https://underactuated.mit.edu/humanoids.html) provide the external forces through which it can move and support its body. ### Underactuated does not mean uncontrollable Coupled dynamics can allow a robot to reach useful states over time even when it cannot independently select every acceleration now. Swinging an unpowered joint by moving another joint illustrates the difference between instantaneous actuation and longer-term control. ### The chosen model matters [MIT's definition](https://underactuated.mit.edu/intro.html) emphasizes that actuation depends on the model and state. A rigid-link model may appear fully actuated while a more detailed model includes unactuated flexible modes. Contacts can also change the feasible motions. Counting motors is therefore a useful first check, but it does not replace examining input directions, constraints, and the modeled [degrees of freedom](https://xumanoids.com/glossary/degrees-of-freedom). ### Sources - [MIT Underactuated Robotics: Fully-actuated vs Underactuated Systems](https://underactuated.mit.edu/intro.html) - [MIT Underactuated Robotics: Highly-articulated Legged Robots](https://underactuated.mit.edu/humanoids.html) --- ## URDF https://xumanoids.com/glossary/unified-robot-description-format URDF is an XML format for describing a robot's links, joints, geometry, and associated physical properties. ROS tools use it to represent a robot model for visualization, kinematics, and related applications. Updated: 2026-10-05 Unified Robot Description Format ### Links and joints describe the robot The [ROS 2 URDF documentation](https://docs.ros.org/en/jazzy/Tutorials/Intermediate/URDF/URDF-Main.html) introduces the format as a description of robot geometry and organization. Links represent body components, and joints specify their relationships and permitted relative movement. A humanoid model can describe a torso connected to arm and leg links. The [movable-model tutorial](https://docs.ros.org/en/jazzy/Tutorials/Intermediate/URDF/Building-a-Movable-Robot-Model-with-URDF.html) explains joint types and limits used to make such a model articulate. ### Visual appearance and physical geometry differ A visual mesh specifies appearance. Collision geometry specifies shapes for collision checking. The [physical-properties tutorial](https://docs.ros.org/en/jazzy/Tutorials/Intermediate/URDF/Adding-Physical-and-Collision-Properties-to-a-URDF-Model.html) treats these separately and adds mass, centre of mass, and inertia for simulation. Simpler collision shapes can reduce computation, while accurate inertial values matter when dynamics enter the task. ### A description does not control the hardware Loading a valid URDF does not make a physical robot move or establish that its simulated behaviour matches reality. Controllers, live joint states, actuator interfaces, and appropriate simulator integration are separate parts of the system. A visually correct model can still have incorrect joint axes, limits, collision shapes, or inertial parameters. ### Sources - [ROS 2 Documentation: Unified Robot Description Format](https://docs.ros.org/en/jazzy/Tutorials/Intermediate/URDF/URDF-Main.html) - [ROS 2 Documentation: Building a Movable Robot Model](https://docs.ros.org/en/jazzy/Tutorials/Intermediate/URDF/Building-a-Movable-Robot-Model-with-URDF.html) - [ROS 2 Documentation: Adding Physical and Collision Properties to a URDF Model](https://docs.ros.org/en/jazzy/Tutorials/Intermediate/URDF/Adding-Physical-and-Collision-Properties-to-a-URDF-Model.html) --- ## Vision-language model https://xumanoids.com/glossary/vision-language-model A vision-language model processes visual information and natural language in a shared system. Depending on its design, it may connect images with text representations or generate text from visual and textual inputs. Updated: 2026-10-05 Also known as: VLM, Vision language model ### The output depends on the architecture [CLIP](https://arxiv.org/abs/2103.00020) learns image and text representations by identifying which captions belong with which images. It can compare a visual observation with descriptions without being a conversational text generator. [PaLI](https://arxiv.org/abs/2209.06794) instead generates text from visual and textual inputs, supporting tasks such as image captioning and visual question answering. Both are vision-language models, so the label alone does not specify the output interface. ### Visual semantics can help robot tasks A robot may need to distinguish the object described as a red cup from neighboring objects. A vision-language representation can provide semantic information for that decision. [CLIPort](https://cliport.github.io/) combines CLIP's visual-language features with a spatial manipulation architecture to learn language-conditioned picking and placing. The robot policy adds action-relevant structure beyond the pretrained visual-language representation. ### Understanding an image does not execute an action A [vision-language-action model](https://xumanoids.com/glossary/vision-language-action-model) explicitly includes robot action output. A VLM can instead serve as a perception component, task planner, or source of pretrained features in a larger robot system. Textual accuracy and physical task success therefore require different evaluations. Naming the correct object does not by itself show that the robot can reach it, grasp it, or complete the requested operation. ### Sources - [CLIP: Learning Transferable Visual Models From Natural Language Supervision](https://arxiv.org/abs/2103.00020) - [PaLI: A Jointly-Scaled Multilingual Language-Image Model](https://arxiv.org/abs/2209.06794) - [CLIPort: What and Where Pathways for Robotic Manipulation](https://cliport.github.io/) --- ## Vision-language-action model https://xumanoids.com/glossary/vision-language-action-model A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute. Updated: 2026-10-05 Also known as: VLA, VLA model, Vision language action model ### Action is part of the model output A vision-language model can describe an image or answer a question about it. A VLA also produces robot action outputs. [RT-2](https://deepmind.google/blog/rt-2-new-model-translates-vision-and-language-into-action/) represents actions as tokens, allowing a model trained with visual and language data to generate instructions for robotic control. For example, a robot may receive a camera view and an instruction to move an object. The output needs to guide movement, not merely describe the object. ### Different models can use different robot bodies [Gemini Robotics](https://deepmind.google/blog/gemini-robotics-brings-ai-into-the-physical-world/) is described by its developers as a VLA model. Their work includes two-arm platforms and adaptation to other robot forms. A VLA is therefore a model type, not a synonym for a [humanoid robot](https://xumanoids.com/glossary/humanoid-robot). ### A model is part of a robot system Camera input, action representation, controllers, hardware, and operating conditions affect what a system can do. The RT-2 and Gemini Robotics reports describe particular training setups and evaluations. Their results should be read in that context, rather than as evidence that every VLA can perform every physical task. ### Sources - [Google DeepMind: RT-2 translates vision and language into action](https://deepmind.google/blog/rt-2-new-model-translates-vision-and-language-into-action/) - [Google DeepMind: Gemini Robotics brings AI into the physical world](https://deepmind.google/blog/gemini-robotics-brings-ai-into-the-physical-world/) --- ## Visual servoing https://xumanoids.com/glossary/visual-servoing Visual servoing uses visual measurements inside a feedback loop to control robot motion. The controller updates movement to reduce an error defined from image features or visually estimated pose. Updated: 2026-10-05 Also known as: Visual servo control ### Vision stays in the feedback loop A robot can repeatedly observe a target and adjust its motion as the observed error changes. This is different from taking one image, planning once, and then executing without visual correction. [Chaumette and Hutchinson's tutorial](https://inria.hal.science/inria-00350283) develops the control formulation for this closed-loop use of visual information. For example, a wrist camera can guide a gripper toward an object by aligning selected visual features with their desired positions. ### Image-based and position-based control Image-based visual servoing defines its error directly from image features. Position-based visual servoing uses visual data to estimate a three-dimensional pose and controls an error in that space. The tutorial compares the two approaches and their stability and performance characteristics. ### Geometry and visibility constrain performance The relationship between camera motion and image-feature motion enters the control law. Calibration, feature depth, and the current configuration affect this relationship. Features can also leave the camera's field of view during movement. A visually decreasing error is therefore not enough to establish collision-free motion or joint-limit compliance. Those constraints need to be addressed by the larger robot-control system. ### Sources - [Chaumette and Hutchinson: Visual Servo Control, Part I, Basic Approaches](https://inria.hal.science/inria-00350283) --- ## Visual-inertial odometry https://xumanoids.com/glossary/visual-inertial-odometry Visual-inertial odometry estimates a moving system's motion by combining camera observations with inertial measurements. It typically estimates position, orientation, velocity, and sensor biases over time. Updated: 2026-10-05 Also known as: VIO ### Cameras and inertial sensors contribute different information Images constrain motion through visible scene features. An [inertial measurement unit](https://xumanoids.com/glossary/inertial-measurement-unit) supplies rapid measurements of angular velocity and specific force. [Forster et al.](https://arxiv.org/abs/1512.02363) explain how these complementary inputs can support motion estimation and recover metric scale that a monocular camera alone cannot directly observe. A head-mounted camera and IMU can, for example, help estimate a humanoid's motion through a room. This estimate still needs consistent sensor frames and calibration. ### Preintegration keeps the problem manageable Inertial readings often arrive more frequently than camera frames. The cited method combines many readings between selected image frames into relative-motion constraints. Its estimator also accounts for IMU bias, rather than treating every acceleration measurement as exact. ### Local motion is not a permanent global reference VIO is a form of [odometry](https://xumanoids.com/glossary/odometry). Without additional global constraints, accumulated error can grow. [ROS frame conventions](https://www.ros.org/reps/rep-0105.html) distinguish a continuous local odometry frame from a globally corrected map frame. Adding place recognition and loop closure changes the larger system into a visual-inertial SLAM pipeline. ### Sources - [Forster et al.: On-Manifold Preintegration for Real-Time Visual-Inertial Odometry](https://arxiv.org/abs/1512.02363) - [ROS REP 105: Coordinate Frames for Mobile Platforms](https://www.ros.org/reps/rep-0105.html) --- ## Whole-body control https://xumanoids.com/glossary/whole-body-control Whole-body control coordinates a robot’s joints and contacts to satisfy several motion and force objectives together. In humanoids, it commonly combines balance, foot motion, hand tasks, and posture subject to physical constraints. Updated: 2026-10-05 Also known as: WBC, Whole body control ### Coordinating tasks that share joints A humanoid reaching for a shelf cannot choose arm motion independently of balance and leg motion. Whole-body control expresses these requirements together, often assigning priorities or weights so that a hand task can yield when maintaining support requires it. The [survey by Wensing and colleagues](https://arxiv.org/abs/2211.11644) describes widely used optimization formulations that compute joint commands from the current state. These can include contact constraints, joint limits, actuator effort limits, and desired task-space accelerations. Quadratic programs are common, but the term covers a broader family of control methods. ### From a plan to motor commands A higher-level planner may provide footsteps and a [centroidal](https://xumanoids.com/glossary/centroidal-dynamics) motion plan. The whole-body controller then uses the full robot model to realize those references through joint motion and contact forces. [Operational-space control](https://xumanoids.com/glossary/operational-space-control) supplies a foundation for expressing task priorities, as described in [MIT's manipulation notes](https://manipulation.mit.edu/force.html). ### Coordination cannot remove physical conflicts If a hand target requires a joint to exceed its range while both feet remain fixed, all objectives cannot be satisfied exactly. The result depends on which constraints are hard and which goals are allowed error. Whole-body control also depends on the assumed contacts and estimated state; it does not itself guarantee that a planned foothold is present or that a slipping foot remains fixed. ### Sources - [Wensing et al.: Optimization-Based Control for Dynamic Legged Robots](https://arxiv.org/abs/2211.11644) - [MIT Robotic Manipulation: Force Control](https://manipulation.mit.edu/force.html) --- ## Workspace https://xumanoids.com/glossary/workspace A robot workspace is the set of positions or poses its end-effector can reach under specified geometric and joint constraints. Its meaning depends on whether orientation is included and which base and tool configuration are assumed. Updated: 2026-10-05 Also known as: Robot workspace, Reachable workspace ### Reach includes assumptions A reach envelope shows where a robot's [end-effector](https://xumanoids.com/glossary/end-effector) can go. Link geometry and joint travel limits shape that envelope. Specifying the robot base and tool frame matters because changing either changes the positions being described. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-5-task-space-and-workspace/) distinguishes a position-only workspace from a workspace that includes orientation. A hand may reach a point without being able to approach it from every direction. ### Reachable and dexterous workspaces The dexterous workspace is commonly the subset of positions reachable with all end-effector orientations. It is therefore more restrictive than simply reaching each position with at least one orientation. Task space is different again: it is the space in which a task is naturally expressed. Drawing on a board can use a two-dimensional task space even when the robot itself has more [degrees of freedom](https://xumanoids.com/glossary/degrees-of-freedom). ### Reachability is not a motion guarantee A reachable target does not identify which joint posture to use or a route to that posture. Those are [inverse kinematics](https://xumanoids.com/glossary/inverse-kinematics) and [motion planning](https://xumanoids.com/glossary/motion-planning) questions. Reach also does not imply equal motion capability everywhere: the [singularity lesson](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) shows how a fully extended arm can lose an instantaneous direction of end-effector motion. ### Sources - [Modern Robotics: Task Space and Workspace](https://modernrobotics.northwestern.edu/nu-gm-book-resource/2-5-task-space-and-workspace/) - [Modern Robotics: Singularities](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) --- ## World model https://xumanoids.com/glossary/world-model A world model is an internal predictive model of an environment and how it changes. In robot learning, it can predict future states or observations under possible actions to support planning or policy training. Updated: 2026-10-05 Also known as: World models ### Predict what may happen next [Ha and Schmidhuber's World Models](https://worldmodels.github.io/) learns a compressed representation of observations and a recurrent model of how that representation evolves. A controller can use features from the model, and the authors also train an agent inside a generated environment. For a robot, a corresponding use would be predicting how an object might move after a push. The prediction could concern a compact state representation instead of a photorealistic future image. ### Prediction and control remain different roles A world model predicts outcomes; a policy chooses actions. One system can contain both, but a predictive video model does not automatically provide a robot controller. The cited work evaluates simulated reinforcement-learning environments. Its generated-environment training demonstrates the method in that setting, rather than establishing that any visually plausible model accurately predicts physical robot contact. ### Controllers can exploit model errors The authors describe a controller finding behavior that exploits mistakes in its learned environment. Those opportunities did not behave the same way in the original environment. This makes model fidelity relevant to the task and to the actions considered. A model that reconstructs familiar observations well can still make poor predictions for unfamiliar actions. [System identification](https://xumanoids.com/glossary/system-identification) and world-model learning overlap where both estimate dynamics from data, although world-model architectures need not use explicit mechanical equations. ### Sources - [Ha and Schmidhuber: World Models](https://worldmodels.github.io/) --- ## Wrench https://xumanoids.com/glossary/wrench A wrench is a six-component representation of force and moment acting on a rigid body. It combines three force components with three moment components about a specified reference point. Updated: 2026-10-05 Also known as: Force-moment wrench, Spatial wrench ### A force can also produce a moment A downward load at a robot hand produces a force at the wrist and may also produce a turning moment, depending on its offset from the wrist. A wrench records both effects. [Modern Robotics](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-4-wrenches/) illustrates this with an apple held by a hand fitted with a force-torque sensor. The same gravitational load has zero moment about the apple's center of mass but a nonzero moment about an offset sensor frame. ### The reference point matters Forces and moments must be expressed in a declared frame. Moving the reference point changes the moment through the force's moment arm, even when the physical loading stays the same. Moment and force also have different units, conventionally newton-metres and newtons. A wrench and a [twist](https://xumanoids.com/glossary/twist), expressed with compatible conventions, have a dot product equal to instantaneous mechanical power. That power is independent of the coordinate frame. ### Connecting contact loads to joints The [Jacobian-transpose relationship](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) maps an end-effector wrench to corresponding joint forces and torques under the stated static convention. It is useful in [force control](https://xumanoids.com/glossary/force-control) and load analysis. During acceleration, inertial and other dynamic terms also matter; a wrench transformation alone is not a complete dynamics calculation. ### Sources - [Modern Robotics: Wrenches](https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-4-wrenches/) - [Modern Robotics: Singularities](https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/) --- ## Xumanoid https://xumanoids.com/glossary/xumanoid Xumanoid is a coined term used on this site for a humanoid robot that understands context, learns from experience, and works alongside people. It describes an intended combination of humanlike form and collaborative behaviour, rather than an established technical class. Updated: 2026-10-05 Also known as: Xumanoids ### How this site uses the word The [original definition](/) uses xumanoids, the plural of xumanoid, for humanoid robots that understand context, learn from experience, and work alongside people. Its broader meaning describes machines in human form built to expand people's abilities and serve them with care. ### A coined label, not a standard The word expresses this site's intended meaning. It is not a certification, a robotics standard, or evidence that a particular robot already has those capabilities. Claims about a robot still need evidence about its actual behaviour and operating conditions. ### Related technical concepts [Humanoid robot](https://xumanoids.com/glossary/humanoid-robot) describes physical form. [Embodied AI](https://xumanoids.com/glossary/embodied-ai) describes intelligence that perceives and acts through a body. [Robotic manipulation](https://xumanoids.com/glossary/robotic-manipulation) describes interaction with objects. Those concepts help explain a robot's parts and capabilities without treating the coined label as a technical specification. ### Sources - [The site's original definition](https://xumanoids.com/) --- ## Yaw https://xumanoids.com/glossary/yaw Yaw is the rotation angle about the z-axis in a specified roll-pitch-yaw convention. For a level robot in a z-up frame, it describes heading in the horizontal plane. Updated: 2026-10-05 Also known as: Yaw angle ### Turning in a declared frame For an upright humanoid, a change in yaw can describe turning to face left or right. Roll and pitch describe the other two rotational components. [ROS REP 103](https://www.ros.org/reps/rep-0103.html) defines its standard body axes as x forward, y left, and z up, and associates fixed-axis roll, pitch, and yaw with x, y, and z respectively. The reference frame and rotation convention are part of the meaning. A yaw value without them is an incomplete orientation specification. ### Positive yaw and compass bearings differ Under ROS's right-handed convention, positive yaw turns counterclockwise when viewed from above a z-up frame. In a geographic east-north-up frame, zero yaw points east. REP 103 explicitly contrasts this with a compass bearing, which starts at north and increases clockwise. A navigation system must convert between these conventions before comparing sensor readings or commanding a turn. ### Heading is only one orientation component Yaw alone does not describe a tilted robot's full orientation. [Euler angles](https://xumanoids.com/glossary/euler-angles) require a specified sequence, while a [quaternion](https://xumanoids.com/glossary/quaternion) represents the complete rotation without angle-coordinate singularities. REP 103 recommends quaternions or rotation matrices ahead of angle representations for exchanging orientation data. ### Sources - [ROS REP 103: Standard Units of Measure and Coordinate Conventions](https://www.ros.org/reps/rep-0103.html) --- ## Young's modulus https://xumanoids.com/glossary/youngs-modulus Young's modulus is a measure of material stiffness equal to axial stress divided by axial strain in the linear elastic regime. It describes resistance to elastic stretching or compression, rather than the load at which a part fails. Updated: 2026-10-05 Also known as: Young modulus, Young’s modulus ### The slope of the elastic stress-strain curve In a uniaxial test, stress is force divided by cross-sectional area and strain is extension divided by original length. [MIT's stress-strain notes](https://web.mit.edu/course/3/3.11/www/modules/ss.pdf) define Young's modulus, usually written `E`, as the proportionality constant where stress and strain have a linear relationship. Its units are pascals because strain is dimensionless. Material selection for a robot's structural or elastic components needs this relationship to estimate deformation under load. ### Material modulus and part stiffness differ Combining the notes' stress and strain definitions gives the axial stiffness of a uniform, linearly elastic bar as `k = E A / L`, where `A` is area and `L` is length. Geometry therefore matters as well as material: a longer bar stretches more under the same force if its material and area stay unchanged. For a [series elastic actuator](https://xumanoids.com/glossary/series-elastic-actuator), the complete spring geometry determines the effective stiffness used to relate deflection to force or torque. ### Stiffness does not establish strength The MIT notes distinguish elastic slope, yield stress, and ultimate tensile strength. A high modulus does not by itself establish a high failure load. The linear relation also stops describing the whole response when the material enters nonlinear deformation or plastic flow. Use the relevant stress-strain range rather than extrapolating one modulus to every load. ### Sources - [MIT, David Roylance: Stress-Strain Curves](https://web.mit.edu/course/3/3.11/www/modules/ss.pdf) --- ## Zero-moment point https://xumanoids.com/glossary/zero-moment-point The zero-moment point is a point on a chosen support plane where the net moment associated with the ground reaction wrench has zero components parallel to that plane. In flat-ground walking with the usual contact assumptions, it coincides with the center of pressure. Updated: 2026-10-05 Also known as: ZMP, Zero moment point ### Which moment becomes zero The name does not mean that every torque acting on the robot vanishes. In three dimensions, the two moment components parallel to the support plane are zero; a moment about the plane's normal can remain. [MIT's derivation](https://underactuated.mit.edu/humanoids.html) distinguishes this from the simpler planar case. For coplanar ground contacts that push without adhesion, with a nonzero total normal force, the center of pressure lies inside the convex hull of the active contact points. Under these assumptions, that center of pressure is the physical ZMP. ### Planning center-of-mass motion A walking planner can choose a ZMP trajectory inside the [support polygon](https://xumanoids.com/glossary/support-polygon) and derive a compatible [center-of-mass](https://xumanoids.com/glossary/center-of-mass) trajectory. Constant center-of-mass height and negligible change in angular momentum lead to the familiar [linear inverted pendulum model](https://xumanoids.com/glossary/linear-inverted-pendulum-model). ### Limits of the balance criterion A desired ZMP inside the feet is a contact-moment condition, not a complete guarantee of balance. Friction, joint torque, reachability, and future motion must also be feasible. The [MIT notes](https://underactuated.mit.edu/humanoids.html) warn that simplified planning can miss joint constraints. Uneven contacts, hand support, flight, and substantial angular-momentum changes require a model that represents those conditions explicitly. ### Sources - [MIT Underactuated Robotics: Highly-articulated Legged Robots](https://underactuated.mit.edu/humanoids.html)