Robotics

Bundle adjustment

Definition

Bundle adjustment is the joint nonlinear refinement of camera or robot poses, three-dimensional scene points, and sometimes calibration parameters by minimising image reprojection error. It is used to make a multi-view reconstruction geometrically consistent.

Also known as: BA

Updated

Adjusting views and structure together

A feature observed in several images gives a measured pixel location in each view. A camera model predicts where a candidate 3D point would appear from a candidate camera pose. Bundle adjustment changes the poses and point positions together so that the predicted and measured image positions agree as closely as possible.

The residual between those positions is the reprojection error. The Ceres Solver example defines bundle adjustment as finding 3D points and camera parameters that minimise this error. Implementations may also optimise focal length, lens distortion, rig extrinsics, or other calibration variables.

Use in robot mapping

Visual simultaneous localization and mapping can run local bundle adjustment over a recent window of keyframes or a larger optimisation after a loop closure. Multi-camera robots can use it to refine both the trajectory and landmark map. Structure-from-motion systems use the same principle when reconstructing a scene from an unordered image collection.

Triggs and colleagues present bundle adjustment as a sparse nonlinear least-squares problem that jointly refines viewing and scene parameters. It can be represented as a factor graph, but the terms are not equivalent. A factor graph is a general dependency representation; bundle adjustment names the particular multi-view refinement problem.

Bundle adjustment also differs from point cloud registration. Registration aligns geometric point sets, while bundle adjustment normally works from image observations and a projection model. A reconstruction pipeline may use both.

Initialisation and observations matter

The objective is nonlinear and can converge to a poor local solution if poses, correspondences, or calibration start far from the correct values. Outlier feature matches can pull the estimate away from the scene, so practical solvers use outlier-resistant loss functions and careful track filtering.

Some degrees of freedom are unobservable without a gauge choice. For example, a monocular reconstruction may be determined only up to a global scale, rotation, and translation. Rolling shutters, moving objects, timing errors, or an inaccurate lens model also violate the usual static-scene projection assumptions. Large maps require sparse linear algebra, windowing, or marginalisation to control computation and memory.

Sources