NOTE / 8/22/2018
[Fourteen Lectures 1] Summary of Chapters 1–2
I have recently been reading Sixteen Lectures on Visual SLAM. Here are my notes on several key ideas.
Chapter 1: Prerequisites
- SLAM stands for Simultaneous Localization and Mapping.
- Its main task is to help a device—such as a robot or phone—enter an unknown environment and: (1) estimate its own position (localization), and (2) perceive and model its surroundings (mapping).
- What makes the problem distinctive is that it is usually real-time, has little or no prior knowledge of the environment, and solves localization and mapping together, with each reinforcing the other.
- It can use many sensors: cameras, LiDAR, IMU, encoders, ultrasonic sensors, and more. SLAM can be viewed as a specialized state-estimation or multi-sensor-fusion problem.
One exercise asks how a system should be solved, which is a recurring question. Different parameter states call for different techniques, such as SVD and QR decomposition; bundle adjustment also uses the Schur complement for speed. I will discuss this separately later.
Chapter 2: A first look at SLAM
SLAM addresses localization together with environment modeling.
Compared with approaches that install external markers or sensors in an environment, SLAM is more general and flexible because it does not require modifying the environment. This is why modern AR localization systems, including Apple ARKit, Google ARCore, and Baidu DuMix, have sought to solve localization with SLAM.
Visual sensors: monocular, stereo, multi-camera, depth, event-based, infrared, and more.
- A monocular camera lacks absolute scale.
- Stereo cameras provide scale, but accuracy depends on baseline and object distance; they are computationally heavier and matching is difficult.
- RGB-D cameras provide depth at a limited range and are generally used indoors.
The classic visual-SLAM architecture

- Front end: abstracts sensor measurements, estimates the pose between successive frames as visual odometry, and builds a graph for back-end optimization. Because visual odometry estimates only frame-to-frame motion, drift accumulates; back-end optimization and loop closure address it.
- Back end: uses odometry and loop-closure information for global optimization, producing a globally consistent map and trajectory. It solves a maximum-a-posteriori estimation problem, mainly with filtering and nonlinear optimization.
- Loop closure: detects revisited places and supplies loop constraints.
- Mapping: constructs an environment model from the estimated trajectory. The kind of map is task-driven: metric, topological, semantic, or hybrid.
Classical SLAM commonly assumes: (1) a static environment; (2) a rigid environment; (3) limited lighting variation; and (4) no human interference.
A general mathematical expression
Different sensors and devices use different parameterizations.