Perspective-n-Point (PnP): given correspondences between 3D reference points and their 2D image projections, their world and image coordinates, and camera intrinsics K, estimate the pose transformation between world and camera frames. PnP is used for camera and object tracking, AR/VR, robot manipulation, and SLAM pose initialization. Common solvers include DLT, P3P, EPnP, and UPnP.
1. Direct Linear Transform
The original DLT does not require known camera intrinsics; see Multiple View Geometry. Since intrinsics are usually known, this derivation includes them.
1.1 Derivation
Write a homogeneous 3D world reference point as
c=xyz1,
and its homogeneous 2D image projection as
u=uv1.
Camera intrinsics are
K=fx000fy0cxcy1.
The 3D-to-2D projection is
λuv1=K[R∣t]xyz1.(1)
Although [R∣t] has six degrees of freedom, DLT initially ignores the orthogonality constraint on R and treats it as twelve unknowns x=[a1,…,a12]T:
For an experiment, DLT was inserted into the ORB-SLAM2 frontend and compared with motion-only bundle adjustment (MOB). MOB produces a visibly smoother trajectory. At turns, fewer and more concentrated 3D points lead to obvious DLT error. DLT also does not address outliers, whereas MOB uses a Huber function, making MOB more stable. RANSAC can later address DLT outliers. With 200–300 correspondences, DLT runtime is about 0.07 ms on an Intel i7-8700 under Ubuntu 16.04.
References
14 Lectures on Visual SLAM.
Hartley R, Zisserman A. Multiple View Geometry in Computer Vision. 2003.