Skip to main content
Open In Colab

Camera Models

Camera models describe how a camera projects points in the 3D world onto a 2D image plane. Each model makes different assumptions. Common camera models include:
  1. The Pinhole Camera Model: The pinhole model is the simplest camera model. It assumes that light rays pass through a single point, called the pinhole, before reaching the image plane. Its parameters include the focal length and the position of the pinhole.
  2. The Perspective Camera Model: The perspective model extends the pinhole model to include lens distortion and other optical effects. It is used in computer vision and graphics to model camera behavior.
  3. The Orthographic Camera Model: In the orthographic model, parallel lines in the 3D world remain parallel in the 2D image. This property is useful for technical drawings and architectural visualizations that require accurate measurements.
  4. The Spherical Camera Model: The spherical model uses a spherical image sensor to capture a 360-degree view of a scene. It is used in virtual reality and panoramic photography.
  5. The Omnidirectional Camera Model: The omnidirectional model uses multiple lenses or a fisheye lens to capture a wide field of view (FOV). It is used in robotics and surveillance.
Each model defines equations and parameters for projecting 3D points onto the 2D image plane. These models are used for image rectification, 3D reconstruction, and camera calibration. This section focuses on the pinhole camera model. Pinhole camera model. The image is formed on the image plane by light rays passing through a small aperture (the pinhole) at the center of projection. The image is inverted and smaller than the object Pinhole camera model. The image is formed on the image plane by light rays passing through a small aperture (the pinhole) at the center of projection. The image is inverted and smaller than the object.

Camera Model Fundamentals

Pinhole Camera Model

The functions in this section use the pinhole camera model. A perspective transformation projects a 3D scene point PwP_w onto the image plane at pixel pp. The variables PwP_w and pp are represented in homogeneous coordinates as 3D and 2D vectors, respectively. The pinhole camera model gives the following distortion-free projective transformation: λ  p=K[R∣t]Pw\lambda \; p = K \begin{bmatrix} R|t \end{bmatrix} P_w where:
  • PwP_w is a 3D point expressed with respect to the world coordinate system
  • pp is a 2D pixel in the image plane
  • KK is the camera intrinsic matrix
  • RR and tt are the rotation and translation that describe the change of coordinates from world to camera coordinate systems
  • λ\lambda is the projective transformation’s arbitrary scaling

Camera Intrinsic Matrix

The camera intrinsic matrix KK maps 3D points in the camera coordinate system to 2D pixel coordinates, up to a scale factor: λ p=KPc,p=[uv1],λ=Zc\lambda\, p = K P_c, \qquad p = \begin{bmatrix} u \\ v \\ 1 \end{bmatrix}, \quad \lambda = Z_c Dividing by λ\lambda is the perspective division that turns KPcK P_c into pixel coordinates. The matrix KK contains the focal lengths fxf_x and fyf_y, expressed in pixels, and the principal point (cx,cy)(c_x, c_y), which is usually near the image center: K=[fx0cx0fycy001]K = \begin{bmatrix} f_x & 0 & c_x \\ 0 & f_y & c_y \\ 0 & 0 & 1 \end{bmatrix} Therefore: λ[uv1]=[fx0cx0fycy001][XcYcZc]\lambda \begin{bmatrix} u \\ v \\ 1 \end{bmatrix} = \begin{bmatrix} f_x & 0 & c_x \\ 0 & f_y & c_y \\ 0 & 0 & 1 \end{bmatrix} \begin{bmatrix} X_c \\ Y_c \\ Z_c \end{bmatrix}

Coordinate Transformations

The joint rotation-translation matrix [R∣t][R|t] factors into a 3-by-4 projection and a 4-by-4 homogeneous transformation (Hartley & Zisserman, 2004; Szeliski, n.d.): [R∣t]=[I0][Rt01][R|t] = \begin{bmatrix} I & 0 \end{bmatrix} \begin{bmatrix} R & t \\ 0 & 1 \end{bmatrix} The homogeneous transformation, built from the extrinsic parameters, maps a world point to camera coordinates (Xc,Yc,Zc)(X_c, Y_c, Z_c). The projection [I0]\begin{bmatrix} I & 0 \end{bmatrix} keeps these three coordinates and drops the homogeneous one. Read as a homogeneous 2D point, the result gives the normalized camera coordinates x′=Xc/Zcx' = X_c / Z_c and y′=Yc/Zcy' = Y_c / Z_c on the image plane. The extrinsic parameters RR and tt define the homogeneous transformation from the world coordinate system ww to the camera coordinate system cc: Pc=[Rt01]PwP_c = \begin{bmatrix} R & t \\ 0 & 1 \end{bmatrix} P_w The complete transformation is: λ[uv1]=[fx0cx0fycy001][r11r12r13txr21r22r23tyr31r32r33tz][XwYwZw1]\lambda \begin{bmatrix} u \\ v \\ 1 \end{bmatrix} = \begin{bmatrix} f_x & 0 & c_x \\ 0 & f_y & c_y \\ 0 & 0 & 1 \end{bmatrix} \begin{bmatrix} r_{11} & r_{12} & r_{13} & t_x \\ r_{21} & r_{22} & r_{23} & t_y \\ r_{31} & r_{32} & r_{33} & t_z \end{bmatrix} \begin{bmatrix} X_w \\ Y_w \\ Z_w \\ 1 \end{bmatrix} If Zc≠0Z_c \neq 0, this is equivalent to: [uv]=[fxXc/Zc+cxfyYc/Zc+cy]\begin{bmatrix} u \\ v \end{bmatrix} = \begin{bmatrix} f_x X_c/Z_c + c_x \\ f_y Y_c/Z_c + c_y \end{bmatrix}

Lens Distortion Model

Real lenses introduce radial and tangential distortion. Distortion examples: barrel, pincushion, and tangential distortions Distortion examples: barrel, pincushion, and tangential distortions. The extended camera model includes this distortion: [uv]=[fxx′′+cxfyy′′+cy]\begin{bmatrix} u \\ v \end{bmatrix} = \begin{bmatrix} f_x x'' + c_x \\ f_y y'' + c_y \end{bmatrix} where: [x′′y′′]=[x′1+k1r2+k2r4+k3r61+k4r2+k5r4+k6r6+2p1x′y′+p2(r2+2x′2)+s1r2+s2r4y′1+k1r2+k2r4+k3r61+k4r2+k5r4+k6r6+p1(r2+2y′2)+2p2x′y′+s3r2+s4r4]\begin{bmatrix} x'' \\ y'' \end{bmatrix} = \begin{bmatrix} x' \frac{1 + k_1 r^2 + k_2 r^4 + k_3 r^6}{1 + k_4 r^2 + k_5 r^4 + k_6 r^6} + 2 p_1 x' y' + p_2(r^2 + 2 x'^2) + s_1 r^2 + s_2 r^4 \\ y' \frac{1 + k_1 r^2 + k_2 r^4 + k_3 r^6}{1 + k_4 r^2 + k_5 r^4 + k_6 r^6} + p_1 (r^2 + 2 y'^2) + 2 p_2 x' y' + s_3 r^2 + s_4 r^4 \end{bmatrix} with r2=x′2+y′2r^2 = x'^2 + y'^2 and [x′y′]=[Xc/ZcYc/Zc]\begin{bmatrix} x' \\ y' \end{bmatrix} = \begin{bmatrix} X_c/Z_c \\ Y_c/Z_c \end{bmatrix} if Zc≠0Z_c \neq 0. Distortion parameters:
  • Radial coefficients: k1k_1, k2k_2, k3k_3, k4k_4, k5k_5, k6k_6
  • Tangential coefficients: p1p_1, p2p_2
  • Thin prism coefficients: s1s_1, s2s_2, s3s_3, s4s_4
The distortion coefficients are passed as: (k1,k2,p1,p2[,k3[,k4,k5,k6[,s1,s2,s3,s4[,τx,τy]]]])(k_1, k_2, p_1, p_2[, k_3[, k_4, k_5, k_6 [, s_1, s_2, s_3, s_4[, \tau_x, \tau_y]]]]) Types of distortion: The radial part of the model scales each point by the full radial factor ρ(r)=1+k1r2+k2r4+k3r61+k4r2+k5r4+k6r6,\rho(r) = \frac{1 + k_1 r^2 + k_2 r^4 + k_3 r^6}{1 + k_4 r^2 + k_5 r^4 + k_6 r^6}, so the type of radial distortion depends on how the whole ratio changes with rr, not on the numerator alone:
  • Barrel distortion: ρ(r)\rho(r) decreases as rr grows, so points far from the center move inward.
  • Pincushion distortion: ρ(r)\rho(r) increases as rr grows, so points far from the center move outward.

Right-handed and Left-handed Coordinate Systems

Right-handed and left-handed coordinate systems use different conventions for orienting axes in 3D space. Applications may use either convention, so you must know which one applies.

ROS2

ROS2 RViz2 uses a right-handed coordinate system. The right-hand rule determines the axis directions. Point your thumb along the X-axis and your index finger along the Y-axis. Your middle finger then points along the Z-axis.

Sensor Coordinate Systems

Each sensor uses a coordinate system documented by the vendor. It may be right-handed or left-handed. RealSense cameras use a right-handed coordinate system.
  • Point of view:
    • Imagine that you are standing behind the camera and looking forward.
    • Use this point of view when describing coordinates, left and right infrared sensors, and sensor positions.
image
  • ROS2 Coordinate System: (X: Forward, Y: Left, Z: Up)
  • Camera Optical Coordinate System: (X: Right, Y: Down, Z: Forward)
  • References: REP-0103, REP-0105
  • All data published by the RealSense wrapper is optical data taken directly from the camera sensors.
  • Static and dynamic TF topics publish the optical and ROS coordinate systems. You can use these transforms to move from one coordinate system to the other.

TF from Coordinate A to Coordinate B

  • A TF message gives the pose of the coordinate frame child_frame_id in the coordinate frame “header.frame_id”. Applied to a point, it maps coordinates in child_frame_id into coordinates in header.frame_id; mapping the other way needs its inverse. The ROS documentation describes this as a transform from header.frame_id to child_frame_id, which refers to the frames, not to the direction in which coordinates are mapped. See the reference.
  • In RealSense cameras, the origin (0,0,0) is the position of the left infrared sensor (infra1). This origin defines the “camera_link” frame.
  • The depth, left infrared, and “camera_link” coordinate frames coincide.
  • The wrapper provides static TFs from each sensor coordinate frame to the camera base (camera_link).
  • It also provides TFs from each sensor ROS coordinate frame to the corresponding optical coordinate frame.
  • The following image shows the static TFs for the RGB sensor and the Infra2 (right infrared) sensor of a D435i module in RViz2: example

Extrinsics from Sensor A to Sensor B

  • The extrinsic transform from sensor A to sensor B gives the position and orientation of sensor A relative to sensor B.
  • If B is the origin (0,0,0), then Extrinsics(A->B) describes the position of sensor A relative to sensor B.
  • For example, consider depth_to_color for the D435i:
    • Viewed from behind the D435i, the extrinsic transform from depth to color specifies the position of the depth sensor relative to the color sensor.
    • Consider only the X coordinates in the optical coordinate system, viewed from behind. If the COLOR (RGB) sensor is at (0,0,0), the DEPTH sensor is 0.0148 m (1.48 cm) to its right.
d435i
  • An extrinsic message has two parts:
    • float64[9] rotation (column-major 3x3 rotation matrix)
    • float64[3] translation (three-element translation vector in meters)
When exporting vertices to PLY format, RealSense SDK versions 2.19.0 and later convert the points to a left-handed coordinate system. PLY is a common 3D file format. The conversion provides compatibility with the default left-handed viewpoint in MeshLab.

References

  • Hartley, R., Zisserman, A. (2004). Multiple View Geometry in Computer Vision.
  • Szeliski, R. (n.d.). Computer Vision: Algorithms and Applications.