Camera models describe how a camera projects points in the 3D world onto a 2D image plane. Each model makes different assumptions. Common camera models include:
The Pinhole Camera Model: The pinhole model is the simplest camera model. It assumes that light rays pass through a single point, called the pinhole, before reaching the image plane. Its parameters include the focal length and the position of the pinhole.
The Perspective Camera Model: The perspective model extends the pinhole model to include lens distortion and other optical effects. It is used in computer vision and graphics to model camera behavior.
The Orthographic Camera Model: In the orthographic model, parallel lines in the 3D world remain parallel in the 2D image. This property is useful for technical drawings and architectural visualizations that require accurate measurements.
The Spherical Camera Model: The spherical model uses a spherical image sensor to capture a 360-degree view of a scene. It is used in virtual reality and panoramic photography.
The Omnidirectional Camera Model: The omnidirectional model uses multiple lenses or a fisheye lens to capture a wide field of view (FOV). It is used in robotics and surveillance.
Each model defines equations and parameters for projecting 3D points onto the 2D image plane. These models are used for image rectification, 3D reconstruction, and camera calibration.This section focuses on the pinhole camera model.Pinhole camera model. The image is formed on the image plane by light rays passing through a small aperture (the pinhole) at the center of projection. The image is inverted and smaller than the object.
The functions in this section use the pinhole camera model. A perspective transformation projects a 3D scene point Pw onto the image plane at pixel p. The variables Pw and p are represented in homogeneous coordinates as 3D and 2D vectors, respectively.The pinhole camera model gives the following distortion-free projective transformation:λp=K[R∣t]Pwwhere:
Pw is a 3D point expressed with respect to the world coordinate system
p is a 2D pixel in the image plane
K is the camera intrinsic matrix
R and t are the rotation and translation that describe the change of coordinates from world to camera coordinate systems
λ is the projective transformation’s arbitrary scaling
The camera intrinsic matrix K maps 3D points in the camera coordinate system to 2D pixel coordinates, up to a scale factor:λp=KPc,p=uv1,λ=ZcDividing by λ is the perspective division that turns KPc into pixel coordinates.The matrix K contains the focal lengths fx and fy, expressed in pixels, and the principal point (cx,cy), which is usually near the image center:K=fx000fy0cxcy1Therefore:λuv1=fx000fy0cxcy1XcYcZc
The joint rotation-translation matrix [R∣t] factors into a 3-by-4 projection and a 4-by-4 homogeneous transformation (Hartley & Zisserman, 2004; Szeliski, n.d.):[R∣t]=[I0][R0t1]The homogeneous transformation, built from the extrinsic parameters, maps a world point to camera coordinates (Xc,Yc,Zc). The projection [I0] keeps these three coordinates and drops the homogeneous one. Read as a homogeneous 2D point, the result gives the normalized camera coordinates x′=Xc/Zc and y′=Yc/Zc on the image plane.The extrinsic parameters R and t define the homogeneous transformation from the world coordinate system w to the camera coordinate system c:Pc=[R0t1]PwThe complete transformation is:λuv1=fx000fy0cxcy1r11r21r31r12r22r32r13r23r33txtytzXwYwZw1If Zc=0, this is equivalent to:[uv]=[fxXc/Zc+cxfyYc/Zc+cy]
Real lenses introduce radial and tangential distortion.Distortion examples: barrel, pincushion, and tangential distortions.The extended camera model includes this distortion:[uv]=[fxx′′+cxfyy′′+cy]where:[x′′y′′]=[x′1+k4r2+k5r4+k6r61+k1r2+k2r4+k3r6+2p1x′y′+p2(r2+2x′2)+s1r2+s2r4y′1+k4r2+k5r4+k6r61+k1r2+k2r4+k3r6+p1(r2+2y′2)+2p2x′y′+s3r2+s4r4]with r2=x′2+y′2 and [x′y′]=[Xc/ZcYc/Zc] if Zc=0.Distortion parameters:
Radial coefficients: k1, k2, k3, k4, k5, k6
Tangential coefficients: p1, p2
Thin prism coefficients: s1, s2, s3, s4
The distortion coefficients are passed as:(k1,k2,p1,p2[,k3[,k4,k5,k6[,s1,s2,s3,s4[,τx,τy]]]])Types of distortion:The radial part of the model scales each point by the full radial factorρ(r)=1+k4r2+k5r4+k6r61+k1r2+k2r4+k3r6,so the type of radial distortion depends on how the whole ratio changes with r, not on the numerator alone:
Barrel distortion: ρ(r) decreases as r grows, so points far from the center move inward.
Pincushion distortion: ρ(r) increases as r grows, so points far from the center move outward.
Right-handed and left-handed coordinate systems use different conventions for orienting axes in 3D space. Applications may use either convention, so you must know which one applies.
ROS2 RViz2 uses a right-handed coordinate system. The right-hand rule determines the axis directions. Point your thumb along the X-axis and your index finger along the Y-axis. Your middle finger then points along the Z-axis.
Each sensor uses a coordinate system documented by the vendor. It may be right-handed or left-handed. RealSense cameras use a right-handed coordinate system.
Point of view:
Imagine that you are standing behind the camera and looking forward.
Use this point of view when describing coordinates, left and right infrared sensors, and sensor positions.
All data published by the RealSense wrapper is optical data taken directly from the camera sensors.
Static and dynamic TF topics publish the optical and ROS coordinate systems. You can use these transforms to move from one coordinate system to the other.
A TF message gives the pose of the coordinate frame child_frame_id in the coordinate frame “header.frame_id”. Applied to a point, it maps coordinates in child_frame_id into coordinates in header.frame_id; mapping the other way needs its inverse. The ROS documentation describes this as a transform from header.frame_id to child_frame_id, which refers to the frames, not to the direction in which coordinates are mapped. See the reference.
In RealSense cameras, the origin (0,0,0) is the position of the left infrared sensor (infra1). This origin defines the “camera_link” frame.
The depth, left infrared, and “camera_link” coordinate frames coincide.
The wrapper provides static TFs from each sensor coordinate frame to the camera base (camera_link).
It also provides TFs from each sensor ROS coordinate frame to the corresponding optical coordinate frame.
The following image shows the static TFs for the RGB sensor and the Infra2 (right infrared) sensor of a D435i module in RViz2:
The extrinsic transform from sensor A to sensor B gives the position and orientation of sensor A relative to sensor B.
If B is the origin (0,0,0), then Extrinsics(A->B) describes the position of sensor A relative to sensor B.
For example, consider depth_to_color for the D435i:
Viewed from behind the D435i, the extrinsic transform from depth to color specifies the position of the depth sensor relative to the color sensor.
Consider only the X coordinates in the optical coordinate system, viewed from behind. If the COLOR (RGB) sensor is at (0,0,0), the DEPTH sensor is 0.0148 m (1.48 cm) to its right.
float64[3] translation (three-element translation vector in meters)
When exporting vertices to PLY format, RealSense SDK versions 2.19.0 and later convert the points to a left-handed coordinate system. PLY is a common 3D file format. The conversion provides compatibility with the default left-handed viewpoint in MeshLab.