Материал: [2.1] 3D Imaging, Analysis and Applications-Springer-Verlag London (2012)

Внимание! Если размещение файла нарушает Ваши авторские права, то обязательно сообщите нам

44

S. Se and N. Pears

2.A perspective projection from the 3D world to the 2D image plane.

3.A mapping from metric image coordinates to pixel coordinates.

We now discuss each of these projective mappings in turn.

2.3.2.1 Camera Modeling: The Coordinate Transformation

As shown in Fig. 2.4, the camera frame has its (X, Y ) plane parallel to the image plane and Z is in the direction of the principal axis of the lens and encodes depth

˜

from the camera. Suppose that the camera center has inhomogeneous position C in the world frame3 and the rotation of the camera frame is Rc relative to the world frame orientation. This means that we can express any inhomogeneous camera frame points as:

X˜ c = RcT (X˜ − C˜ ) = RX˜ + t.

(2.1)

= T = − T ˜

Here R Rc represents the rigid rotation and t Rc C represents the rigid translation that maps a scene point expressed in the world coordinate frame into a camera-centered coordinate frame. Equation (2.1) can be expressed as a projective mapping, namely one that is linear in homogeneous coordinates, to give:

Xc

 

R

 

Yc

 

Zc

=

0T

 

 

 

 

 

1

 

 

X

tY

. 1 Z

1

We denote Pr as the 4 × 4 homogeneous matrix representing the rigid coordinate transformation in the above equation.

2.3.2.2 Camera Modeling: Perspective Projection

Observing the similar triangles in the geometry of perspective imaging, we have

xc

=

Xc

,

yc

=

Yc

,

(2.2)

f

Zc

f

Zc

where (xc , yc ) is the position (metric units) of a point in the camera’s image plane and f is the distance (metric units) of the image plane to the camera center. (This is usually set to the focal length of the camera lens.) The two equations above can be written in linear form as:

 

 

xc

 

 

f

0

0

0

 

Xc

 

 

0

f 0

0

 

Yc

.

Zc

 

yc

 

=

 

Zc

 

1

0

0

1

0

 

1

 

 

 

 

 

 

 

 

 

 

 

 

 

3We use a tilde to differentiate n-tuple inhomogeneous coordinates from (n + 1)-tuple homogeneous coordinates.

2 Passive 3D Imaging

45

We denote Pp as the 3 × 4 perspective projection matrix, defined by the value of f , in the above equation. If we consider an abstract image plane at f = 1, then points on this plane are termed normalized image coordinates4 and from Eq. (2.2), these are given by

xn =

Xc

 

yn =

Yc

 

,

 

.

Zc

Zc

2.3.2.3 Camera Modeling: Image Sampling

Typically, the image on the image plane is sampled by an image sensor, such as a CCD or CMOS device, at the locations defined by an array of pixels. The final part of camera modeling defines how that array is positioned on the [xc , yc ]T image plane, so that pixel coordinates can be generated. In general, pixels in an image sensor are not square and the number of pixels per unit distance varies between the xc and yc directions; we will call these scalings mx and my . Note that pixel positions have their origin at the corner of the sensor and so the position of the principal point (where the principal axis intersects the image plane) is modeled with pixel coordinates [x0, y0]T . Finally, many camera models also cater for any skew,5 s, so that the mapping into pixels is given by:

x mx

s

x0 xc

 

y

 

=

0

m

y

y

1

0

0y

10

1c .

We denote Pc as the 3 × 3 projective matrix defined by the five parameters mx , my , s, x0 and y0 in the above equation.

2.3.2.4 Camera Modeling: Concatenating the Projective Mappings

We can concatenate the three stages described in the three previous subsections to give

λx = Pc Pp Pr X

or simply

λx = PX,

(2.3)

where λ is non-zero and positive. We note the following points concerning the above equation

4We need to use a variety of image coordinate normalizations in this chapter. For simplicity, we will use the same subscript n, but it will be clear about how the normalization is achieved.

5Skew models a lack of orthogonality between the two image sensor sampling directions. For most imaging situations it is zero.

46

S. Se and N. Pears

1.For any homogeneous image point scaled to λ[x, y, 1]T , the scale λ is equal to the imaged point’s depth in the camera centered frame (λ = Zc ).

2.Any non-zero scaling of the projection matrix λP P performs the same projection since, in Eq. (2.3), any non-zero scaling of homogeneous image coordinates is equivalent.

3.A camera with projection matrix P, or some non-zero scalar multiple of that, is informally referred to as camera P in the computer vision literature and, because of point 2 above, it is referred to as being defined up to scale.

The matrix P is a 3 × 4 projective camera matrix with the following structure:

P = K[R|t].

(2.4)

The parameters within K are the camera’s intrinsic parameters. These parameters are those combined from Sects. 2.3.2.2 and 2.3.2.3 above, so that:

 

 

αx

s

x0

 

K

=

0

α

y

,

 

0

0y

10

where αx = f mx and αy = f my represent the focal length in pixels in the x and y directions respectively. Together, the rotation and translation in Eq. (2.4) are termed the camera’s extrinsic parameters. Since there are 5 DOF from intrinsic parameters and 6 DOF from extrinsic parameters, a camera projection matrix has only 11 DOF, not the full 12 of a general 3 × 4 matrix. This is also evident from the fact that we are dealing with homogeneous coordinates and so the overall scale of P does not matter.

By expanding Eq. (2.3), we have:

 

homogeneous

 

 

 

 

 

 

 

 

 

 

 

 

 

 

homogeneous

 

 

 

 

 

intrinsic

 

 

 

extrinsic

 

 

 

world

 

 

 

image

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

camera

 

 

 

camera

 

 

coordinates

 

 

coordinates

 

 

parameters

 

 

parameters

 

 

 

 

 

 

 

 

 

 

 

x

 

 

αx

 

s

x0

 

r11

r12

r13

tx

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

X

 

 

 

 

 

 

 

0

 

 

 

 

 

 

 

 

 

 

 

 

Y

 

 

 

λ

 

y

=

αy

y0

 

r21

r22

r23

ty

 

Z

,

(2.5)

 

1

0

0 1

r31

r32

r33

tz

 

1

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

which indicates that both the intrinsic and extrinsic camera parameters are necessary to fully define a ray (metrically, not just in pixel units) in 3D space and hence make absolute measurements in multiple-view 3D reconstruction. Finally, we note that any non-zero scaling of scene homogeneous coordinates [X, Y, Z, 1]T in Eq. (2.5) gives the same image coordinates6 which, for a single image, can be interpreted as ambiguity between the scene scale and the translation vector t.

6The same homogeneous image coordinates up to scale or the same inhomogeneous image coordinates.

2 Passive 3D Imaging

47

Fig. 2.5 Examples of radial distortion effects in lenses: (a) No distortion (b) Pincushion distortion (c) Barrel distortion (d) Fisheye distortion

2.3.3 Radial Distortion

Typical cameras have a lens distortion, which disrupts the assumed linear projective model. Thus a camera may not be accurately represented by the pinhole camera model that we have described, particularly if a low-cost lens or a wide field-of- view (short focal length) lens such as a fisheye lens is employed. Some examples of lens distortion effects are shown in Fig. 2.5. Note that the effect is non-linear and, if significant, it must be corrected so that the camera can again be modeled as a linear device. The estimation of the required distortion parameters to do this is often encompassed within a camera calibration procedure, which is described in Sect. 2.4. With reference to our previous three-stage development of a projective camera in Sect. 2.3.2, lens distortion occurs at the second stage, which is the 3D to 2D projection, and this distortion is sampled by the image sensor.

Detailed distortion models contain a large number of parameters that model both radial and tangential distortion [7]. However, radial distortion is the dominant factor and usually it is considered sufficiently accurate to model this distortion only, using a low-order polynomial such as:

xnd

xn

xn

 

k1r

2

+

k2r4

,

ynd

=

yn

+

yn

 

 

 

 

where [xn, yn]T is the undistorted image position (i.e. that obeys our linear projection model) in normalized coordinates, [xnd , ynd ]T is the distorted image position in normalized coordinates, k1 and k2 are the unknown radial distortion parameters,

and r = xn2 + yn2. Assuming zero skew, we also have

xd

x

 

(x

− x0)

k1r2

+

k2r4

,

(2.6)

yd

= y

+

(y

−

y0)

 

 

 

 

 

 

 

 

 

 

 

 

 

 

whereT the distorted position [xd , yd ]T is now expressed in pixel coordinates and

[x, y]

are the usual pixel coordinates predicted by the linear pinhole model. Note

that r is still defined in normalized image coordinates and so a non-unity aspect ratio (mx = my ) in the image sensor does not invalidate this equation. Also note that both Eq. (2.6) and Fig. 2.5 indicate that distortion increases away from the center of the image. In the barrel distortion, shown in Fig. 2.5(c), distortion correction requires that image points are moved slightly towards the center of the image, more so if they are near the edges of the image. Correction could be applied to the whole image, as

48

S. Se and N. Pears

in dense stereo, or just a set of relevant features, such as extracted corner points. Clearly, the latter process is computationally cheaper.

Now that we have discussed the modeling of a camera’s image formation process in detail, we now need to understand how to estimate the parameters within this model. This is the focus of the next section, which details camera calibration.

2.4 Camera Calibration

Camera calibration [8] is the process of finding the parameters of the camera that produced a given image of a scene. This includes both extrinsic parameters R, t and intrinsic parameters, comprising those within the matrix K and radial distortion parameters, k1, k2. Once the intrinsic and extrinsic camera parameters are known, we know the camera projection matrix P and, taking into account of any radial distortion present, we can back-project any image pixel to a 3D ray in space. Clearly, as the intrinsic camera calibration parameters are tied to the focal length, changing the zoom on the lens would make the calibration invalid. It is also worth noting that calibration is not always required. For example, we may be more interested in approximate shape, where we need to know what objects in a scene are co-planar, rather than their absolute 3D position measurements. However, for stereo systems at least, camera calibration is commonplace.

Generally, it is not possible for an end-user to get the required calibration information to the required accuracy from camera manufacturer’s specifications and external measurement of the position of cameras in some frame. Hence some sort of camera calibration procedure is required, of which there are several different categories. The longest established of these is photogrammetric calibration, where calibration is performed using a scene object of precisely known physical dimensions. Typically, several images of a special 3D target, such as three orthogonal planes with calibration grids (chessboard patterns of black and white squares), are captured and precise known translations may be used [58]. Although this gives accurate calibration results, it lacks flexibility due to the need for precise scene knowledge.

At the other end of the spectrum is self-calibration (auto-calibration) [21, 35], where no calibration target is used. The correspondences across three images of the same rigid scene provide enough constraints to recover a set of camera parameters which allow 3D reconstruction up to a similarity transform. Although this approach is flexible, there are many parameters to estimate and reliable calibrations cannot always be obtained.

Between these two extremes are ‘desktop’ camera calibration approaches that use images of planar calibration grids, captured at several unknown positions and orientations (i.e. a single planar chessboard pattern is manually held at several random poses and calibration images are captured and stored). This gives a good compromise between the accuracy of photogrammetric calibration and the ease of use of self-calibration. A seminal example is given by Zhang [64].

Although there are a number of publicly available camera calibration packages on the web, such as the Caltech camera calibration toolbox for MATLAB [9] and in

Источник: https://studfile.net/preview/16498100/