2 Passive 3D Imaging |
89 |
One of the key challenges for 3D vision researchers is to develop algorithms to recover accurate 3D information robustly under a wide range of illumination conditions which can be done by humans so effortlessly. While 3D passive vision algorithms have been maturing over the years, this is still an active topic in the research community and at major computer vision conferences. Many algorithms perform reasonably well with test data but there are still challenges to handle scenes with uncontrolled illumination. Other open issues include efficient global dense stereo matching, multi-image matching and fully automated accurate 3D reconstruction from images.
Passive 3D imaging systems are becoming more prevalent as cameras are getting cheaper and computers are fast enough to handle the intensive processing requirements. Thanks to hardware acceleration and GPUs, real-time applications are more common, leading to a growing number of real-world applications.
After working through this chapter, you should be able to:
•Explain the fundamental concepts and challenges of passive 3D imaging systems.
•Explain the principles of epipolar geometry.
•Solve the correspondence problem by correlation-based and feature-based techniques (using off-the-shelf feature extractors).
•Estimate the fundamental matrix from correspondences.
•Perform dense stereo matching and compute a 3D point cloud.
•Explain the principles of structure from motion.
•Provide example applications of passive 3D imaging systems.
Two-view geometry is studied extensively in [21], which also covers the equivalent of epipolar geometry for three or more images. The eight-point algorithm was proposed in [19] to compute the fundamental matrix, while the five-point algorithm was proposed in [39] for calibrated cameras. Reference [57] provides a good tutorial and survey on bundle adjustment, which is also covered in textbooks [15, 21] and a recent survey article [35].
Surveys such as [46] serve as a guide to the extensive literature on stereo imaging. Structure from motion is extensively covered in review articles such as [35]. A step- by-step guide to 3D modeling from images is described in detail in [30]. Non-rigid structure from motion for dynamic scenes is discussed in [56].
Multiple-view 3D vision continues to be a highly active research topic and some of the major computer vision conferences include: the International Conference on Computer Vision (ICCV), IEEE Conference on Computer Vision and Pattern Recognition (CVPR) and the European Conference on Computer Vision (ECCV). Some of the relevant major journals include: International Journal of Computer
90 |
S. Se and N. Pears |
Vision (IJCV), IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI) and Image and Vision Computing (IVC).
The International Society for Photogrammetry and Remote Sensing (ISPRS) proceedings and archives provide extensive literature on photogrammetry and related topics.
The following web sites provide comprehensive on-line resources for computer vision including 3D passive vision topics and are being updated regularly.
•CVonline (http://homepages.inf.ed.ac.uk/rbf/CVonline/) provides an on-line compendium of computer vision.
•VisionBib.Com (http://www.visionbib.com) contains annotated bibliography on a wide range of computer vision topics, as well as references to available datasets.
•Computer Vision online (http://www.computervisiononline.com) is a portal with links to software, hardware and datasets.
•OpenCV (http://opencv.willowgarage.com) is an open-source computer vision library.
1.What are the differences between passive and active 3D vision systems?
2.Name two approaches to recover 3D from single images and two approaches to recover 3D from multiple images.
3.What is the epipolar constraint and how can you use it to speed up the search for correspondences?
4.What are the differences between essential and fundamental matrices?
5.What is the purpose of rectification?
6.What are the differences between correlation-based and feature-based methods for finding correspondences?
7.What are the differences between local and global methods for dense stereo matching?
8.What are the differences between stereo and structure from motion?
9.What are the factors that affect the accuracy of stereo vision systems?
Experimenting with stereo imaging requires that you have two images of a scene from slightly different viewpoints, with a good overlap between the views, and a significant number of well distributed corner features that can be matched. You will also need a corner detector. There are many stereo image pairs and corner detector implementations available on the web [40]. Of course, you can collect your own images either with a pre-packaged stereo camera or with a pair of standard digital cameras. The following programming exercises should be implemented in a language of your choice.
2 Passive 3D Imaging |
91 |
1.Fundamental matrix with manual correspondences. Run a corner detector on the image pair. Use a point-and-click GUI to manually label around 20 well distributed correspondences. Compute the fundamental matrix and plot the conjugate pair of epipolar lines on the images for each correspondence. Experiment with different numbers and combinations of correspondences, using a minimum of eight in the eight-point algorithm. Observe and comment on the sensitivity of the epipolar lines with respect to the set of correspondences chosen.
2.Fundamental matrix estimation with outlier removal. Add 4 incorrect corner correspondences to your list of 20 correct ones. Observe the effect on the computed fundamental matrix and the associated (corrupted) epipolar lines. Augment your implementation of fundamental matrix estimation with the RANSAC algorithm. Use a graphical overlay on your images to show that RANSAC has correctly identified the outliers, and verify that the fundamental matrix and its associated epipolar lines can now be computed without the corrupting effect of the outliers.
3.Automatic feature correspondences. Implement a function to automatically match corners between two images according to the Sum of Squared Differences (SSD) measure. Also, implement a function for the Normalized CrossCorrelation (NCC) measure. Compare the matching results with test images of similar brightness and also of different brightness.
4.Fundamental matrix from automatic correspondences. Use your fundamental matrix computation (with RANSAC) with the automatic feature correspondences. Determine the positions of the epipoles and, again, plot the epipolar lines.
The following additional exercises require the use of a stereo rig, which could be a pre-packaged stereo pair or a home-made rig with a pair of standard digital cameras. The cameras should have a small amount of vergence to overlap their fields of view.
5.Calibration. Create your own calibration target by printing off a chessboard pattern and pasting it to a flat piece of wood. Use a point-and-click GUI to semiautomate the corner correspondences between the calibration target and a set of captured calibration images. Implement a camera calibration procedure for a stereo pair to determine the intrinsic and extrinsic parameters of the stereo rig. If you have less time available you may choose to use some of the calibration libraries available on the web [9, 40].
6.Rectification. Compute an image warping (homography) to apply to each image in the stereo image pair, such that conjugate epipolar lines are horizontal (parallel to the x-axis) and have the same y-coordinate. Plot a set of epipolar lines to check that this rectification is correct.
7.Dense stereo matching. Implement a function to perform local dense stereo matching between left and right rectified images, using NCC as the similarity measure, and hence generate a disparity map for the stereo pair. Capture stereo images for a selection of scenes with varying amounts of texture within them and at varying distances from the cameras, and compare their disparity maps.
8.3D reconstruction. Implement a function to perform a 3D reconstruction from your disparity maps and camera calibration information. Use a graphics tool to visualize the reconstructions. Comment on the performance of the reconstructions for different scenes and for different distances from the stereo rig.
92 |
S. Se and N. Pears |
1.Barfoot, T., Se, S., Jasiobedzki, P.: Vision-based localization and terrain modelling for planetary rovers. In: Howard, A., Tunstel, E. (eds.) Intelligence for Space Robotics, pp. 71–92. TSI Press, Albuquerque (2006)
2.Bay, H., Ess, A., Tuytelaars, T., Van Gool, L.: SURF: speeded up robust features. Comput. Vis. Image Underst. 110(3), 346–359 (2008)
3.Belhumeur, P.N.: A Bayesian approach to binocular stereopsis. Int. J. Comput. Vis. 19(3), 237–260 (1996)
4.Bergen, J.R., Anandan, P., Hanna, K.J., Hinogorani, R.: Hierarchical model-based motion estimation. In: European Conference on Computer Vision (ECCV), Italy, pp. 237–252 (1992)
5.Birchfield, S., Tomasi, C.: Depth discontinuities by pixel-to-pixel stereo. Int. J. Comput. Vis. 35(3), 269–293 (1999)
6.Boykov, Y., Veksler, O., Zabih, R.: Fast approximate energy minimization via graph cuts. IEEE Trans. Pattern Anal. Mach. Intell. 23(11), 1222–1239 (2001)
7.Brown, D.C.: Decentering distortion of lenses. Photogramm. Eng. 32(3), 444–462 (1966)
8.Brown, D.C.: Close-range camera calibration. Photogramm. Eng. 37(8), 855–866 (1971)
9.Camera Calibration Toolbox for MATLAB: http://www.vision.caltech.edu/bouguetj/ calib_doc/. Accessed 20th October 2011
10.Chen, Z., Pears, N.E., Liang, B.: Monocular obstacle detection using Reciprocal-Polar rectification. Image Vis. Comput. 24(12), 1301–1312 (2006)
11.Cox, I.J., Hingorani, S.L., Rao, S.B., Maggs, B.M.: A maximum likelihood stereo algorithm. Comput. Vis. Image Underst. 63(3), 542–567 (1996)
12.Coxeter, H.S.M.: Projective Geometry, 2nd edn. Springer, Berlin (2003)
13.Criminisi, A., Reid, I., Zisserman, A.: Single view metrology. Int. J. Comput. Vis. 40(2), 123– 148 (2000)
14.Davison, A., Reid, I., Molton, N., Stasse, O.: MonoSLAM: real-time single camera SLAM. IEEE Trans. Pattern Anal. Mach. Intell. 29(6), 1052–1067 (2007)
15.Faugeras, O., Luong, Q.T.: The Geometry of Multiple Images. MIT Press, Cambridge (2001)
16.Fischler, M.A., Bolles, R.C.: Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 24(6), 381–395 (1981)
17.Garding, J.: Shape from texture for smooth curved surfaces in perspective projection. J. Math. Imaging Vis. 2, 329–352 (1992)
18.Harris, C., Stephens, M.J.: A combined corner and edge detector. In: Alvey Vision Conference, pp. 147–152 (1988)
19.Harley, R.I.: In defense of the 8-point algorithm. IEEE Trans. Pattern Anal. Mach. Intell. 19(6), 580–593 (1997)
20.Hartley, R.I.: Theory and practice of projective rectification. Int. J. Comput. Vis. 35(2), 115– 127 (1999)
21.Hartley, R.I., Zisserman, A.: Multiple View Geometry in Computer Vision, 2nd edn. Cambridge University Press, Cambridge (2004)
22.Horn, B.K.P., Brooks, M.J. (eds.): Shape from Shading. MIT Press, Cambridge (1989)
23.Horn, B.K.P., Schunck, B.G.: Determining optical flow. Artif. Intell. 17, 185–203 (1981)
24.Huang, R., P, S.W.A.: A shape-from-shading framework for satisfying data-closeness and structure-preserving smoothness constraints. In: Proceedings of the British Machine Vision Conference (2009)
25.Kolmogorov, V., Zabih, R.: Computing visual correspondence with occlusions using graph cuts. In: International Conference on Computer Vision (ICCV), Vancouver, pp. 508–515 (2001)
26.Konolige, K.: Small vision system: hardware and implementation. In: Proc. Int. Symp. on Robotics Research, Hayama, Japan, pp. 111–116 (1997)
27.Longuet-Higgins, H.C.: A computer algorithm for re-constructing a scene from two projections. Nature 293, 133–135 (1981)
2 Passive 3D Imaging |
93 |
28.Lowe, D.G.: Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. 60(2), 91–110 (2004)
29.Lucas, B.D., Kanade, T.: An interactive image registration technique with an application in stereo vision. In: International Joint Conference on Artificial Intelligence (IJCAI), Vancouver,
pp.674–679 (1981)
30.Ma, Y., Soatto, S., Kosecka, J., Sastry, S.S.: An Invitation to 3-D Vision. Springer, New York (2003)
31.Maimone, M., Biesiadecki, J., Tunstel, E., Cheng, Y., Leger, C.: Surface navigation and mobility intelligence on the Mars Exploration Rovers. In: Howard, A., Tunstel, E. (eds.) Intelligence for Space Robotics, pp. 45–69. TSI Press, Albuquerque (2006)
32.Maimone, M., Cheng, Y., Matthies, L.: Two years of visual odometry on the mars exploration rovers. J. Field Robot. 24(3), 169–186 (2007)
33.Mallon, J., Whelan, P.F.: Projective rectification from the fundamental matrix. In: Image and Vision Computing, pp. 643–650 (2005)
34.The Middlebury stereo vision page: http://vision.middlebury.edu/stereo/. Accessed 16th November 2011
35.Moons, T., Van Gool, L., Vergauwen, M.: 3D reconstruction from multiple images. Found. Trends Comput. Graph. Vis. 4(4), 287–404 (2010)
36.Mordohai, P., et al.: Real-time video-based reconstruction of urban environments. In: International Workshop on 3D Virtual Reconstruction and Visualization of Complex Architectures (3D-ARCH), Zurich, Switzerland (2007)
37.Nayar, S.K., Nakagawa, Y.: Shape from focus. IEEE Trans. Pattern Anal. Mach. Intell. 16(8), 824–831 (1994)
38.Nister, D.: Automatic passive recovery of 3D from images and video. In: International Symposium on 3D Data Processing, Visualization and Transmission (3DPVT), Thessaloniki, Greece,
pp.438–445 (2004)
39.Nister, D.: An efficient solution to the five-point relative pose problem. IEEE Trans. Pattern Anal. Mach. Intell. 26(6), 756–770 (2004)
40.Open source computer vision library: http://opencv.willowgarage.com/wiki/. Accessed 20th October 2011
41.Pentland, A.P.: A new sense for depth of field. IEEE Trans. Pattern Anal. Mach. Intell. 9(4), 523–531 (1987)
42.Pollefeys, M., Koch, R., Van Gool, L.: A simple and efficient rectification method for general motion. In: International Conference on Compute Vision (ICCV), Kerkyra, Greece, pp. 496– 501 (1999)
43.Pollefeys, M., Van Gool, L., Vergauwen, M., Verbiest, F., Cornelis, K., Tops, J., Koch, R.: Visual modeling with a hand-held camera. Int. J. Comput. Vis. 59(3), 207–232 (2004)
44.Quam, L.H.: Hierarchical warp stereo. In: Image Understanding Workshop, New Orleans,
pp.149–155 (1984)
45.Roy, S., Cox, I.J.: A maximum-flow formulation of the N -camera stereo correspondence problem. In: International Conference on Computer Vision (ICCV), Bombay, pp. 492–499 (1998)
46.Scharstein, D., Szeliski, R.: A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. Int. J. Comput. Vis. 47(1/2/3), 7–42 (2002)
47.Se, S., Jasiobedzki, P.: Photo-realistic 3D model reconstruction. In: IEEE International Conference on Robotics and Automation, Orlando, Florida, pp. 3076–3082 (2006)
48.Se, S., Jasiobedzki, P.: Stereo-vision based 3D modeling and localization for unmanned vehicles. Int. J. Intell. Control Syst. 13(1), 47–58 (2008)
49.Se, S., Lowe, D., Little, J.: Mobile robot localization and mapping with uncertainty using scale-invariant visual landmarks. Int. J. Robot. Res. 21(8), 735–758 (2002)
50.Se, S., Firoozfam, P., Goldstein, N., Dutkiewicz, M., Pace, P.: Automated UAV-based video exploitation for mapping and surveillance. In: International Society for Photogrammetry and Remote Sensing (ISPRS) Commission I Symposium, Calgary (2010)
51.Seitz, S., Curless, B., Diebel, J., Scharstein, D., Szeliski, R.: A comparison and evaluation of multi-view stereo reconstruction algorithms. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), New York, pp. 519–526 (2006)