Related Experiment Video
Updated: Apr 15, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.7K
PM-PM: PatchMatch with Potts Model for object segmentation and stereo matching
Summary
This study introduces a novel method for joint object segmentation and stereo matching, achieving accurate depth-map generation by representing objects with depth planes. The approach enhances multi-view reconstruction efficiency and accuracy, outperforming existing methods.
Area of Science:
- Computer Vision
- 3D Reconstruction
- Image Processing
Background:
- Traditional stereo matching methods often struggle with accuracy and efficiency.
- Object-level depth representation is crucial for detailed 3D scene understanding.
- Existing algorithms may introduce artifacts like discretization or staircasing.
Purpose of the Study:
- To develop a unified variational formulation for joint object segmentation and stereo matching.
- To improve the accuracy and efficiency of depth-map generation.
- To enable robust multi-view reconstruction without common artifacts.
Main Methods:
- A novel object-level depth representation using image space perimeter, depth plane, and planar bias.
- Convex formulation of the multilabel Potts Model integrated with PatchMatch stereo techniques.
- Energy minimization optimized via a fast primal-dual algorithm for efficient computation.
Main Results:
- The proposed method achieves subpixel accurate disparity estimation, outperforming traditional and PatchMatch variants.
- Demonstrated accurate multi-view reconstruction using induced homography without discretization artifacts.
- Efficiently handles several hundred object depth segments.
Conclusions:
- The unified formulation offers a significant advancement in joint object segmentation and stereo matching.
- The object-level depth representation leads to superior performance on benchmark datasets (Middlebury, KITTI).
- The method provides a robust and efficient solution for real-world multi-view reconstruction challenges.

