UniTT-Stereo: Unified Training of Transformer for Enhanced Stereo Matching

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

1

초록

Unlike other vision tasks where Transformer-based approaches are becoming increasingly common, stereo depth estimation is still dominated by convolution-based models. This is mainly due to the limited availability of real-world ground truth for stereo matching, which hinders the performance improvement of transformer-based stereo approaches. In this paper, we propose UniTT-Stereo, a method to maximize the potential of Transformer-based stereo architectures by unifying self-supervised learning for pre-training with stereo matching framework based on supervised learning. Specifically, we design a dual-task learning scheme that reconstructs masked regions of an input image while simultaneously predicting corresponding points in the paired image. We demonstrate that this approach encourages the model to learn locality-aware representations, which are critical to overcoming the data inefficiency of Transformers. Moreover, to address these challenging tasks of reconstruction-and-prediction, we propose a variable masking ratio strategy that promotes robustness to varying levels of visual information. Additionally, we introduce losses that exploit stereo geometry and correspondence at the appearance, feature, and disparity levels. To further validate the effectiveness of our design, we conduct frequency decomposition and attention map visualization, which reveal how the model effectively captures fine-grained structures and cross-view correspondences. State-of-the-art performance of UniTT-Stereo is validated on various benchmarks such as the ETH3D, KITTI 2012, and KITTI 2015 datasets. Code is available at: https://github.com/00kim/UniTT-Stereo

키워드

TransformersImage reconstructionDepth measurementCostsTrainingFeature extractionSelf-supervised learningDecodingAccuracyTraining dataMasked image modelingstereo depth estimationsupervised learningself-supervised learningtransformer
제목
UniTT-Stereo: Unified Training of Transformer for Enhanced Stereo Matching
저자
Kim, SoominChoi, HyesongAhn, JihyeMin, Dongbo
DOI
10.1109/ACCESS.2025.3633291
발행일
2025-11
유형
Article
저널명
IEEE Access
13
페이지
204695 ~ 204707