Learning matching costs has been shown to be critical to the success of the\nstate-of-the-art deep stereo matching methods, in which 3D convolutions are\napplied on a 4D feature volume to learn a 3D cost volume. However, this\nmechanism has never been employed for the optical flow task. This is mainly due\nto the significantly increased search dimension in the case of optical flow\ncomputation, ie, a straightforward extension would require dense 4D\nconvolutions in order to process a 5D feature volume, which is computationally\nprohibitive. This paper proposes a novel solution that is able to bypass the\nrequirement of building a 5D feature volume while still allowing the network to\nlearn suitable matching costs from data. Our key innovation is to decouple the\nconnection between 2D displacements and learn the matching costs at each 2D\ndisplacement hypothesis independently, ie, displacement-invariant cost\nlearning. Specifically, we apply the same 2D convolution-based matching net\nindependently on each 2D displacement hypothesis to learn a 4D cost volume.\nMoreover, we propose a displacement-aware projection layer to scale the learned\ncost volume, which reconsiders the correlation between different displacement\ncandidates and mitigates the multi-modal problem in the learned cost volume.\nThe cost volume is then projected to optical flow estimation through a 2D\nsoft-argmin layer. Extensive experiments show that our approach achieves\nstate-of-the-art accuracy on various datasets, and outperforms all published\noptical flow methods on the Sintel benchmark.\n
Paper
References (49)
Scroll for more · 37 remaining