3D single object tracking(SOT)based on point clouds is a fundamental task for environmental perception in autonomous driving and dynamic scene understanding in robotics.Recent technological advancements in this field ...3D single object tracking(SOT)based on point clouds is a fundamental task for environmental perception in autonomous driving and dynamic scene understanding in robotics.Recent technological advancements in this field have significantly bolstered the environmental interaction capabilities of intelligent systems.This field faces persistent challenges,including feature degradation induced by point cloud sparsity,representation drift caused by non-rigid deformation,and occlusion in complex scenarios.Traditional appearance matching methods,particularly those relying on Siamese networks,are severely constrained by point cloud characteristics,often failing under rapid motions or structural ambiguities among similar objects.In response,the research paradigm has progressively evolved toward motion-centric modeling approaches.These emerging frameworks utilize spatio-temporal joint modeling and geometric shape completion to attain notable performance gains.Furthermore,the incorporation of attention mechanisms and State Space Model(SSM)has enabled more effective multi-scale spatio-temporal feature association,which is particularly beneficial for long-term tracking scenarios.To the best of our knowledge,this is the first comprehensive survey dedicated to 3D single object tracking in point clouds.We provide a detailed analysis of current tracking methods,scrutinizing their limitations regarding multi-object interference and analyzing the trade-off between accuracy and computational efficiency.Finally,we discuss potential future directions,including the development of lightweight models for edge deployment and the integration of cross-modal fusion strategies.展开更多
Center point localization is a major factor affecting the performance of 3D single object tracking.Point clouds themselves are a set of discrete points on the local surface of an object,and there is also a lot of nois...Center point localization is a major factor affecting the performance of 3D single object tracking.Point clouds themselves are a set of discrete points on the local surface of an object,and there is also a lot of noise in the labeling.Therefore,directly regressing the center coordinates is not very reasonable.Existing methods usually use volumetric-based,point-based,and view-based methods,with a relatively single modality.In addition,the sampling strategies commonly used usually result in the loss of object information,and holistic and detailed information is beneficial for object localization.To address these challenges,we propose a novel Multi-view unsupervised center Uncertainty 3D single object Tracker(MUT).MUT models the potential uncertainty of center coordinates localization using an unsupervised manner,allowing the model to learn the true distribution.By projecting point clouds,MUT can obtain multi-view depth map features,realize efficient knowledge transfer from 2D to 3D,and provide another modality information for the tracker.We also propose a former attraction probability sampling strategy that preserves object information.By using both holistic and detailed descriptors of point clouds,the tracker can have a more comprehensive understanding of the tracking environment.Experimental results show that the proposed MUT network outperforms the baseline models on the KITTI dataset by 0.8%and 0.6%in precision and success rate,respectively,and on the NuScenes dataset by 1.4%,and 6.1%in precision and success rate,respectively.The code is made available at http://gffzz188fe103f8f1460aswqvub9obuqux690w.ffgz.tsg.suse.edu.cn/abchears/MUT.git.展开更多
基金supported by the National Natural Science Foundation of China(Nos.62306049,92471207 and W2421089)the General Program of Chongqing Natural Science Foundation(No.CSTB2023NSCQMSX0665)the Fundamental Research Funds for the Central Universities(No.2024CDJXY008).
摘要3D single object tracking(SOT)based on point clouds is a fundamental task for environmental perception in autonomous driving and dynamic scene understanding in robotics.Recent technological advancements in this field have significantly bolstered the environmental interaction capabilities of intelligent systems.This field faces persistent challenges,including feature degradation induced by point cloud sparsity,representation drift caused by non-rigid deformation,and occlusion in complex scenarios.Traditional appearance matching methods,particularly those relying on Siamese networks,are severely constrained by point cloud characteristics,often failing under rapid motions or structural ambiguities among similar objects.In response,the research paradigm has progressively evolved toward motion-centric modeling approaches.These emerging frameworks utilize spatio-temporal joint modeling and geometric shape completion to attain notable performance gains.Furthermore,the incorporation of attention mechanisms and State Space Model(SSM)has enabled more effective multi-scale spatio-temporal feature association,which is particularly beneficial for long-term tracking scenarios.To the best of our knowledge,this is the first comprehensive survey dedicated to 3D single object tracking in point clouds.We provide a detailed analysis of current tracking methods,scrutinizing their limitations regarding multi-object interference and analyzing the trade-off between accuracy and computational efficiency.Finally,we discuss potential future directions,including the development of lightweight models for edge deployment and the integration of cross-modal fusion strategies.
摘要Center point localization is a major factor affecting the performance of 3D single object tracking.Point clouds themselves are a set of discrete points on the local surface of an object,and there is also a lot of noise in the labeling.Therefore,directly regressing the center coordinates is not very reasonable.Existing methods usually use volumetric-based,point-based,and view-based methods,with a relatively single modality.In addition,the sampling strategies commonly used usually result in the loss of object information,and holistic and detailed information is beneficial for object localization.To address these challenges,we propose a novel Multi-view unsupervised center Uncertainty 3D single object Tracker(MUT).MUT models the potential uncertainty of center coordinates localization using an unsupervised manner,allowing the model to learn the true distribution.By projecting point clouds,MUT can obtain multi-view depth map features,realize efficient knowledge transfer from 2D to 3D,and provide another modality information for the tracker.We also propose a former attraction probability sampling strategy that preserves object information.By using both holistic and detailed descriptors of point clouds,the tracker can have a more comprehensive understanding of the tracking environment.Experimental results show that the proposed MUT network outperforms the baseline models on the KITTI dataset by 0.8%and 0.6%in precision and success rate,respectively,and on the NuScenes dataset by 1.4%,and 6.1%in precision and success rate,respectively.The code is made available at http://gffzz188fe103f8f1460aswqvub9obuqux690w.ffgz.tsg.suse.edu.cn/abchears/MUT.git.