XU Ling-zhi,FU Qin-wei,TAO Wei,et al.Monocular vehicle pose estimation based on 3D model[J].Optics and Precision Engineering,2021,29(06):1346-1355.DOI: 10.37188/OPE.20212906.1346.
Monocular vehicle pose estimation based on 3D model
Vehicle pose estimation is an important component of intelligent transportation systems. However, the complex scenes and loss of depth information are challenging problems in the estimation. This paper proposes a method that combines monocular pose estimation and a 3D vehicle model to estimate vehicle pose. First, a multi-scale vehicle are normalized, and then the coordinates of key points are predicted in the form of a vector field to increase the accuracy of the pose estimation for the truncated and occluded vehicle. Furthermore, a distance-based loss function for the vector field and key point error minimization voting method is established to further improve the accuracy of the pose estimation algorithm. In addition, we propose a synthetic vehicle pose estimation dataset with rich annotation information. The verification results show that the average position and angle errors of our algorithm are 0.162 m and 4.692°, respectively. Our method provides significant improvements over existing methods and has considerable practical application value.
关键词
Keywords
references
ZHOU Y , SUN P , ZHANG Y , et al . End-to-end multi-view fusion for 3D object detection in LiDAR point clouds [EB/OL]. 2019: arXiv : 1910 .
BELTRÁN J , GUINDEL C , MORENO F M , et al . BirdNet: a 3D object detection framework from LiDAR information [C]. 2018 21st International Conference on Intelligent Transportation Systems (ITSC) . 47,2018 , Maui , HI, USA . IEEE , 2018 : 3517 - 3523 .
ALI W , ABDELKARIM S , ZAHRAN M , et al . YOLO3D: end-to-end real-time 3D oriented object bounding box detection from LiDAR point cloud [EB/OL]. 2018: arXiv : 1808 .
WU D , ZHUANG Z , XIANG C , et al . 6D-VNet: End-to-end 6DoF Vehicle Pose Estimation from Monocular RGB Images [C]. Proceedings of the Computer Vision and Pattern Recognition Workshops, Long Beach , United States: CVPR , 2019 : 0 - 0 .
KEHL W , MANHARDT F , TOMBARI F , et al . SSD-6D: making RGB-based 3D detection and 6D pose estimation great again [C]. 2017 IEEE International Conference on Computer Vision (ICCV) . 2229,2017 , Venice, Italy . IEEE , 2017 : 1530 - 1538 .
CHEN Y J , TAI L , SUN K , et al . MonoPair: monocular 3D object detection using pairwise spatial relationships [C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 1319,2020 , Seattle , WA , USA . IEEE , 2020 : 12090 - 12099 .
HINTERSTOISSER S , HOLZER S , CAGNIART C , et al . Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes [C]. 2011 International Conference on Computer Vision . 613,2011 , Barcelona, Spain . IEEE , 2011 : 858 - 865 .
LU R Q , MA H M . Template matching with multi-scale saliency [J]. Opt. Precision Eng. , 2018 , 26 ( 11 ): 2776 - 2784 . (in Chinese)
LI Q , HU R , CHEN Y , et al . Vehicle Pose Estimation Using Mask Matching [C]. 2019 IEEE International Conference on Acoustics, Speech and Signal Processing , Brighton , UK: ICASSP , 2019 : 1972 - 1976 .
BAROWSKI T , SZCZOT M , HOUBEN S . 6DoF vehicle pose estimation using segmentation-based part correspondences [C]. 2019 IEEE Intelligent Transportation Systems Conference (ITSC) . 2730,2019 , Auckland , New Zealand. IEEE , 2019 : 573 - 580 .
CHABOT F , CHAOUCH M , RABARISOA J , et al . Deep MANTA: a coarse-to-fine many-task network for joint 2D and 3D vehicle analysis from monocular image [C]. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2126,2017 , Honolulu, HI, USA . IEEE , 2017 : 1827 - 1836 .
TEKIN B , SINHA S N , FUA P . Real-time seamless single shot 6D object pose prediction [C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1823,2018 , Salt Lake City, UT, USA . IEEE , 2018 : 292 - 301 .
QU Y P , HOU W . Attitude accuracy analysis of PnP based on error propagation theory [J]. Opt. Precision Eng. , 2019 , 27 ( 2 ): 479 - 487 . (in Chinese)
FAN L L , ZHAO H W , ZHAO H Y , et al . Survey of target detection based on deep convolutional neural networks [J]. Opt. Precision Eng. , 2020 , 28 ( 5 ): 1152 - 1164 . (in Chinese)
CAO Z , SIMON T , WEI S H , et al . Realtime multi-person 2D pose estimation using part affinity fields [C]. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2126,2017 , Honolulu, HI, USA . IEEE , 2017 : 1302 - 1310 .
XIANG Y , SCHMIDT T , NARAYANAN V , et al . PoseCNN: a convolutional neural network for 6D object pose estimation in cluttered scenes [C]. Robotics : Science and Systems XIV. Robotics : Science and Systems Foundation , 2018 : 176 - 185 .
PENG S D , LIU Y , HUANG Q X , et al . PVNet: pixel-wise voting network for 6DoF pose estimation [C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 1520,2019 , Long Beach, CA, USA . IEEE , 2019 : 4556 - 4565 .
RONNEBERGER O , FISCHER P , BROX T . U-net: convolutional networks for biomedical image segmentation [C]. Medical Image Computing and Computer-Assisted Intervention-MICCAI , 2015 , 2015 : 234 - 241 .
LEPETIT V , MORENO-NOGUER F , FUA P . EPnP: an accurate O(n) solution to the PnP problem [J]. International Journal of Computer Vision , 2008 , 81 ( 2 ): 155 - 166 .
GEIGER A , LENZ P , STILLER C , et al . Vision meets robotics: The KITTI dataset [J]. The International Journal of Robotics Research , 2013 , 32 ( 11 ): 1231 - 1237 .
GÄHLERT N , JOURDAN N , CORDTS M , et al . Cityscapes 3 D: dataset and benchmark for 9 DoF vehicle detection[EB/OL]. arXiv preprint arXiv , 2020 : 2006 . 07864 .
CHANG A X , FUNKHOUSER T , GUIBAS L , et al . ShapeNet: an information-rich 3D model repository [EB/OL]. 2015: arXiv : 1512 .
ZHOU B L , LAPEDRIZA A , KHOSLA A , et al . Places: a 10 million image database for scene recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence , 2018 , 40 ( 6 ): 1452 - 1464 .
HE K M , ZHANG X Y , REN S Q , et al . Deep residual learning for image recognition [C]. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2730,2016 , Las Vegas, NV, USA . IEEE , 2016 : 770 - 778 .
XIANG Y , CHOI W , LIN Y Q , et al . Data-driven 3D Voxel Patterns for object category recognition [C]. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 712,2015 , Boston, MA, USA . IEEE , 2015 : 1903 - 1911 .
Dual-modal precision pose measurement for thick-walled cylindrical holes under monocular backlight vision based on grayscale integral projection
Monocular vision-based spatial registration for implant surgery robots
Automatic alignment of container corner based on monocular vision
Hidden point coordinate measurement based on inertial-visual fused attitude estimation
Research on measurement technology of rocket recovery height based on monocular vision
Related Author
CHEN Suyu
CHENG Lei
LI Wu
LIAO Xiaobo
ZHANG Xianmin
YAO Yanbing
DU Yimin
QIAN Jiake
Related Institution
College of Engineering Technology, Southwest University
Aircraft Fluid Physics Key Laboratory , China Aerodynamics Research and Development Center
Key Laboratory of Manufacturing Process Testing Technology of the Ministry of Education, Southwest University of Science and Technology
Guangdong Province Key Laboratory of Precision Equipment and Manufacturing Technology, School of Mechanical and Automotive Engineering, South China University of Technology
School of Instrumentation Science and Engineering, Harbin Institute of Technology