Li-ViP3D++: 用于端到端感知与轨迹预测的查询门控可变形摄像-激光雷达融合
摘要
从原始传感器数据进行端到端感知与轨迹预测是自动驾驶的关键能力之一。模块化管线限制信息流动,可能放大上游误差。最近的基于查询的、全可微的感知-预测(PnP)模型缓解了这些问题,但摄像机和激光雷达在查询空间中的互补性尚未得到充分探索。模型常依赖引入启发式对齐和离散选择步骤的融合方案,这会限制可用信息的充分利用并引入不必要的偏差。我们提出了 Li-ViP3D++,一种基于查询的多模态 PnP 框架,引入查询门控可变形融合(QGDF)在查询空间集成多视角 RGB 和激光雷达。QGDF(1)通过带掩码注意力跨摄像机和特征层级聚合图像证据,(2)通过带学习型查询偏移的全可微 BEV 采样提取激光雷达上下文,(3)应用查询条件门控自适应加权每个代理人的视觉线索和几何线索。 resulting architecture jointly optimizes detection, tracking, and multi-hypothesis trajectory forecasting in a single end-to-end model. On nuScenes, Li-ViP3D++ improves end-to-end behavior and detection quality, achieving higher EPA (0.335) and mAP (0.502) while substantially reducing false positives (FP ratio 0.147), and it is faster than the prior Li-ViP3D variant (139.82 ms vs. 145.91 ms). These results indicate that query-space, fully differentiable camera-LiDAR fusion can increase robustness of end-to-end PnP without sacrificing deployability.
引用
@article{arxiv.2601.20720,
title = {Li-ViP3D++: Query-Gated Deformable Camera-LiDAR Fusion for End-to-End Perception and Trajectory Prediction},
author = {Matej Halinkovic and Nina Masarykova and Alexey Vinel and Marek Galinski},
journal= {arXiv preprint arXiv:2601.20720},
year = {2026}
}
备注
This work has been submitted to the IEEE for possible publication