中文
相关论文

相关论文: MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregat…

200 篇论文

There have been significant advances in neural networks for both 3D object detection using LiDAR and 2D object detection using video. However, it has been surprisingly difficult to train networks to effectively use both modalities in a way…

计算机视觉与模式识别 · 计算机科学 2020-09-03 Su Pang , Daniel Morris , Hayder Radha

The integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task…

机器人学 · 计算机科学 2024-03-25 Shaowei Fu , Yifan Duan , Yao Li , Chengzhen Meng , Yingjie Wang , Jianmin Ji , Yanyong Zhang

Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view-invariant paradigm, aligning shared features across…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Quan Zhang , Zeqiang Cai , Peiming Zhao , Jingze Wu , Cailun Wu , Hongbo Chen , Jianhuang Lai

Cross-view image matching for geo-localisation is a challenging problem due to the significant visual difference between aerial and ground-level viewpoints. The method provides localisation capabilities from geo-referenced images,…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Tavis Shore , Simon Hadfield , Oscar Mendez

Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave infrared (LWIR) sensing preserves contrast in low light, and event cameras retain…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Craig Iaboni , Pramod Abichandani

The field of autonomous vehicles (AVs) predominantly leverages multi-modal integration of LiDAR and camera data to achieve better performance compared to using a single modality. However, the fusion process encounters challenges in…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Sanjay Bhargav Dharavath , Tanmoy Dam , Supriyo Chakraborty , Prithwiraj Roy , Aniruddha Maiti

This paper introduces VLMFusionOcc3D, a robust multimodal framework for dense 3D semantic occupancy prediction in autonomous driving. Current voxel-based occupancy models often struggle with semantic ambiguity in sparse geometric grids and…

计算机视觉与模式识别 · 计算机科学 2026-03-04 A. Enes Doruk , Hasan F. Ates

We propose DeepFusion, a modular multi-modal architecture to fuse lidars, cameras and radars in different combinations for 3D object detection. Specialized feature extractors take advantage of each modality and can be exchanged easily,…

计算机视觉与模式识别 · 计算机科学 2022-09-28 Florian Drews , Di Feng , Florian Faion , Lars Rosenbaum , Michael Ulrich , Claudius Gläser

Modern Unmanned Aerial Vehicles equipped with state of the art artificial intelligence (AI) technologies are opening to a wide plethora of novel and interesting applications. While this field received a strong impact from the recent AI…

计算机视觉与模式识别 · 计算机科学 2020-04-15 Enkhtogtokh Togootogtokh , Christian Micheloni , Gian Luca Foresti , Niki Martinel

Multi-view detection incorporates multiple camera views to alleviate occlusion in crowded scenes, where the state-of-the-art approaches adopt homography transformations to project multi-view features to the ground plane. However, we find…

计算机视觉与模式识别 · 计算机科学 2023-01-05 Jiahao Ma , Jinguang Tong , Shan Wang , Wei Zhao , Zicheng Duan , Chuong Nguyen

Many recent works on 3D object detection have focused on designing neural network architectures that can consume point cloud data. While these approaches demonstrate encouraging performance, they are typically based on a single modality and…

计算机视觉与模式识别 · 计算机科学 2019-04-04 Vishwanath A. Sindagi , Yin Zhou , Oncel Tuzel

Cross-modal data registration has long been a critical task in computer vision, with extensive applications in autonomous driving and robotics. Accurate and robust registration methods are essential for aligning data from different…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Yuanchao Yue , Hui Yuan , Qinglong Miao , Xiaolong Mao , Raouf Hamzaoui , Peter Eisert

Accurate 3D lane estimation is crucial for ensuring safety in autonomous driving. However, prevailing monocular techniques suffer from depth loss and lighting variations, hampering accurate 3D lane detection. In contrast, LiDAR points offer…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Yueru Luo , Shuguang Cui , Zhen Li

Visual Grounding (VG) aims to locate the most relevant region in an image, based on a flexible natural language query but not a pre-defined label, thus it can be a more useful technique than object detection in practice. Most…

计算机视觉与模式识别 · 计算机科学 2019-03-19 Chaorui Deng , Qi Wu , Guanghui Xu , Zhuliang Yu , Yanwu Xu , Kui Jia , Mingkui Tan

Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Unmanned Aerial Vehicle (UAV) queries by matching them against an extensive geo-tagged…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Bowen Liu , Pengyue Jia , Wanyu Wang , Derong Xu , Jiawei Cheng , Jiancheng Dong , Xiao Han , Zimo Zhao , Chao Zhang , Bowen Yu , Fangyu Hong , Xiangyu Zhao

Visual Place Recognition (VPR) has been traditionally formulated as a single-image retrieval task. Using multiple views offers clear advantages, yet this setting remains relatively underexplored and existing methods often struggle to…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Tianchen Deng , Xun Chen , Ziming Li , Hongming Shen , Danwei Wang , Javier Civera , Hesheng Wang

Drone-view geo-localization (DVGL) aims to match images of the same geographic location captured from drone and satellite perspectives. Despite recent advances, DVGL remains challenging due to significant appearance changes and spatial…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Ke Li , Di Wang , Xiaowei Wang , Zhihong Wu , Yiming Zhang , Yifeng Wang , Quan Wang

Video retrieval requires aligning visual content with corresponding natural language descriptions. In this paper, we introduce Modality Auxiliary Concepts for Video Retrieval (MAC-VR), a novel approach that leverages modality-specific tags…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Adriano Fragomeni , Dima Damen , Michael Wray

Masked Image Modeling (MIM) with Vector Quantization (VQ) has achieved great success in both self-supervised pre-training and image generation. However, most existing methods struggle to address the trade-off in shared latent space for…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Siyuan Li , Luyuan Zhang , Zedong Wang , Juanxi Tian , Cheng Tan , Zicheng Liu , Chang Yu , Qingsong Xie , Haonan Lu , Haoqian Wang , Zhen Lei

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robustness to noisy…