中文
相关论文

相关论文: MoniRefer: A Real-world Large-scale Multi-modal Da…

200 篇论文

3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing methods rely on a pre-defined Object Lookup Table (OLT) to query Visual Language Models (VLMs) for reasoning about object locations,…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Wenyuan Huang , Zhao Wang , Zhou Wei , Ting Huang , Fang Zhao , Jian Yang , Zhenyu Zhang

3D object detection using LiDAR data is an indispensable component for autonomous driving systems. Yet, only a few LiDAR-based 3D object detection methods leverage segmentation information to further guide the detection process. In this…

计算机视觉与模式识别 · 计算机科学 2022-03-07 Hamidreza Fazlali , Yixuan Xu , Yuan Ren , Bingbing Liu

The development of computer vision algorithms for Unmanned Aerial Vehicles (UAVs) imagery heavily relies on the availability of annotated high-resolution aerial data. However, the scarcity of large-scale real datasets with pixel-level…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Giulia Rizzoli , Francesco Barbato , Matteo Caligiuri , Pietro Zanuttigh

Accurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understanding that textual information can provide. To address this…

计算工程、金融与科学 · 计算机科学 2025-12-11 Xi Xiao , Yunbei Zhang , Janet Wang , Lin Zhao , Yuxiang Wei , Hengjia Li , Yanshu Li , Xinyuan Song , Xiao Wang , Swalpa Kumar Roy , Hao Xu , Tianyang Wang

In this technical study, we introduce VFusedSeg3D, an innovative multi-modal fusion system created by the VisionRD team that combines camera and LiDAR data to significantly enhance the accuracy of 3D perception. VFusedSeg3D uses the rich…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Osama Amjad , Ammad Nadeem

Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding performance by training large models with large-scale…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Yangxiao Lu , Ruosen Li , Liqiang Jing , Jikai Wang , Xinya Du , Yunhui Guo , Nicholas Ruozzi , Yu Xiang

Multi-view imaging systems enable uniform coverage of 3D space and reduce the impact of occlusion, which is beneficial for 3D object detection and tracking accuracy. However, existing imaging systems built with multi-view cameras or depth…

计算机视觉与模式识别 · 计算机科学 2023-02-22 Meng Zhang , Wenxuan Guo , Bohao Fan , Yifan Chen , Jianjiang Feng , Jie Zhou

Vision-and-Language Navigation (VLN) requires grounding instructions, such as "turn right and stop at the door", to routes in a visual environment. The actual grounding can connect language to the environment through multiple modalities,…

计算与语言 · 计算机科学 2019-06-11 Ronghang Hu , Daniel Fried , Anna Rohrbach , Dan Klein , Trevor Darrell , Kate Saenko

Recent advances in multimodal large language models(MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring such capabilities…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Peirong Zhang , Yidan Zhang , Luxiao Xu , Jinliang Lin , Zonghao Guo , Fengxiang Wang , Xue Yang , Kaiwen Wei , Lei Wang

We study an important, yet largely unexplored problem of large-scale cross-modal visual localization by matching ground RGB images to a geo-referenced aerial LIDAR 3D point cloud (rendered as depth images). Prior works were demonstrated on…

计算机视觉与模式识别 · 计算机科学 2020-09-15 Niluthpol Chowdhury Mithun , Karan Sikka , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

Seemingly simple natural language requests to a robot are generally underspecified, for example "Can you bring me the wireless mouse?" Flat images of candidate mice may not provide the discriminative information needed for "wireless." The…

计算与语言 · 计算机科学 2021-09-16 Jesse Thomason , Mohit Shridhar , Yonatan Bisk , Chris Paxton , Luke Zettlemoyer

Many existing motion prediction approaches rely on symbolic perception outputs to generate agent trajectories, such as bounding boxes, road graph information and traffic lights. This symbolic representation is a high-level abstraction of…

The ability to decompose scenes in terms of abstract building blocks is crucial for general intelligence. Where those basic building blocks share meaningful properties, interactions and other regularities across scenes, such decompositions…

计算机视觉与模式识别 · 计算机科学 2019-02-01 Christopher P. Burgess , Loic Matthey , Nicholas Watters , Rishabh Kabra , Irina Higgins , Matt Botvinick , Alexander Lerchner

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Sayan Deb Sarkar , Ondrej Miksik , Marc Pollefeys , Daniel Barath , Iro Armeni

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding,…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Henry Zheng , Hao Shi , Qihang Peng , Yong Xien Chng , Rui Huang , Yepeng Weng , Zhongchao Shi , Gao Huang

Traditional approaches for learning 3D object categories have been predominantly trained and evaluated on synthetic datasets due to the unavailability of real 3D-annotated category-centric data. Our main goal is to facilitate advances in…

计算机视觉与模式识别 · 计算机科学 2021-09-02 Jeremy Reizenstein , Roman Shapovalov , Philipp Henzler , Luca Sbordone , Patrick Labatut , David Novotny

Autonomous driving and assistance systems rely on annotated data from traffic and road scenarios to model and learn the various object relations in complex real-world scenarios. Preparation and training of deploy-able deep learning…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Shubham Dokania , A. H. Abdul Hafez , Anbumani Subramanian , Manmohan Chandraker , C. V. Jawahar

Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward…

Navigational signs enable humans to navigate unfamiliar environments without maps. This work studies how robots can similarly exploit signs for mapless navigation in the open world. A central challenge lies in interpreting signs: real-world…

机器人学 · 计算机科学 2026-02-16 Nicky Zimmerman , Joel Loo , Benjamin Koh , Zishuo Wang , David Hsu