中文
相关论文

相关论文: MoniRefer: A Real-world Large-scale Multi-modal Da…

200 篇论文

Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and…

In Global Navigation Satellite System (GNSS)-denied environments such as indoor parking structures or dense urban canyons, achieving accurate and robust vehicle positioning remains a significant challenge. This paper proposes a…

机器人学 · 计算机科学 2025-06-25 Ola Elmaghraby , Eslam Mounier , Paulo Ricardo Marques de Araujo , Aboelmagd Noureldin

Although the majority of recent autonomous driving systems concentrate on developing perception methods based on ego-vehicle sensors, there is an overlooked alternative approach that involves leveraging intelligent roadside cameras to help…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Lei Yang , Jiaxin Yu , Xinyu Zhang , Jun Li , Li Wang , Yi Huang , Chuang Zhang , Hong Wang , Yiming Li

Visual Grounding (VG) aims to utilize given natural language queries to locate specific target objects within images. While current transformer-based approaches demonstrate strong localization performance in standard scene (i.e, scenarios…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Jiangnan Xie , Xiaolong Zheng , Liang Zheng

This paper presents a framework for jointly grounding objects that follow certain semantic relationship constraints given in a scene graph. A typical natural scene contains several objects, often exhibiting visual relationships of varied…

计算机视觉与模式识别 · 计算机科学 2022-11-04 Aditay Tripathi , Anand Mishra , Anirban Chakraborty

Grounding natural language in 3D environments is a critical step toward achieving robust 3D vision-language alignment. Current datasets and models for 3D visual grounding predominantly focus on identifying and localizing objects from…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Zhuofan Zhang , Ziyu Zhu , Junhao Li , Pengxiang Li , Tianxu Wang , Tengyu Liu , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Siyuan Huang , Qing Li

With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visual modality. However, most existing tokenizers are designed…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Dong Zhuo , Wenzhao Zheng , Sicheng Zuo , Siming Yan , Lu Hou , Jie Zhou , Jiwen Lu

The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate efficacy in a range of…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Bangning Wei , Joshua Maraval , Meriem Outtas , Kidiyo Kpalma , Nicolas Ramin , Lu Zhang

This paper tackles the challenging task of 3D visual grounding-locating a specific object in a 3D point cloud scene based on text descriptions. Existing methods fall into two categories: top-down and bottom-up methods. Top-down methods rely…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Yang Liu , Daizong Liu , Wei Hu

Visually poor scenarios are one of the main sources of failure in visual localization systems in outdoor environments. To address this challenge, we present MOZARD, a multi-modal localization system for urban outdoor environments using…

机器人学 · 计算机科学 2020-03-04 Lukas Schaupp , Patrick Pfreundschuh , Mathias Buerki , Cesar Cadena , Roland Siegwart , Juan Nieto

The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing studies are facing two common challenges: 1) they are short of…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Xueying Jiang , Lewei Lu , Ling Shao , Shijian Lu

Determining the location of an image anywhere on Earth is a complex visual task, which makes it particularly relevant for evaluating computer vision algorithms. Yet, the absence of standard, large-scale, open-access datasets with reliably…

Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human environments and frequently critical to understand video. To…

计算机视觉与模式识别 · 计算机科学 2023-05-08 Weijia Wu , Yuzhong Zhao , Zhuang Li , Jiahong Li , Hong Zhou , Mike Zheng Shou , Xiang Bai

Concurrent perception datasets for autonomous driving are mainly limited to frontal view with sensors mounted on the vehicle. None of them is designed for the overlooked roadside perception tasks. On the other hand, the data captured from…

计算机视觉与模式识别 · 计算机科学 2022-03-28 Xiaoqing Ye , Mao Shu , Hanyu Li , Yifeng Shi , Yingying Li , Guangjie Wang , Xiao Tan , Errui Ding

Intelligent Transportation Systems (ITS) require reliable environmental perception to support safe and efficient transportation. With the rapid development of Vehicle-to-everything (V2X), roadside perception has become an effective means to…

机器人学 · 计算机科学 2026-05-08 Yuhan Xia , Runxin Zhao , Hanyang Zhuang , Chunxiang Wang , Ming Yang

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act…

We propose Visual Query Detection (VQD), a new visual grounding task. In VQD, a system is guided by natural language to localize a variable number of objects in an image. VQD is related to visual referring expression recognition, where the…

计算机视觉与模式识别 · 计算机科学 2019-04-15 Manoj Acharya , Karan Jariwala , Christopher Kanan

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Oskar Kristoffersen , Alba Reinders Sánchez , Morten Rieger Hannemose , Anders Bjorholm Dahl , Dim P. Papadopoulos

Neural Radiance Fields (NeRF) has achieved impressive results in single object scene reconstruction and novel view synthesis, which have been demonstrated on many single modality and single object focused indoor scene datasets like DTU,…

计算机视觉与模式识别 · 计算机科学 2023-01-18 Chongshan Lu , Fukun Yin , Xin Chen , Tao Chen , Gang YU , Jiayuan Fan

We introduce Referring 3D Gaussian Splatting Segmentation (R3DGS), a new task that aims to segment target objects in a 3D Gaussian scene based on natural language descriptions, which often contain spatial relationships or object attributes.…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Shuting He , Guangquan Jie , Changshuo Wang , Yun Zhou , Shuming Hu , Guanbin Li , Henghui Ding