中文
相关论文

相关论文: Unified Representation Space for 3D Visual Groundi…

200 篇论文

3D Vision-Language Pre-training (3D-VLP) aims to provide a pre-train model which can bridge 3D scenes with natural language, which is an important technique for embodied intelligence. However, current 3D-VLP datasets are hindered by limited…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Dejie Yang , Zhu Xu , Wentao Mo , Qingchao Chen , Siyuan Huang , Yang Liu

Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Yaqi Zhao , Wang Lin , Zijian Zhang , Miles Yang , Jingyuan Chen , Wentao Zhang , Zhao Zhong , Liefeng Bo

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Mingtao Feng , Zhen Li , Qi Li , Liang Zhang , XiangDong Zhang , Guangming Zhu , Hui Zhang , Yaonan Wang , Ajmal Mian

Following rapid advancements in text and image generation, research has increasingly shifted towards 3D generation. Unlike the well-established pixel-based representation in images, 3D representations remain diverse and fragmented,…

Comprehensive visual, geometric, and semantic understanding of a 3D scene is crucial for successful execution of robotic tasks, especially in unstructured and complex environments. Additionally, to make robust decisions, it is necessary for…

机器人学 · 计算机科学 2026-03-13 Christian Maurer , Snehal Jauhri , Sophie Lueth , Georgia Chalvatzaki

3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language descriptions. Recent zero-shot methods leverage 2D vision-language models (LVLMs). However,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Cuong Huynh , Maxim Popov , Denis Gridusov , Sergey Kolyubin

Visual grounding aims to align visual information of specific regions of images with corresponding natural language expressions. Current visual grounding methods leverage pre-trained visual and language backbones independently to obtain…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Jiaxi Wang , Wenhui Hu , Xueyang Liu , Beihu Wu , Yuting Qiu , YingYing Cai

We propose a system that learns to detect objects and infer their 3D poses in RGB-D images. Many existing systems can identify objects and infer 3D poses, but they heavily rely on human labels and 3D annotations. The challenge here is to…

计算机视觉与模式识别 · 计算机科学 2020-11-02 Mihir Prabhudesai , Shamit Lal , Hsiao-Yu Fish Tung , Adam W. Harley , Shubhankar Potdar , Katerina Fragkiadaki

Typical inverse rendering methods focus on learning implicit neural scene representations by modeling the geometry, materials and illumination separately, which entails significant computations for optimization. In this work we design a…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Shuang Wu , Songlin Tang , Guangming Lu , Jianzhuang Liu , Wenjie Pei

Vision-Language Pretraining (VLP) has demonstrated remarkable capabilities in learning visual representations from textual descriptions of images without annotations. Yet, effective VLP demands large-scale image-text pairs, a resource that…

计算机视觉与模式识别 · 计算机科学 2023-06-09 Yinda Chen , Che Liu , Wei Huang , Sibo Cheng , Rossella Arcucci , Zhiwei Xiong

In this paper, we propose an efficient multi-level convolution architecture for 3D visual grounding. Conventional methods are difficult to meet the requirements of real-time inference due to the two-stage or point-based architecture.…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Wenxuan Guo , Xiuwei Xu , Ziwei Wang , Jianjiang Feng , Jie Zhou , Jiwen Lu

Throughout the history of computer vision, while research has explored the integration of images (visual) and point clouds (geometric), many advancements in image and 3D object recognition have tended to process these modalities separately.…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Ryosuke Yamada , Kensho Hara , Hirokatsu Kataoka , Koshi Makihara , Nakamasa Inoue , Rio Yokota , Yutaka Satoh

Learning generalizable visual representations from Internet data has yielded promising results for robotics. Yet, prevailing approaches focus on pre-training 2D representations, being sub-optimal to deal with occlusions and accurately…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Shizhe Chen , Ricardo Garcia , Ivan Laptev , Cordelia Schmid

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic environments and enable open-vocabulary querying in complex…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Xianfeng Wu , Yajing Bai , Minghan Li , Xianzu Wu , Xueqi Zhao , Zhongyuan Lai , Wenyu Liu , Xinggang Wang

Generating high-quality novel view renderings of 3D Gaussian Splatting (3DGS) in scenes featuring transient objects is challenging. We propose a novel hybrid representation, termed as HybridGS, using 2D Gaussians for transient objects per…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Jingyu Lin , Jiaqi Gu , Lubin Fan , Bojian Wu , Yujing Lou , Renjie Chen , Ligang Liu , Jieping Ye

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, which struggles to exploit…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Ke Li , Ting Wang , Di Wang , Yongshan Zhu , Yiming Zhang , Tao Lei , Quan Wang

The unification of disparate maps is crucial for enabling scalable robot operation across multiple sessions and collaborative multi-robot scenarios. However, achieving a unified map robust to sensor modalities and dynamic environments…

机器人学 · 计算机科学 2025-12-24 Gilhwan Kang , Hogyun Kim , Byunghee Choi , Seokhwan Jeong , Young-Sik Shin , Younggun Cho

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplored. In this paper, we introduce Anywhere3D-Bench, a holistic…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Tianxu Wang , Zhuofan Zhang , Ziyu Zhu , Yue Fan , Jing Xiong , Pengxiang Li , Xiaojian Ma , Qing Li

Visual correspondence across image-to-image (2D-2D), image-to-point cloud (2D-3D), and point cloud-to-point cloud (3D-3D) geometric matching forms the foundation for numerous 3D vision tasks. Despite sharing a similar problem structure,…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Prajnan Goswami , Tianye Ding , Feng Liu , Huaizu Jiang

We present UniBind, a flexible and efficient approach that learns a unified representation space for seven diverse modalities -- images, text, audio, point cloud, thermal, video, and event data. Existing works, eg., ImageBind, treat the…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Yuanhuiyi Lyu , Xu Zheng , Jiazhou Zhou , Lin Wang