English
Related papers

Related papers: AugRefer: Advancing 3D Visual Grounding via Cross-…

200 papers

Recent advances in multimodal large language models(MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring such capabilities…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Peirong Zhang , Yidan Zhang , Luxiao Xu , Jinliang Lin , Zonghao Guo , Fengxiang Wang , Xue Yang , Kaiwen Wei , Lei Wang

Visual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and scalable strategy for learning visual grounding is to utilize…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Yongfei Liu , Bo Wan , Lin Ma , Xuming He

Different from Object Detection, Visual Grounding deals with detecting a bounding box for each text-image pair. This one box for each text-image data provides sparse supervision signals. Although previous works achieve impressive results,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Weitai Kang , Gaowen Liu , Mubarak Shah , Yan Yan

Most models tasked to ground referential utterances in 2D and 3D scenes learn to select the referred object from a pool of object proposals provided by a pre-trained detector. This is limiting because an utterance may refer to visual…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Ayush Jain , Nikolaos Gkanatsios , Ishita Mediratta , Katerina Fragkiadaki

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Ziyu Zhu , Xilin Wang , Yixuan Li , Zhuofan Zhang , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Wei Liang , Qian Yu , Zhidong Deng , Siyuan Huang , Qing Li

With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can operate in diverse environments given natural language…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Haochen Zhang , Nader Zantout , Pujith Kachana , Ji Zhang , Wenshan Wang

Recent advances in imitation learning have shown significant promise for robotic control and embodied intelligence. However, achieving robust generalization across diverse mounted camera observations remains a critical challenge. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Travis Davies , Jiahuan Yan , Xiang Chen , Yu Tian , Yueting Zhuang , Yiqi Huang , Luhui Hu

Generalised 3D Referring Expression Segmentation (3D-GRES) localizes objects in 3D scenes based on natural language, even when descriptions match multiple or zero targets. Existing methods rely solely on sparse point clouds, lacking rich…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Keshen Zhou , Runnan Chen , Mingming Gong , Tongliang Liu

Transferring human motion and appearance between videos of human actors remains one of the key challenges in Computer Vision. Despite the advances from recent image-to-image translation approaches, there are several transferring contexts…

Computer Vision and Pattern Recognition · Computer Science 2021-04-29 Thiago L. Gomes , Renato Martins , João Ferreira , Rafael Azevedo , Guilherme Torres , Erickson R. Nascimento

Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization paradigm by enabling effective multi-view reconstruction in a single forward pass.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 David Huang , Guile Wu , Chengjie Huang , Bingbing Liu , Dongfeng Bai

There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve…

Computer Vision and Pattern Recognition · Computer Science 2020-03-27 Gunnar A. Sigurdsson , Jean-Baptiste Alayrac , Aida Nematzadeh , Lucas Smaira , Mateusz Malinowski , João Carreira , Phil Blunsom , Andrew Zisserman

Automating teaching presents unique challenges, as replicating human interaction and adaptability is complex. Automated systems cannot often provide nuanced, real-time feedback that aligns with students' individual learning paces or…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Ruslan Gokhman , Jialu Li , Youshan Zhang

Video Referring Expression Comprehension (REC) aims to localize a target object in videos based on the queried natural language. Recent improvements in video REC have been made using Transformer-based methods with learnable queries.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-26 Ji Jiang , Meng Cao , Tengtao Song , Long Chen , Yi Wang , Yuexian Zou

Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Dianyi Yang , Xihan Wang , Yu Gao , Shiyang Liu , Bohan Ren , Yufeng Yue , Yi Yang

Narrated instructional videos often show and describe manipulations of similar objects, e.g., repairing a particular model of a car or laptop. In this work we aim to reconstruct such objects and to localize associated narrations in 3D.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-13 Dimitri Zhukov , Ignacio Rocco , Ivan Laptev , Josef Sivic , Johannes L. Schönberger , Bugra Tekin , Marc Pollefeys

Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Hongbing Li , Linhui Xiao , Zihan Zhao , Qi Shen , Yixiang Huang , Bo Xiao , Zhanyu Ma

Open-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learning 3D language…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Zhenyang Liu , Sixiao Zheng , Siyu Chen , Cairong Zhao , Longfei Liang , Xiangyang Xue , Yanwei Fu

As a novel and challenging task, referring segmentation combines computer vision and natural language processing to localize and segment objects based on textual descriptions. While referring image segmentation (RIS) has been extensively…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Rui Li , Xiaowei Zhao

Text-to-image retrieval is a fundamental task in multimedia processing, aiming to retrieve semantically relevant cross-modal content. Traditional studies have typically approached this task as a discriminative problem, matching the text and…

Multimedia · Computer Science 2024-07-25 Yongqi Li , Hongru Cai , Wenjie Wang , Leigang Qu , Yinwei Wei , Wenjie Li , Liqiang Nie , Tat-Seng Chua

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretations. This limits their applicability in real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Wenxuan Wang , Zijia Zhao , Yisi Zhang , Yepeng Tang , Erdong Hu , Xinlong Wang , Jing Liu
‹ Prev 1 4 5 6 7 8 10 Next ›