中文
相关论文

相关论文: PoseRefer: Pathway-Local Parameters for Semantical…

200 篇论文

Object pose estimation is a fundamental problem in computer vision and plays a critical role in virtual reality and embodied intelligence, where agents must understand and interact with objects in 3D space. Recently, score based generative…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Diya He , Qingchen Liu , Cong Zhang , Jiahu Qin

Current RGB-based 6D object pose estimation methods have achieved noticeable performance on datasets and real world applications. However, predicting 6D pose from single 2D image features is susceptible to disturbance from changing of…

计算机视觉与模式识别 · 计算机科学 2022-07-04 Jun Wu , Lilu Liu , Yue Wang , Rong Xiong

Low-overhead visual place recognition (VPR) is a highly active research topic. Mobile robotics applications often operate under low-end hardware, and even more hardware capable systems can still benefit from freeing up onboard system…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Bruno Arcanjo , Bruno Ferrarini , Michael Milford , Klaus D. McDonald-Maier , Shoaib Ehsan

As human-robot collaboration advances, natural and flexible communication methods are essential for effective robot control. Traditional methods relying on a single modality or rigid rules struggle with noisy or misaligned data as well as…

机器人学 · 计算机科学 2025-04-03 Petr Vanc , Karla Stepanova

Robots are finding wider adoption in human environments, increasing the need for natural human-robot interaction. However, understanding a natural language command requires the robot to infer the intended task and how to decompose it into…

机器人学 · 计算机科学 2026-02-05 Julia Kuhn , Francesco Verdoja , Tsvetomila Mihaylova , Ville Kyrki

Seemingly simple natural language requests to a robot are generally underspecified, for example "Can you bring me the wireless mouse?" Flat images of candidate mice may not provide the discriminative information needed for "wireless." The…

计算与语言 · 计算机科学 2021-09-16 Jesse Thomason , Mohit Shridhar , Yonatan Bisk , Chris Paxton , Luke Zettlemoyer

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Wenbin Tan , Jiawen Lin , Fangyong Wang , Yuan Xie , Yong Xie , Yachao Zhang , Yanyun Qu

Spatial perception aims to estimate camera motion and scene structure from visual observations, a problem traditionally addressed through geometric modeling and physical consistency constraints. Recent learning-based methods have…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Haichao Zhu , Zhaorui Yang , Qian Zhang

Satellite imagery differs fundamentally from natural images: its aerial viewpoint, very high resolution, diverse scale variations, and abundance of small objects demand both region-level spatial reasoning and holistic scene understanding.…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Emanuel Sánchez Aimar , Gulnaz Zhambulova , Fahad Shahbaz Khan , Yonghao Xu , Michael Felsberg

Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Minghui Hou , Wei-Hsing Huang , Shaofeng Liang , Daizong Liu , Tai-Hao Wen , Gang Wang , Runwei Guan , Weiping Ding

Visual grounding is a task that aims to locate a target object according to a natural language expression. As a multi-modal task, feature interaction between textual and visual inputs is vital. However, previous solutions mainly handle each…

计算机视觉与模式识别 · 计算机科学 2022-06-23 Chonghan Chen , Qi Jiang , Chih-Hao Wang , Noel Chen , Haohan Wang , Xiang Li , Bhiksha Raj

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We introduce m2sv, a…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Yosub Shin , Michael Buriek , Igor Molybog

This paper presents a framework for jointly grounding objects that follow certain semantic relationship constraints given in a scene graph. A typical natural scene contains several objects, often exhibiting visual relationships of varied…

计算机视觉与模式识别 · 计算机科学 2022-11-04 Aditay Tripathi , Anand Mishra , Anirban Chakraborty

In this paper, we present a neat yet effective transformer-based framework for visual grounding, namely TransVG, to address the task of grounding a language query to the corresponding region onto an image. The state-of-the-art methods,…

计算机视觉与模式识别 · 计算机科学 2022-01-17 Jiajun Deng , Zhengyuan Yang , Tianlang Chen , Wengang Zhou , Houqiang Li

Vision-language models encode continuous geometry that their text pathway fails to express: a 6,000-parameter linear probe extracts hand joint angles at 6.1 degrees MAE from frozen features, while the best text output achieves only 20.0…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Yakov Pyotr Shkolnikov

The bundle of geometry and appearance in computer vision has proven to be a promising solution for robots across a wide variety of applications. Stereo cameras and RGB-D sensors are widely used to realise fast 3D reconstruction and…

计算机视觉与模式识别 · 计算机科学 2016-11-15 Xuanpeng Li , Rachid Belaroussi

Over the past decades, the addition of hundreds of sensors to modern vehicles has led to an exponential increase in their capabilities. This allows for novel approaches to interaction with the vehicle that go beyond traditional touch-based…

人机交互 · 计算机科学 2021-11-04 Amr Gomaa , Guillermo Reyes , Michael Feld

As robots become more ubiquitous and capable, it becomes ever more important to enable untrained users to easily interact with them. Recently, this has led to study of the language grounding problem, where the goal is to extract…

计算与语言 · 计算机科学 2012-07-03 Cynthia Matuszek , Nicholas FitzGerald , Luke Zettlemoyer , Liefeng Bo , Dieter Fox

Functional affordance grounding requires more than recognizing an object: an agent must localize the specific region that supports an interaction, such as the handle to pull or the button to press. This is difficult for training-free…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Qirui Wang , Jingyi He , Yining Pan , Xulei Yang , Shijie Li

Vision-language foundation models have shown remarkable performance in various zero-shot settings such as image retrieval, classification, or captioning. But so far, those models seem to fall behind when it comes to zero-shot localization…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Walid Bousselham , Felix Petersen , Vittorio Ferrari , Hilde Kuehne