English
Related papers

Related papers: SUGAMAN: Describing Floor Plans for Visually Impai…

200 papers

The 3D scene graph models spatial relationships between objects, enabling the agent to efficiently navigate in a partially observable environment and predict the location of the target object.This paper proposes an original framework named…

Robotics · Computer Science 2025-06-06 Nikita Oskolkov , Huzhenyu Zhang , Dmitry Makarov , Dmitry Yudin , Aleksandr Panov

This paper tackles a novel yet challenging problem: how to transfer knowledge from the emerging Segment Anything Model (SAM) -- which reveals impressive zero-shot instance segmentation capacity -- to learn a compact panoramic semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Weiming Zhang , Yexin Liu , Xu Zheng , Lin Wang

When navigating in a man-made environment they haven't visited before--like an office building--humans employ behaviors such as reading signs and asking others for directions. These behaviors help humans reach their destinations efficiently…

Robotics · Computer Science 2025-09-26 Bhargav Chandaka , Gloria X. Wang , Haozhe Chen , Henry Che , Albert J. Zhai , Shenlong Wang

Large language models (LLMs) have demonstrated impressive results in developing generalist planning agents for diverse tasks. However, grounding these plans in expansive, multi-floor, and multi-room environments presents a significant…

Robotics · Computer Science 2023-09-29 Krishan Rana , Jesse Haviland , Sourav Garg , Jad Abou-Chakra , Ian Reid , Niko Suenderhauf

Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Gong Jingyu , Tong Kunkun , Chen Zhuoran , Yuan Chuanhan , Chen Mingang , Zhang Zhizhong , Tan Xin , Xie Yuan

Cross-lingual summarization consists of generating a summary in one language given an input document in a different language, allowing for the dissemination of relevant content across speakers of other languages. The task is challenging…

Computation and Language · Computer Science 2024-02-01 Fantine Huot , Joshua Maynez , Chris Alberti , Reinald Kim Amplayo , Priyanka Agrawal , Constanza Fierro , Shashi Narayan , Mirella Lapata

Object navigation is a core capability of embodied intelligence, enabling an agent to locate target objects in unknown environments. Recent advances in vision-language models (VLMs) have facilitated zero-shot object navigation (ZSON).…

Robotics · Computer Science 2026-02-13 Wancai Zheng , Hao Chen , Xianlong Lu , Linlin Ou , Xinyi Yu

Despite considerable progress in image classification tasks, classification models seem unaffected by the images that significantly deviate from those that appear natural to human eyes. Specifically, while human perception can easily…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Chun Tao , Timur Ibrayev , Kaushik Roy

When humans perform a task with an articulated object, they interact with the object only in a handful of ways, while the space of all possible interactions is nearly endless. This is because humans have prior knowledge about what…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Liquan Wang , Nikita Dvornik , Rafael Dubeau , Mayank Mittal , Animesh Garg

Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabeled document datasets have opened up promising directions…

Computation and Language · Computer Science 2022-04-29 Jiuxiang Gu , Jason Kuen , Vlad I. Morariu , Handong Zhao , Nikolaos Barmpalios , Rajiv Jain , Ani Nenkova , Tong Sun

Neural Architecture Search (NAS) provides state-of-the-art results when trained on well-curated datasets with annotated labels. However, annotating data or even having balanced number of samples can be a luxury for practitioners from…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Aleksandr Timofeev , Grigorios G. Chrysos , Volkan Cevher

State-of-the-art computer vision approaches rely on huge amounts of annotated data. The collection of such data is a time consuming process since it is mainly performed by humans. The literature shows that semi-automatic annotation…

Computer Vision and Pattern Recognition · Computer Science 2019-11-05 Jonas Jäger , Gereon Reus , Joachim Denzler , Viviane Wolff , Klaus Fricke-Neuderth

Learning medical visual representations from paired images and reports is a promising direction in representation learning. However, current vision-language pretraining methods in the medical domain often simplify clinical reports into…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Wei Li , Xun Gong , Jiao Li , Xiaobin Sun

Geometry problem solving (GPS) is a challenging mathematical reasoning task requiring multi-modal understanding, fusion, and reasoning. Existing neural solvers take GPS as a vision-language task but are short in the representation of…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Zhong-Zhi Li , Ming-Liang Zhang , Fei Yin , Cheng-Lin Liu

Text-goal instance navigation (TGIN) asks an agent to resolve a single, free-form description into actions that reach the correct object instance among same-category distractors. We present \textit{Context-Nav}, which elevates long,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Won Shik Jang , Ue-Hwan Kim

Visual Simultaneous Localization and Mapping (vSLAM) is a widely used technique in robotics and computer vision that enables a robot to create a map of an unfamiliar environment using a camera sensor while simultaneously tracking its…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Yasaman Haghighi , Suryansh Kumar , Jean-Philippe Thiran , Luc Van Gool

Enabling robots to autonomously discover high-level spatial concepts (e.g., rooms and walls) from primitive geometric observations (e.g., planar surfaces) within 3D Scene Graphs is essential for robust indoor navigation and mapping. These…

The acquisition of grammar has been a central question to adjudicate between theories of language acquisition. In order to conduct faster, more reproducible, and larger-scale corpus studies on grammaticality in child-caregiver…

Computation and Language · Computer Science 2024-03-22 Mitja Nikolaus , Abhishek Agrawal , Petros Kaklamanis , Alex Warstadt , Abdellah Fourtassi

The nonliteral interpretation of a text is hard to be understood by machine models due to its high context-sensitivity and heavy usage of figurative language. In this study, inspired by human reading comprehension, we propose a novel,…

Computation and Language · Computer Science 2020-01-17 Guoxiu He , Zhe Gao , Zhuoren Jiang , Yangyang Kang , Changlong Sun , Xiaozhong Liu , Wei Lu

Scene graph generation (SGG) is a sophisticated task that suffers from both complex visual features and dataset long-tail problem. Recently, various unbiased strategies have been proposed by designing novel loss functions and data balancing…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Xiaoguang Chang , Teng Wang , Shaowei Cai , Changyin Sun
‹ Prev 1 8 9 10 Next ›