English
Related papers

Related papers: SUGAMAN: Describing Floor Plans for Visually Impai…

200 papers

This paper addresses the problem of building augmented metric representations of scenes with semantic information from RGB-D images. We propose a complete framework to create an enhanced map representation of the environment with…

Computer Vision and Pattern Recognition · Computer Science 2020-03-16 Renato Martins , Dhiego Bersan , Mario F. M. Campos , Erickson R. Nascimento

Neural architecture search (NAS) in expressive search spaces is a computationally hard problem, but it also holds the potential to automatically discover completely novel and performant architectures. To achieve this we need effective…

Neural and Evolutionary Computing · Computer Science 2025-12-05 Adri Gómez Martín , Felix Möller , Steven McDonagh , Monica Abella , Manuel Desco , Elliot J. Crowley , Aaron Klein , Linus Ericsson

Prior highly-tuned image parsing models are usually studied in a certain domain with a specific set of semantic labels and can hardly be adapted into other scenarios (e.g., sharing discrepant label granularity) without extensive…

Computer Vision and Pattern Recognition · Computer Science 2021-01-27 Liang Lin , Yiming Gao , Ke Gong , Meng Wang , Xiaodan Liang

Panoptic maps enable robots to reason about both geometry and semantics. However, open-vocabulary models repeatedly produce closely related labels that split panoptic entities and degrade volumetric consistency. The proposed UPPM advances…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Mohamad Al Mdfaa , Raghad Salameh , Geesara Kulathunga , Sergey Zagoruyko , Gonzalo Ferrer

This work provides a unified framework for addressing the problem of visual supervised domain adaptation and generalization with deep models. The main idea is to exploit the Siamese architecture to learn an embedding subspace that is…

Computer Vision and Pattern Recognition · Computer Science 2017-10-02 Saeid Motiian , Marco Piccirilli , Donald A. Adjeroh , Gianfranco Doretto

Sign language is a gesture-based symbolic communication medium among speech and hearing impaired people. It also serves as a communication bridge between non-impaired and impaired populations. Unfortunately, in most situations, a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Prasun Roy , Saumik Bhattacharya , Partha Pratim Roy , Umapada Pal

Natural language instructions for visual navigation often use scene descriptions (e.g., "bedroom") and object references (e.g., "green chairs") to provide a breadcrumb trail to a goal location. This work presents a transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Abhinav Moudgil , Arjun Majumdar , Harsh Agrawal , Stefan Lee , Dhruv Batra

Maps --- specifically floor plans --- are useful for a variety of tasks from arranging furniture to designating conceptual or functional spaces (e.g., kitchen, walkway). We present a simple algorithm for quickly laying a floor plan (or…

Human-Computer Interaction · Computer Science 2016-06-16 Leo Bowen-Biggs , Suzanne Dazo , Yili Zhang , Alex Hubers , Matthew Rueben , Ross Sowell , William D. Smart , Cindy Grimm

Abstractive text summarization is one of the areas influenced by the emergence of pre-trained language models. Current pre-training works in abstractive summarization give more points to the summaries with more words in common with the main…

Computation and Language · Computer Science 2021-09-10 Alireza Salemi , Emad Kebriaei , Ghazal Neisi Minaei , Azadeh Shakery

Transient objects in casual multi-view captures cause ghosting artifacts in 3D Gaussian Splatting (3DGS) reconstruction. Existing solutions relied on scene decomposition at significant memory cost or on motion-based heuristics that were…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Aditi Prabakaran , Priyesh Shukla

We introduce an open-source system called SIGMA (short for "Situated Interactive Guidance, Monitoring, and Assistance") as a platform for conducting research on task-assistive agents in mixed-reality scenarios. The system leverages the…

Human-Computer Interaction · Computer Science 2024-05-24 Dan Bohus , Sean Andrist , Nick Saw , Ann Paradiso , Ishani Chakraborty , Mahdi Rad

Humans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior…

Robotics · Computer Science 2025-03-10 Jiaxin Li , Weiqi Huang , Zan Wang , Wei Liang , Huijun Di , Feng Liu

Video summarization is a crucial technique for social understanding, enabling efficient browsing of massive multimedia content and extraction of key information from social platforms. Most existing unsupervised summarization methods rely on…

Artificial Intelligence · Computer Science 2026-01-22 Haizhou Liu , Haodong Jin , Yiming Wang , Hui Yu

Large Language Models (LLMs) are unable to reliably reason about specific physical systems. Attempts to imbue LLMs with knowledge of the necessary physics concepts have shown great promise, but explainability and validation remain open…

Artificial Intelligence · Computer Science 2026-05-22 Sean Memery , Kartic Subr

3D semantic field learning is crucial for applications like autonomous navigation, AR/VR, and robotics, where accurate comprehension of 3D scenes from limited viewpoints is essential. Existing methods struggle under sparse view conditions,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Kangjie Chen , BingQuan Dai , Minghan Qin , Dongbin Zhang , Peihao Li , Yingshuang Zou , Haoqian Wang

Recently, salient object detection (SOD) methods have achieved impressive performance. However, salient regions predicted by existing methods usually contain unsaturated regions and shadows, which limits the model for reliable fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Yao Yuan , Pan Gao , Qun Dai , Jie Qin , Wei Xiang

This paper studies the problem of image-goal navigation which involves navigating to the location indicated by a goal image in a novel previously unseen environment. To tackle this problem, we design topological representations for space…

Computer Vision and Pattern Recognition · Computer Science 2020-06-01 Devendra Singh Chaplot , Ruslan Salakhutdinov , Abhinav Gupta , Saurabh Gupta

Modern graph neural networks (GNNs) can be sensitive to changes in the input graph structure and node features, potentially resulting in unpredictable behavior and degraded performance. In this work, we introduce a spectral framework known…

Machine Learning · Computer Science 2024-10-11 Wuxinlin Cheng , Chenhui Deng , Ali Aghdaei , Zhiru Zhang , Zhuo Feng

Koopman analysis of a general dynamics system provides a linear Koopman operator and an embedded eigenfunction space, enabling the application of standard techniques from linear analysis. However, in practice, deriving exact operators and…

Systems and Control · Electrical Eng. & Systems 2025-04-29 Alexander Estornell , Leonard Jung , Alenna Spiro , Mario Sznaier , Michael Everett

An estimated 253 million people have visual impairments. These visual impairments affect everyday lives, and limit their understanding of the outside world. This can pose a risk to health from falling or collisions. We propose a solution to…

Human-Computer Interaction · Computer Science 2023-03-30 Alexander Mehta , Ritik Jalisatgi