English
Related papers

Related papers: SUGAMAN: Describing Floor Plans for Visually Impai…

200 papers

Buildings are a central feature of human culture and are increasingly being analyzed with computational methods. However, recent works on computational building understanding have largely focused on natural imagery of buildings, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Keren Ganon , Morris Alper , Rachel Mikulinsky , Hadar Averbuch-Elor

Mobile applications (apps) are integral to our daily lives, offering diverse services and functionalities. They enable sighted users to access information coherently in an extremely convenient manner. However, it remains unclear if visually…

Software Engineering · Computer Science 2025-02-24 Mengxi Zhang , Huaxiao Liu , Yuheng Zhou , Chunyang Chen , Pei Huang , Jian Zhao

Unsupervised semantic segmentation aims to obtain high-level semantic representation on low-level visual features without manual annotations. Most existing methods are bottom-up approaches that try to group pixels into regions based on…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Zhaoyuan Yin , Pichao Wang , Fan Wang , Xianzhe Xu , Hanling Zhang , Hao Li , Rong Jin

Most previous studies integrate cognitive language processing signals (e.g., eye-tracking or EEG data) into neural models of natural language processing (NLP) just by directly concatenating word embeddings with cognitive features, ignoring…

Computation and Language · Computer Science 2023-11-15 Yuqi Ren , Deyi Xiong

We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates…

Machine Learning · Computer Science 2019-04-09 Khanh Nguyen , Debadeepta Dey , Chris Brockett , Bill Dolan

Aligning 3D scene graphs is a crucial initial step for several applications in robot navigation and embodied perception. Current methods in 3D scene graph alignment often rely on single-modality point cloud data and struggle with incomplete…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Binod Singh , Sayan Deb Sarkar , Iro Armeni

There has been a growing adoption of computer vision tools and technologies in architectural design workflows over the past decade. Notable use cases include point cloud generation, visual content analysis, and spatial awareness for robotic…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Demircan Tas , Rohit Priyadarshi Sanatani

Learning medical visual representations through vision-language pre-training has reached remarkable progress. Despite the promising performance, it still faces challenges, i.e., local alignment lacks interpretability and clinical relevance,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Qingqiu Li , Xiaohan Yan , Jilan Xu , Runtian Yuan , Yuejie Zhang , Rui Feng , Quanli Shen , Xiaobo Zhang , Shujun Wang

The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language instruction, in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Gengze Zhou , Yicong Hong , Zun Wang , Chongyang Zhao , Mohit Bansal , Qi Wu

Sparsity-based representations have recently led to notable results in various visual recognition tasks. In a separate line of research, Riemannian manifolds have been shown useful for dealing with features and models that do not lie in…

Machine Learning · Computer Science 2015-05-21 Mehrtash Harandi , Richard Hartley , Chunhua Shen , Brian Lovell , Conrad Sanderson

We introduce ShelfGaussian, an open-vocabulary multi-modal Gaussian-based 3D scene understanding framework supervised by off-the-shelf vision foundation models (VFMs). Gaussian-based methods have demonstrated superior performance and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Lingjun Zhao , Yandong Luo , James Hays , Lu Gan

Exam documents are essential educational materials for exam preparation. However, they pose a significant academic barrier for blind and visually impaired students, as they are often created without accessibility considerations. Typically,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 David Wilkening , Omar Moured , Thorsten Schwarz , Karin Muller , Rainer Stiefelhagen

We present a strong baseline that surpasses the performance of previously published methods on the Habitat Challenge task of navigating to a target object in indoor environments. Our method is motivated from primary failure modes of prior…

Robotics · Computer Science 2022-03-15 Haokuan Luo , Albert Yue , Zhang-Wei Hong , Pulkit Agrawal

We present our approach to improve room classification task on floor plan maps of buildings by representing floor plans as undirected graphs and leveraging graph neural networks to predict the room categories. Rooms in the floor plans are…

Machine Learning · Computer Science 2021-08-16 Abhishek Paudel , Roshan Dhakal , Sakshat Bhattarai

The goal of weakly-supervised video moment retrieval is to localize the video segment most relevant to the given natural language query without access to temporal annotations during training. Prior strongly- and weakly-supervised approaches…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Reuben Tan , Huijuan Xu , Kate Saenko , Bryan A. Plummer

In this work, we propose a geometry-aware semi-supervised framework for fine-grained building function recognition, utilizing geometric relationships among multi-source data to enhance pseudo-label accuracy in semi-supervised learning,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Weijia Li , Jinhua Yu , Dairong Chen , Yi Lin , Runmin Dong , Xiang Zhang , Conghui He , Haohuan Fu

This paper presents an innovative approach called BGTAI to simplify multimodal understanding by utilizing gloss-based annotation as an intermediate step in aligning Text and Audio with Images. While the dynamic temporal factors in textual…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Sen Fang , Sizhou Chen , Yalin Feng , Xiaofeng Zhang , Teik Toe Teoh

We present a universal framework to model contextualized sentence representations with visual awareness that is motivated to overcome the shortcomings of the multimodal parallel data with manual annotations. For each sentence, we first…

Computation and Language · Computer Science 2019-11-12 Zhuosheng Zhang , Rui Wang , Kehai Chen , Masao Utiyama , Eiichiro Sumita , Hai Zhao

In this paper, we teach machines to understand visuals and natural language by learning the mapping between sentences and noisy video snippets without explicit annotations. Firstly, we define a self-supervised learning framework that…

Computer Vision and Pattern Recognition · Computer Science 2021-01-12 Yujie Zhong , Linhai Xie , Sen Wang , Lucia Specia , Yishu Miao

Diagrammatic reasoning (DR) is pervasive in human problem solving as a powerful adjunct to symbolic reasoning based on language-like representations. The research reported in this paper is a contribution to building a general purpose DR…

Artificial Intelligence · Computer Science 2014-01-17 Bonny Banerjee , B. Chandrasekaran