English
Related papers

Related papers: SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic V…

200 papers

To endow machines with the ability to perceive the real-world in a three dimensional representation as we do as humans is a fundamental and long-standing topic in Artificial Intelligence. Given different types of visual inputs such as…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Bo Yang

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech…

Computation and Language · Computer Science 2025-04-29 Tuochao Chen , Qirui Wang , Runlin He , Shyam Gollakota

Artificial intelligence (AI) has revolutionized human cognitive abilities and facilitated the development of new AI entities capable of interacting with humans in both physical and virtual environments. Despite the existence of virtual…

Human-Computer Interaction · Computer Science 2024-01-30 Zhenliang Zhang , Zeyu Zhang , Ziyuan Jiao , Yao Su , Hangxin Liu , Wei Wang , Song-Chun Zhu

We propose LENS, a modular approach for tackling computer vision problems by leveraging the power of large language models (LLMs). Our system uses a language model to reason over outputs from a set of independent and highly descriptive…

Computation and Language · Computer Science 2023-06-29 William Berrios , Gautam Mittal , Tristan Thrush , Douwe Kiela , Amanpreet Singh

Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any objects,…

Computer Vision and Pattern Recognition · Computer Science 2019-11-19 Xiaoze Jiang , Jing Yu , Zengchang Qin , Yingying Zhuang , Xingxing Zhang , Yue Hu , Qi Wu

A chief goal of artificial intelligence is to build machines that think like people. Yet it has been argued that deep neural network architectures fail to accomplish this. Researchers have asserted these models' limitations in the domains…

Machine Learning · Computer Science 2024-08-09 Luca M. Schulze Buschoff , Elif Akata , Matthias Bethge , Eric Schulz

Multimodal large language models (LMMs) excel in world knowledge and problem-solving abilities. Through the use of a world-facing camera and contextual AI, emerging smart accessories aim to provide a seamless interface between humans and…

Human-Computer Interaction · Computer Science 2024-02-01 Robert Konrad , Nitish Padmanaban , J. Gabriel Buckmaster , Kevin C. Boyle , Gordon Wetzstein

People with visual impairments (PVI) use a variety of assistive technologies to navigate their daily lives, and conversational AI (CAI) tools are a growing part of this toolset. Much existing HCI research has focused on the technical…

Human-Computer Interaction · Computer Science 2025-10-15 Jeanne Choi , Dasom Choi , Sejun Jeong , Hwajung Hong , Joseph Seering

We discuss two kinds of semantics relevant to Computer Vision (CV) systems - Visual Semantics and Lexical Semantics. While visual semantics focus on how humans build concepts when using vision to perceive a target reality, lexical semantics…

Computer Vision and Pattern Recognition · Computer Science 2022-12-14 Fausto Giunchiglia , Mayukh Bagchi , Xiaolei Diao

Visual blends combine elements from two distinct visual concepts into a single, integrated image, with the goal of conveying ideas through imaginative and often thought-provoking visuals. Communicating abstract concepts through visual…

Human-Computer Interaction · Computer Science 2025-02-25 Zhida Sun , Zhenyao Zhang , Yue Zhang , Min Lu , Dani Lischinski , Daniel Cohen-Or , Hui Huang

Robots deployed in real-world environments, such as homes, must not only navigate safely but also understand their surroundings and adapt to changes in the environment. To perform tasks efficiently, they must build and maintain a semantic…

The increasing complexity of machine learning models in computer vision, particularly in face verification, requires the development of explainable artificial intelligence (XAI) to enhance interpretability and transparency. This study…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Miriam Doh , Caroline Mazini Rodrigues , N. Boutry , L. Najman , Matei Mancas , Bernard Gosselin

Lightweight augmented reality (AR) glasses are increasingly entering everyday use, extending interaction design beyond short, isolated sessions. However, most existing gesture vocabularies are inherited from VR headsets or early AR goggles.…

Human-Computer Interaction · Computer Science 2026-03-17 Wei Wu , Binyan Xu , Soonhyeon Kweon , Yujie Wang , Leanne Chukoskie , Casper Harteveld

We have developed a system for automatic facial expression recognition, which runs on Google Glass and delivers real-time social cues to the wearer. We evaluate the system as a behavioral aid for children with Autism Spectrum Disorder…

Human-Computer Interaction · Computer Science 2020-02-18 Catalin Voss , Peter Washington , Nick Haber , Aaron Kline , Jena Daniels , Azar Fazel , Titas De , Beth McCarthy , Carl Feinstein , Terry Winograd , Dennis Wall

This paper introduces Teachable Reality, an augmented reality (AR) prototyping tool for creating interactive tangible AR applications with arbitrary everyday objects. Teachable Reality leverages vision-based interactive machine teaching…

Human-Computer Interaction · Computer Science 2023-02-23 Kyzyl Monteiro , Ritik Vatsal , Neil Chulpongsatorn , Aman Parnami , Ryo Suzuki

Current vision and language tasks usually take complete visual data (e.g., raw images or videos) as input, however, practical scenarios may often consist the situations where part of the visual information becomes inaccessible due to…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Ye Zhu , Yu Wu , Yi Yang , Yan Yan

Image-text matching has been a hot research topic bridging the vision and language areas. It remains challenging because the current representation of image usually lacks global semantic concepts as in its corresponding text caption. To…

Computer Vision and Pattern Recognition · Computer Science 2019-09-09 Kunpeng Li , Yulun Zhang , Kai Li , Yuanyuan Li , Yun Fu

Augmented Reality (AR) smartglasses are increasingly regarded as the next generation personal computing platform. However, there is a lack of understanding about how to design communication systems using them. We present ARcall, a novel…

Human-Computer Interaction · Computer Science 2022-03-10 Hemant Bhaskar Surale , Yu Jiang Tham , Brian A. Smith , Rajan Vaish

Language grounded image understanding tasks have often been proposed as a method for evaluating progress in artificial intelligence. Ideally, these tasks should test a plethora of capabilities that integrate computer vision, reasoning, and…

Machine Learning · Computer Science 2019-05-28 Kushal Kafle , Robik Shrestha , Christopher Kanan

This study explores the integration of generative artificial intelligence (AI), specifically large language models, with multi-modal analogical reasoning as an innovative approach to enhance science, technology, engineering, and mathematics…

Artificial Intelligence · Computer Science 2023-08-22 Chen Cao , Zijian Ding , Gyeong-Geon Lee , Jiajun Jiao , Jionghao Lin , Xiaoming Zhai
‹ Prev 1 3 4 5 6 7 10 Next ›