English
Related papers

Related papers: GazePrompt: Enhancing Low Vision People's Reading …

200 papers

Currently, low-light conditions present a significant challenge for machine cognition. In this paper, rather than optimizing models by assuming that human and machine cognition are correlated, we use zero-reference low-light enhancement to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Igor Morawski , Kai He , Shusil Dangi , Winston H. Hsu

Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation stimuli, we captured 1,360 real-world photos of scenes in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Zory Zhang , Pinyuan Feng , Bingyang Wang , Tianwei Zhao , Suyang Yu , Qingying Gao , Hokin Deng , Ziqiao Ma , Yijiang Li , Dezhi Luo

In this paper, we present GazeTrak, the first acoustic-based eye tracking system on glasses. Our system only needs one speaker and four microphones attached to each side of the glasses. These acoustic sensors capture the formations of the…

Human-Computer Interaction · Computer Science 2024-02-27 Ke Li , Ruidong Zhang , Boao Chen , Siyuan Chen , Sicheng Yin , Saif Mahmud , Qikang Liang , François Guimbretière , Cheng Zhang

Individuals with language disorders often face significant communication challenges due to their limited language processing and comprehension abilities, which also affect their interactions with voice-assisted systems that mostly rely on…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-21 Seungbae Kim , Daeun Lee , Brielle Stark , Jinyoung Han

Prompt engineering is a powerful tool used to enhance the performance of pre-trained models on downstream tasks. For example, providing the prompt "Let's think step by step" improved GPT-3's reasoning accuracy to 63% on MutiArith while…

Computer Vision and Pattern Recognition · Computer Science 2023-09-25 Cheng Shi , Sibei Yang

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, and temporal…

It has always been a rather tough task to communicate with someone possessing a hearing impairment. One of the most tested ways to establish such a communication is through the use of sign based languages. However, not many people are aware…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Sharanya Mukherjee , Md Hishaam Akhtar , Kannadasan R

We present See-Through Face Display, an eye-contact display system designed to enhance gaze awareness in both human-to-human and human-to-avatar communication. The system addresses the limitations of existing gaze correction methods by…

Human-Computer Interaction · Computer Science 2024-10-23 Kazuya Izumi , Ryosuke Hyakuta , Ippei Suzuki , Yoichi Ochiai

The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by partial vision…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Wiktor Mucha , Florin Cuconasu , Naome A. Etori , Valia Kalokyri , Giovanni Trappolini

Multimodal Large Language Models (MLLMs) such as GPT-4V and Gemini Pro face challenges in achieving human-level perception in Visual Question Answering (VQA), particularly in object-oriented perception tasks which demand fine-grained…

Computation and Language · Computer Science 2024-04-09 Songtao Jiang , Yan Zhang , Chenyi Zhou , Yeying Jin , Yang Feng , Jian Wu , Zuozhu Liu

Appearance-based supervised methods with full-face image input have made tremendous advances in recent gaze estimation tasks. However, intensive human annotation requirement inhibits current methods from achieving industrial level accuracy…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Yangzhou Jiang , Yinxin Lin , Yaoming Wang , Teng Li , Bilian Ke , Bingbing Ni

We present Lattice Menu, a gaze-based marking menu utilizing a lattice of visual anchors that helps perform accurate gaze pointing for menu item selection. Users who know the location of the desired item can leverage target-assisted gaze…

Human-Computer Interaction · Computer Science 2025-11-27 Taejun Kim , Auejin Ham , Sunggeun Ahn , Geehyuk Lee

Modern information querying systems are progressively incorporating multimodal inputs like vision and audio. However, the integration of gaze -- a modality deeply linked to user intent and increasingly accessible via gaze-tracking wearables…

Human-Computer Interaction · Computer Science 2024-05-14 Zeyu Wang , Yuanchun Shi , Yuntao Wang , Yuchen Yao , Kun Yan , Yuhan Wang , Lei Ji , Xuhai Xu , Chun Yu

Advanced multimodal AI agents can now collaborate with users to solve challenges in the world. Yet, these emerging contextual AI systems rely on explicit communication channels between the user and system. We hypothesize that implicit…

We present GazeGen, a user interaction system that generates visual content (images and videos) for locations indicated by the user's eye gaze. GazeGen allows intuitive manipulation of visual content by targeting regions of interest with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 He-Yen Hsieh , Ziyun Li , Sai Qian Zhang , Wei-Te Mark Ting , Kao-Den Chang , Barbara De Salvo , Chiao Liu , H. T. Kung

Transcripts displayed on dictation interfaces can be hard to read due to recognition errors and disfluencies. LLM-based text auto-correction could help, but changing the text during production could lead to distraction and unintended…

Human-Computer Interaction · Computer Science 2026-03-24 Zhaohui Liang , Yonglin Chen , Naser Al Madi , Can Liu

Mutual gaze detection, i.e., predicting whether or not two people are looking at each other, plays an important role in understanding human interactions. In this work, we focus on the task of image-based mutual gaze detection, and propose a…

Computer Vision and Pattern Recognition · Computer Science 2020-12-23 Bardia Doosti , Ching-Hui Chen , Raviteja Vemulapalli , Xuhui Jia , Yukun Zhu , Bradley Green

In this paper, we propose a Guided Attention (GA) auxiliary training loss, which improves the effectiveness and robustness of automatic speech recognition (ASR) contextual biasing without introducing additional parameters. A common…

Computation and Language · Computer Science 2024-01-18 Jiyang Tang , Kwangyoun Kim , Suwon Shon , Felix Wu , Prashant Sridhar , Shinji Watanabe

Precisely detecting which object a person is paying attention to is critical for human-robot interaction since it provides important cues for the next action from the human user. We propose an end-to-end approach for gaze target detection:…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Zhi-Yi Lin , Jouh Yeong Chew , Jan van Gemert , Xucong Zhang

Neonatal resuscitations demand an exceptional level of attentiveness from providers, who must process multiple streams of information simultaneously. Gaze strongly influences decision making; thus, understanding where a provider is looking…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Felipe Parodi , Jordan Matelsky , Alejandra Regla-Vargas , Elizabeth Foglia , Charis Lim , Danielle Weinberg , Konrad Kording , Heidi Herrick , Michael Platt
‹ Prev 1 4 5 6 7 8 10 Next ›