English
Related papers

Related papers: EasyCom: An Augmented Reality Dataset to Support A…

200 papers

The ability to interpret social cues comes naturally for most people, but for those living with Autism Spectrum Disorder (ASD), some experience a deficiency in this area. This paper presents the development of a multimodal augmented reality…

Computer Vision and Pattern Recognition · Computer Science 2020-10-23 James Ren Hou Lee , Alexander Wong

In order for AI to be safely deployed in real-world scenarios such as hospitals, schools, and the workplace, it must be able to robustly reason about the physical world. Fundamental to this reasoning is physical common sense: understanding…

Machine Learning · Computer Science 2022-08-02 Samuel Yu , Peter Wu , Paul Pu Liang , Ruslan Salakhutdinov , Louis-Philippe Morency

Unlike the free exploration of childhood, the demands of daily life reduce our motivation to explore our surroundings, leading to missed opportunities for informal learning. Traditional tools for knowledge acquisition are reactive, relying…

Human-Computer Interaction · Computer Science 2025-02-25 Runze Cai , Nuwan Janaka , Hyeongcheol Kim , Yang Chen , Shengdong Zhao , Yun Huang , David Hsu

Augmented reality technology has emerged as a promising solution to assist with wayfinding difficulties, bridging the gap between obtaining navigational assistance and maintaining an awareness of one's real-world surroundings. This article…

Human-Computer Interaction · Computer Science 2023-11-21 Zhiwen Qiu , Armin Mostafavi , Saleh Kalantari

Augmented-reality (AR) glasses that will have access to onboard sensors and an ability to display relevant information to the user present an opportunity to provide user assistance in quotidian tasks. Many such tasks can be characterized as…

Human-Computer Interaction · Computer Science 2020-10-16 Benjamin Newman , Kevin Carlberg , Ruta Desai

All-day smart glasses are likely to emerge as platforms capable of continuous contextual sensing, uniquely positioning them for unprecedented assistance in our daily lives. Integrating the multi-modal AI agents required for human memory…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Akshay Paruchuri , Sinan Hersek , Lavisha Aggarwal , Qiao Yang , Xin Liu , Achin Kulshrestha , Andrea Colaco , Henry Fuchs , Ishan Chatterjee

The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded…

Sound · Computer Science 2024-10-02 Bing Yang , Changsheng Quan , Yabo Wang , Pengyu Wang , Yujie Yang , Ying Fang , Nian Shao , Hui Bu , Xin Xu , Xiaofei Li

We present RealityTalk, a system that augments real-time live presentations with speech-driven interactive virtual elements. Augmented presentations leverage embedded visuals and animation for engaging and expressive storytelling. However,…

Human-Computer Interaction · Computer Science 2022-08-15 Jian Liao , Adnan Karim , Shivesh Jadon , Rubaiat Habib Kazi , Ryo Suzuki

Lack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages. Although various data augmentation approaches have been proposed to synthesize training data in low-resource…

Computation and Language · Computer Science 2021-09-06 Yingmei Guo , Linjun Shou , Jian Pei , Ming Gong , Mingxing Xu , Zhiyong Wu , Daxin Jiang

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

This paper introduces the concept of augmented conversation, which aims to support co-located in-person conversations via embedded speech-driven on-the-fly referencing in augmented reality (AR). Today computing technologies like smartphones…

Human-Computer Interaction · Computer Science 2024-05-30 Shivesh Jadon , Mehrad Faridan , Edward Mah , Rajan Vaish , Wesley Willett , Ryo Suzuki

The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were a multilingual speaker. To create this experience, speech…

Computation and Language · Computer Science 2026-05-19 Supriti Sinhamahapatra , Thai-Binh Nguyen , Yiğit Oğuz , Enes Ugan , Jan Niehues , Alexander Waibel

Event camera has significant advantages in capturing dynamic scene information while being prone to noise interference, particularly in challenging conditions like low threshold and low illumination. However, most existing research focuses…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Yuxing Duan , Shihan Peng , Lin Zhu , Wei Zhang , Yi Chang , Sheng Zhong , Luxin Yan

The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-19 Yufeng Yang , Desh Raj , Ju Lin , Niko Moritz , Junteng Jia , Gil Keren , Egor Lakomkin , Yiteng Huang , Jacob Donley , Jay Mahadeokar , Ozlem Kalinli

Efficient face detection is critical to provide natural human-robot interactions. However, computer vision tends to involve a large computational load due to the amount of data (i.e. pixels) that needs to be processed in a short amount of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-19 William Aris , François Grondin

This paper proposes a flexible multichannel speech enhancement system with the main goal of improving robustness of automatic speech recognition (ASR) in noisy conditions. The proposed system combines a flexible neural mask estimator…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Ante Jukić , Jagadeesh Balam , Boris Ginsburg

This paper introduces an augmented reality (AR) captioning framework designed to support Deaf and Hard of Hearing (DHH) learners in STEM classrooms by integrating non-verbal emotional cues into live transcriptions. Unlike conventional…

Human-Computer Interaction · Computer Science 2025-04-29 Sunday David Ubur

Stress during public speaking is common and adversely affects performance and self-confidence. Extensive research has been carried out to develop various models to recognize emotional states. However, minimal research has been conducted to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-03 Arushi , Roberto Dillon , Ai Ni Teoh , Denise Dillon

The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning. This paper introduces the…

Sound · Computer Science 2018-03-29 Jon Barker , Shinji Watanabe , Emmanuel Vincent , Jan Trmal

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges…

‹ Prev 1 4 5 6 7 8 10 Next ›