English
Related papers

Related papers: Interactive spatial speech recognition maps based …

200 papers

Natural language processing (NLP) models trained on people-generated data can be unreliable because, without any constraints, they can learn from spurious correlations that are not relevant to the task. We hypothesize that enriching models…

Computation and Language · Computer Science 2022-03-18 Alissa Ostapenko , Shuly Wintner , Melinda Fricke , Yulia Tsvetkov

Speech and voice conditions can alter the acoustic properties of speech, which could impact the performance of paralinguistic models for affect for people with atypical speech. We evaluate publicly available models for recognizing…

Machine Learning · Computer Science 2025-08-01 Jaya Narain , Amrit Romana , Vikramjit Mitra , Colin Lea , Shirley Ren

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Determining the head orientation of a talker is not only beneficial for various speech signal processing applications, such as source localization or speech enhancement, but also facilitates intuitive voice control and interaction with…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-10 Kaspar Müller , Bilgesu Çakmak , Paul Didier , Simon Doclo , Jan Østergaard , Tobias Wolff

Humanness is core to speech interface design. Yet little is known about how users conceptualise perceptions of humanness and how people define their interaction with speech interfaces through this. To map these perceptions n=21 participants…

Human-Computer Interaction · Computer Science 2019-07-30 Philip R. Doyle , Justin Edwards , Odile Dumbleton , Leigh Clark , Benjamin R. Cowan

In the era of digitalization, as individuals increasingly rely on digital platforms for communication and news consumption, various actors employ linguistic strategies to influence public perception. While models have become proficient at…

Computation and Language · Computer Science 2025-06-18 Sina Abdidizaji , Md Kowsher , Niloofar Yousefi , Ivan Garibay

One of the major problems in modeling natural signals is that signals with very similar structure may locally have completely different measurements, e.g., images taken under different illumination conditions, or the speech signal captured…

Computer Vision and Pattern Recognition · Computer Science 2012-07-19 Nebojsa Jojic , Yaron Caspi , Manuel Reyes-Gomez

Humans use audio signals in the form of spoken language or verbal reactions effectively when teaching new skills or tasks to other humans. While demonstrations allow humans to teach robots in a natural way, learning from trajectories alone…

Robotics · Computer Science 2022-11-02 Akanksha Saran , Kush Desai , Mai Lee Chang , Rudolf Lioutikov , Andrea Thomaz , Scott Niekum

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Speech AI Technologies are largely trained on publicly available datasets or by the massive web-crawling of speech. In both cases, data acquisition focuses on minimizing collection effort, without necessarily taking the data subjects'…

Computers and Society · Computer Science 2023-05-04 Orestis Papakyriakopoulos , Alice Xiang

In this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-13 Raghavendra Pappagari , Tianzi Wang , Jesus Villalba , Nanxin Chen , Najim Dehak

Speech and language technologies offer valuable opportunities for supporting mental health assessment through objective and interpretable cues. We present a systematic feature-based analysis framework leveraging perceptually grounded…

Artificial Intelligence · Computer Science 2026-05-28 Vassilis Lyberatos , Edmund G. Dervakos , Eleni Adamidi , Athanasios Voulodimos , Giorgos Stamou

Voice-based systems like Amazon Alexa, Google Assistant, and Apple Siri, along with the growing popularity of OpenAI's ChatGPT and Microsoft's Copilot, serve diverse populations, including visually impaired and low-literacy communities.…

Human-Computer Interaction · Computer Science 2024-09-04 Sachin Pathiyan Cherumanal , Falk Scholer , Johanne R. Trippas , Damiano Spina

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Xiulong Liu , Anurag Kumar , Paul Calamia , Sebastia V. Amengual , Calvin Murdock , Ishwarya Ananthabhotla , Philip Robinson , Eli Shlizerman , Vamsi Krishna Ithapu , Ruohan Gao

Acoustics-to-word models are end-to-end speech recognizers that use words as targets without relying on pronunciation dictionaries or graphemes. These models are notoriously difficult to train due to the lack of linguistic knowledge. It is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-14 Hao Tang , James Glass

Eliminating the negative effect of non-stationary environmental noise is a long-standing research topic for automatic speech recognition that stills remains an important challenge. Data-driven supervised approaches, including ones based on…

Self-supervised learning enables the training of large neural models without the need for large, labeled datasets. It has been generating breakthroughs in several fields, including computer vision, natural language processing, biology, and…

Computation and Language · Computer Science 2023-12-19 Luis Lugo , Valentin Vielzeuf

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio…

Sound · Computer Science 2019-05-15 Yu-Ding Lu , Hsin-Ying Lee , Hung-Yu Tseng , Ming-Hsuan Yang

Prior research indicates that users prefer assistive technologies whose personalities align with their own. This has sparked interest in automatic personality perception (APP), which aims to predict an individual's perceived personality…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-28 Alice Zhang , Skanda Muralidhar , Daniel Gatica-Perez , Mathew Magimai-Doss

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that the overlapping…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Changil Kim , Hijung Valentina Shin , Tae-Hyun Oh , Alexandre Kaspar , Mohamed Elgharib , Wojciech Matusik
‹ Prev 1 8 9 10 Next ›