English
Related papers

Related papers: Inferring Facing Direction from Voice Signals

200 papers

Smart Home Assistants (SHAs) have become ubiquitous in modern households, offering convenience and efficiency through its voice interface. However, for Deaf and Hard-of-Hearing (DHH) individuals, the reliance on auditory and textual…

Human-Computer Interaction · Computer Science 2024-12-03 Tyrone Justin Sta. Maria , Jordan Aiko Deja

The adoption of modern technologies for use in healthcare has become an inevitable change. The emergence of artificial intelligence drives this digital disruption. Artificial intelligence has augmented machine capabilities to act like and…

Human-Computer Interaction · Computer Science 2023-05-23 Akiri Surely

When human listeners try to guess the spatial position of a speech source, they are influenced by the speaker's production level, regardless of the intensity level reaching their ears. Because the perception of distance is a very difficult…

Robotics · Computer Science 2023-05-23 Ambre Davat , Véronique Aubergé , Gang Feng

Today, technological advancement is increasing day by day. Earlier, there was only a computer system in which we could only perform a few tasks. But now, machine learning, artificial intelligence, deep learning, and a few more technologies…

Human-Computer Interaction · Computer Science 2023-05-30 Sumit Kumar , Varun Gupta , Sankalp Sagar , Sachin Kumar Singh

A room's acoustic properties are a product of the room's geometry, the objects within the room, and their specific positions. A room's acoustic properties can be characterized by its impulse response (RIR) between a source and listener…

Sound · Computer Science 2024-01-17 Mason Wang , Samuel Clarke , Jui-Hsien Wang , Ruohan Gao , Jiajun Wu

As we interact with the world, for example when we communicate with our colleagues in a large open space or meeting room, we continuously analyse the surrounding environment and, in particular, localise and recognise acoustic events. While…

Sound · Computer Science 2019-04-02 Pawel Swietojanski , Ondrej Miksik

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech LLMs to comprehend and process multi-channel audio with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-19 Jiamin Xie , Ju Lin , Yiteng Huang , Tyler Vuong , Zhaojiang Lin , Zhaojun Yang , Peng Su , Prashant Rawat , Sangeeta Srivastava , Ming Sun , Florian Metze

Agent assistance during human-human customer support spoken interactions requires triggering workflows based on the caller's intent (reason for call). Timeliness of prediction is essential for a good user experience. The goal is for a…

Artificial Intelligence · Computer Science 2022-08-16 Mrinal Rawat , Victor Barres

As humans, we hear sound every second of our life. The sound we hear is often affected by the acoustics of the environment surrounding us. For example, a spacious hall leads to more reverberation. Room Impulse Responses (RIR) are commonly…

Artificial Intelligence · Computer Science 2023-10-10 Yinfeng Yu , Changan Chen , Lele Cao , Fangkai Yang , Fuchun Sun

In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical…

Human-Computer Interaction · Computer Science 2025-08-05 Kazushi Kato , Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara

In video conferencing, human faces serve as the primary visual focal points, playing multifaceted roles that enhance visual communication and emotional connection. However, we argue that a human face is also a side channel, which can…

Cryptography and Security · Computer Science 2026-04-09 Yong Huang , Yanzhao Lu , Mingyang Chen , En Zhang , Jiazi Li , Wanqing Tu

Recent neural network based Direction of Arrival (DoA) estimation algorithms have performed well on unknown number of sound sources scenarios. These algorithms are usually achieved by mapping the multi-channel audio input to the single…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-18 Haoran Yin , Meng Ge , Yanjie Fu , Gaoyan Zhang , Longbiao Wang , Lei Zhang , Lin Qiu , Jianwu Dang

In the last years several solutions were proposed to support people with visual impairments or blindness during road crossing. These solutions focus on computer vision techniques for recognizing pedestrian crosswalks and computing their…

Human-Computer Interaction · Computer Science 2015-06-25 Sergio Mascetti , Lorenzo Picinali , Andrea Gerino , Dragan Ahmetovic , Cristian Bernareggi

This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Aswin Shanmugam Subramanian , Chao Weng , Shinji Watanabe , Meng Yu , Yong Xu , Shi-Xiong Zhang , Dong Yu

Speech-to-text capabilities on mobile devices have proven helpful for hearing and speech accessibility, language translation, note-taking, and meeting transcripts. However, our foundational large-scale survey (n=263) shows that the…

Human-Computer Interaction · Computer Science 2025-03-06 Artem Dementyev , Dimitri Kanevsky , Samuel J. Yang , Mathieu Parvaix , Chiong Lai , Alex Olwal

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Adaptive test-time compute for LLM agents aims to invoke extra computation only when it improves performance. Existing methods typically use confidence-, uncertainty-, or difficulty-based gates, assuming a fixed direction from the gating…

Machine Learning · Computer Science 2026-05-11 Ziming Li , Jiatan Huang , Xiaoguang Guo , Guilin Wang , Chuxu Zhang

User identification plays a pivotal role in how we interact with our mobile devices. Many existing authentication approaches require active input from the user or specialized sensing hardware, and studies on mobile device usage show…

Human-Computer Interaction · Computer Science 2020-04-13 Yilin Yang , Chen Wang , Yingying Chen , Yan Wang

A new ESPRIT-based algorithm is proposed to estimate the direction-of-arrival of an arbitrary degree polynomial-phase signal with a single acoustic vector sensor. The proposed approach requires neither a priori knowledge of the…

Applications · Statistics 2013-08-06 Xin Yuan

An indoor localization approach uses Wi-Fi Access Points (APs) to estimate the Direction of Arrival (DoA) of the WiFi signals. This paper demonstrates FIND, a tool for Fine INDoor localization based on a software-defined radio, which…

Networking and Internet Architecture · Computer Science 2021-03-10 Evgeny Khorov , Aleksey Kureev , Vladislav Molodtsov