English
Related papers

Related papers: Estimating Speech Duration by Measuring the Abdomi…

200 papers

Continuous collection of physiological data from wearable sensors enables temporal characterization of individual behaviors. Understanding the relation between an individual's behavioral patterns and psychological states can help identify…

Any audio recording encapsulates the unique fingerprint of the associated acoustic environment, namely the background noise and reverberation. Considering the scenario of a room equipped with a fixed smart speaker device with one or more…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-05 Francesco Nespoli , Daniel Barreda , Patrick A. Naylor

The study of well-being, stress and other human factors has traditionally relied on self-report instruments to assess key variables. However, concerns about potential biases in these instruments, even when thoroughly validated and…

Software Engineering · Computer Science 2025-07-08 Cristina Martinez Montes , Daniela Grassi , Nicole Novielli , Birgit Penzenstadler

Distributed Virtual Reality systems enable globally dispersed users to interact with each other in a shared virtual environment. In such systems, different types of latencies occur. For a good VR experience, they need to be controlled. The…

Graphics · Computer Science 2018-09-18 Armin Becher , Jens Angerer , Thomas Grauschopf

Depression is a global health concern with a critical need for increased patient screening. Speech technology offers advantages for remote screening but must perform robustly across patients. We have described two deep learning models…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Y. Lu , A. Harati , T. Rutowski , R. Oliveira , P. Chlebek , E. Shriberg

Videoconferencing is now a frequent mode of communication in both professional and informal settings, yet it often lacks the fluidity and enjoyment of in-person conversation. This study leverages multimodal machine learning to predict…

Machine Learning · Computer Science 2025-03-11 Andrew Chang , Viswadruth Akkaraju , Ray McFadden Cogliano , David Poeppel , Dustin Freeman

Social interactions play a crucial role in shaping human behavior, relationships, and societies. It encompasses various forms of communication, such as verbal conversation, non-verbal gestures, facial expressions, and body language. In this…

Machine Learning · Computer Science 2026-05-13 Alice Zhang , Callihan Bertley , Dawei Liang , Edison Thomaz

Estimating human vital signs in a contactless non-invasive method using radar provides a convenient method in the medical field to conduct several health checkups easily and quickly. In addition to monitoring while sitting and sleeping, the…

Signal Processing · Electrical Eng. & Systems 2022-06-14 Tassneem Helal , Fady Aziz , Omar Metwally , Marco F. Huber , Dominik Alscher , Christoph Wasser , Urs Schneider

In this paper, we introduce a new task, Reactive Listener Motion Generation from Speaker Utterance, which aims to generate naturalistic listener body motions that appropriately respond to a speaker's utterance. However, modeling such…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Cheng Luo , Bizhu Wu , Bing Li , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen , Bernard Ghanem

Does speaking style variation affect humans' ability to distinguish individuals from their voices? How do humans compare with automatic systems designed to discriminate between voices? In this paper, we attempt to answer these questions by…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Amber Afshan , Jody Kreiman , Abeer Alwan

We propose a Perceiver-based sequence classifier to detect abnormalities in speech reflective of several neurological disorders. We combine this classifier with a Universal Speech Model (USM) that is trained (unsupervised) on 12 million…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-23 Hagen Soltau , Izhak Shafran , Alex Ottenwess , Joseph R. JR Duffy , Rene L. Utianski , Leland R. Barnard , John L. Stricker , Daniela Wiepert , David T. Jones , Hugo Botha

Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram…

Sound · Computer Science 2026-03-26 Qingkun Deng , Saturnino Luz , Sofia de la Fuente Garcia

Synthesizing natural head motion to accompany speech for an embodied conversational agent is necessary for providing a rich interactive experience. Most prior works assess the quality of generated head motion by comparing them against a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-27 Trisha Mittal , Zakaria Aldeneh , Masha Fedzechkina , Anurag Ranjan , Barry-John Theobald

This paper examines the speaker identification potential of breath sounds in continuous speech. Speech is largely produced during exhalation. In order to replenish air in the lungs, speakers must periodically inhale. When inhalation occurs…

Sound · Computer Science 2017-12-05 Wenbo Zhao , Yang Gao , Rita Singh

This position paper describes an experiment conducted to understand the relationships between different physiological measures including pupil Diameter, Blinking Rate, Heart Rate, and Heart Rate Variability in order to develop an estimation…

Human-Computer Interaction · Computer Science 2019-06-26 Ingo Keller , Muneeb Imtiaz Ahmad , Katrin Lohan

Speech separation has been extensively explored to tackle the cocktail party problem. However, these studies are still far from having enough generalization capabilities for real scenarios. In this work, we raise a common strategy named…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-26 Jing Shi , Jiaming Xu , Yusuke Fujita , Shinji Watanabe , Bo Xu

Voice interfaces are increasingly used in high stakes domains such as mobile banking, smart home security, and hands free healthcare. Meanwhile, modern generative models have made high quality voice forgeries inexpensive and easy to create,…

Cryptography and Security · Computer Science 2026-02-23 Ynes Ineza , Muhammad A. Ullah , Abdul Serwadda , Aurore Munyaneza

In most automatic speech recognition (ASR) systems, the audio signal is processed to produce a time series of sensor measurements (e.g., filterbank outputs). This time series encodes semantic information in a speaker-dependent way. An…

Sound · Computer Science 2019-05-10 David N. Levin

In this article, we explore computer vision approaches to detect abnormal head pose during e-learning sessions and we introduce a study on the effects of mobile phone usage during these sessions. We utilize behavioral data collected from…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Álvaro Becerra , Javier Irigoyen , Roberto Daza , Ruth Cobos , Aythami Morales , Julian Fierrez , Mutlu Cukurova

Since the mental states of the speaker modulate speech, stress introduced by cognitive or physical loads could be detected in the voice. The existing voice stress detection benchmark has shown that the audio embeddings extracted from the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-12 Zihan Wu , Neil Scheidwasser-Clow , Karl El Hajal , Milos Cernak