English
Related papers

Related papers: ANGUS: Real-time manipulation of vocal roughness f…

200 papers

Deciphering the acoustic language of chickens offers new opportunities in animal welfare and ecological informatics. Their subtle vocal signals encode health conditions, emotional states, and dynamic interactions within ecosystems.…

Sound · Computer Science 2024-12-24 Venkatraman Manikandan , Suresh Neethirajan

Intro: Vocal cord ultrasound (VCUS) has emerged as a less invasive and better tolerated examination technique, but its accuracy is operator dependent. This research aims to apply a machine learning-assisted algorithm to automatically…

Machine Learning · Computer Science 2025-12-30 Will Sebelik-Lassiter , Evan Schubert , Muhammad Alliyu , Quentin Robbins , Excel Olatunji , Mustafa Barry

In this dissertation the practical speech emotion recognition technology is studied, including several cognitive related emotion types, namely fidgetiness, confidence and tiredness. The high quality of naturalistic emotional speech data is…

Sound · Computer Science 2017-09-28 Chengwei Huang

Active noise control (ANC) has become popular for reducing noise and thus enhancing user comfort in headphones. While feedback control offers an effective way to implement ANC, it is restricted by uncertainty of the controlled system that…

Systems and Control · Electrical Eng. & Systems 2025-09-22 Florian Hilgemann , Egke Chatzimoustafa , Peter Jax

Speech recognition in noisy and channel distorted scenarios is often challenging as the current acoustic modeling schemes are not adaptive to the changes in the signal distribution in the presence of noise. In this work, we develop a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Purvi Agrawal , Sriram Ganapathy

Humans are able to fuse information from both auditory and visual modalities to help with understanding speech. This is demonstrated through a phenomenon known as the McGurk Effect, during which a listener is presented with incongruent…

Sound · Computer Science 2025-10-30 Lukas Grasse , Matthew S. Tata

Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-13 Benjamin van Niekerk , Marc-André Carbonneau , Herman Kamper

The automatic recognition of emotion in speech can inform our understanding of language, emotion, and the brain. It also has practical application to human-machine interactive systems. This paper examines the recognition of emotion in…

Speech anonymisation prevents misuse of spoken data by removing any personal identifier while preserving at least linguistic content. However, emotion preservation is crucial for natural human-computer interaction. The well-known voice…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Suhita Ghosh , Arnab Das , Yamini Sinha , Ingo Siegert , Tim Polzehl , Sebastian Stober

Audio coding is an essential module in the real-time communication system. Neural audio codecs can compress audio samples with a low bitrate due to the strong modeling and generative capabilities of deep neural networks. To address the poor…

Sound · Computer Science 2023-10-18 Wenzhe Liu , Wei Xiao , Meng Wang , Shan Yang , Yupeng Shi , Yuyong Kang , Dan Su , Shidong Shang , Dong Yu

Hallucination is an apparent perception in the absence of real external sensory stimuli. An auditory hallucination is a perception of hearing sounds that are not real. A common form of auditory hallucination is hearing voices in the absence…

Sound · Computer Science 2023-04-24 Shayan Mirjafari , Subigya Nepal , Weichen Wang , Andrew T. Campbell

Emotions are one of the important components of the human being, thus they are a valuable part of daily activities such as interaction with people, decision making and learning. For this reason, it is important to detect, recognize and…

Human-Computer Interaction · Computer Science 2025-12-30 Ricardo Vasquez , Diego Riofrío-Luzcando , Joe Carrion-Jumbo , Cesar Guevara

The use of natural language and voice-based interfaces gradu-ally transforms how consumers search, shop, and express their preferences. The current work explores how changes in the syntactical structure of the interaction with…

Artificial Intelligence · Computer Science 2021-11-04 Christian Hildebrand , Donna Hoffman , Tom Novak

With every advancement in generative AI models, forensics is under increasing pressure. The constant emergence of new generation techniques makes it impossible to collect data for each manipulation to train a deepfake detection model. Thus,…

Artificial Intelligence · Computer Science 2026-05-20 Aritra Marik , Marcel Klemt , Anna Rohrbach

Recognizing emotions in spoken communication is crucial for advanced human-machine interaction. Current emotion detection methodologies often display biases when applied cross-corpus. To address this, our study amalgamates 16 diverse…

Computation and Language · Computer Science 2023-11-16 Mohamed Osman , Tamer Nadeem , Ghada Khoriba

Indeed, these are exciting times. We are in the heart of a digital renaissance. Automation and computer technology allow engineers and scientists to fabricate processes that amalgamate quality of life. We anticipate much growth in medical…

Computer Vision and Pattern Recognition · Computer Science 2013-01-16 H. J. Moukalled

Audiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural network…

Computer Vision and Pattern Recognition · Computer Science 2018-05-01 Michael Wand , Ngoc Thang Vu , Juergen Schmidhuber

Traditional audiometry often fails to fully characterize the functional impact of hearing loss on speech understanding, particularly supra-threshold deficits and frequency-specific perception challenges in conditions like presbycusis. This…

Sound · Computer Science 2025-05-29 Stefan Bleeck

This paper presents our method for the estimation of valence-arousal (VA) in the 8th Affective Behavior Analysis in-the-Wild (ABAW) competition. Our approach integrates visual and audio information through a multimodal framework. The visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Jun Yu , Yongqi Wang , Lei Wang , Yang Zheng , Shengfan Xu

Form about four decades human beings have been dreaming of an intelligent machine which can master the natural speech. In its simplest form, this machine should consist of two subsystems, namely automatic speech recognition (ASR) and speech…

Sound · Computer Science 2013-05-08 Urmila Shrawankar , V. M. Thakare