English
Related papers

Related papers: Single Microphone Own Voice Detection based on Sim…

200 papers

We investigate the objective performance of five high-end commercially available Hearing Aid (HA) devices compared to DNN-based speech enhancement algorithms in complex acoustic environments. To this end, we measure the HRTFs of a single HA…

Sound · Computer Science 2023-07-25 Enric Gusó , Joanna Luberadzka , Martí Baig , Umut Sayin Saraç , Xavier Serra

A new database of head-related transfer functions (HRTFs) for accurate sound source localization is presented through precise measurement and post-processing in terms of improved frequency bandwidth and causality of head-related impulse…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-07 Gyeong-Tae Lee , Sang-Min Choi , Byeong-Yun Ko , Yong-Hwa Park

Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increasing demand for personalized and context-aware technologies,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Satyam Kumar , Sai Srujana Buddi , Utkarsh Oggy Sarawgi , Vineet Garg , Shivesh Ranjan , Ognjen , Rudovic , Ahmed Hussen Abdelaziz , Saurabh Adya

Adapters are an efficient, composable alternative to full fine-tuning of pre-trained models and help scale the deployment of large ASR models to many tasks. In practice, a task ID is commonly prepended to the input during inference to route…

Computation and Language · Computer Science 2023-10-23 Hillary Ngai , Rohan Agrawal , Neeraj Gaur , Ronny Huang , Parisa Haghani , Pedro Moreno Mengibar

Audio Deepfake Detection (ADD) aims to detect the fake audio generated by text-to-speech (TTS), voice conversion (VC) and replay, etc., which is an emerging topic. Traditionally we take the mono signal as input and focus on robust feature…

Sound · Computer Science 2023-05-29 Rui Liu , Jinhua Zhang , Guanglai Gao , Haizhou Li

Due to the successful application of deep learning, audio spoofing detection has made significant progress. Spoofed audio with speech synthesis or voice conversion can be well detected by many countermeasures. However, an automatic speaker…

Sound · Computer Science 2024-01-12 Lian Huang , Chi-Man Pun

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA)…

Sound · Computer Science 2024-03-22 Samuel Pegg , Kai Li , Xiaolin Hu

Accurate sound propagation simulation is essential for delivering immersive experiences in virtual applications, yet industry methods for acoustic modeling often do not account for the full breadth of acoustic wave phenomena. This paper…

Sound · Computer Science 2025-07-15 Bilkent Samsurya

Automated detection of voice disorders with computational methods is a recent research area in the medical domain since it requires a rigorous endoscopy for the accurate diagnosis. Efficient screening methods are required for the diagnosis…

Quantitative Methods · Quantitative Biology 2018-12-06 Vibhuti Gupta

Orientation and mobility (O&M) instruction for blind and low-vision learners is effective but difficult to standardize and repeat at scale due to the reliance on instructor availability, physical mock-ups, and variable real-world outdoor…

Human-Computer Interaction · Computer Science 2026-02-24 Daniel A. Muñoz

Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE methods rely exclusively on frontal-view videos, this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-12 Peijun Yang , Zhan Jin , Juan Liu , Ming Li

Determining the head orientation of a talker is not only beneficial for various speech signal processing applications, such as source localization or speech enhancement, but also facilitates intuitive voice control and interaction with…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-10 Kaspar Müller , Bilgesu Çakmak , Paul Didier , Simon Doclo , Jan Østergaard , Tobias Wolff

Audio-visual automatic speech recognition (AV-ASR) models are very effective at reducing word error rates on noisy speech, but require large amounts of transcribed AV training data. Recently, audio-visual self-supervised learning (SSL)…

Sound · Computer Science 2023-12-18 Avner May , Dmitriy Serdyuk , Ankit Parag Shah , Otavio Braga , Olivier Siohan

In this paper, we propose an enhanced audio-visual deep detection method. Recent methods in audio-visual deepfake detection mostly assess the synchronization between audio and visual features. Although they have shown promising results,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Marcella Astrid , Enjie Ghorbel , Djamila Aouada

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

This paper evaluates the effectiveness of a Cycle-GAN based voice converter (VC) on four speaker identification (SID) systems and an automated speech recognition (ASR) system for various purposes. Audio samples converted by the VC model are…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-30 Gokce Keskin , Tyler Lee , Cory Stephenson , Oguz H. Elibol

Traditionally, Blind Speech Separation techniques are computationally expensive as they update the demixing matrix at every time frame index, making them impractical to use in many Real-Time applications. In this paper, a robust data-driven…

Sound · Computer Science 2018-12-11 Chandan K A Reddy , Gautam Bhat , Nikhil Shankar , Issa Panahi

The Voice Conversion Challenge 2020 is the third edition under its flagship that promotes intra-lingual semiparallel and cross-lingual voice conversion (VC). While the primary evaluation of the challenge submissions was done through…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-09 Rohan Kumar Das , Tomi Kinnunen , Wen-Chin Huang , Zhenhua Ling , Junichi Yamagishi , Yi Zhao , Xiaohai Tian , Tomoki Toda

Voice activity detection (VAD) remains a challenge in noisy environments. With access to multiple microphones, prior studies have attempted to improve the noise robustness of VAD by creating multi-channel VAD (MVAD) methods. However, MVAD…

We propose a data-driven sparse recovery framework for hybrid spherical linear microphone arrays using singular value decomposition (SVD) of the transfer operator. The SVD yields orthogonal microphone and field modes, reducing to spherical…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Shunxi Xu , Thushara Abhayapala , Craig T. Jin
‹ Prev 1 4 5 6 7 8 10 Next ›