中文
相关论文

相关论文: A Multi-stage Low-latency Enhancement System for H…

200 篇论文

Machine Listening, as usually formalized, attempts to perform a task that is, from our perspective, fundamentally human-performable, and performed by humans. Current automated models of Machine Listening vary from purely data-driven…

音频与语音处理 · 电气工程与系统科学 2023-02-27 Laurie M. Heller , Benjamin Elizalde , Bhiksha Raj , Soham Deshmukh

This paper presents an end-to-end text-to-speech system with low latency on a CPU, suitable for real-time applications. The system is composed of an autoregressive attention-based sequence-to-sequence acoustic model and the LPCNet vocoder…

The performance of voice-controlled systems is usually influenced by accented speech. To make these systems more robust, the frontend accent recognition (AR) technologies have received increased attention in recent years. As accent is a…

音频与语音处理 · 电气工程与系统科学 2021-05-06 Zhan Zhang , Xi Chen , Yuehai Wang , Jianyi Yang

Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural…

While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training…

Signal-dependent beamformers are advantageous over signal-independent beamformers when the acoustic scenario - be it real-world or simulated - is straightforward in terms of the number of sound sources, the ambient sound field and their…

音频与语音处理 · 电气工程与系统科学 2023-12-01 Sina Hafezi , Alastair H. Moore , Pierre H. Guiraud , Patrick A. Naylor , Jacob Donley , Vladimir Tourbabin , Thomas Lunner

Automatic speech recognition enables a wide range of current and emerging applications such as automatic transcription, multimedia content analysis, and natural human-computer interfaces. This paper provides a glimpse of the opportunities…

计算与语言 · 计算机科学 2013-05-14 Rashmi Makhijani , Urmila Shrawankar , V M Thakare

Single-channel speech enhancement approaches do not always improve automatic recognition rates in the presence of noise, because they can introduce distortions unhelpful for recognition. Following a trend towards end-to-end training of…

声音 · 计算机科学 2021-12-14 Peter Plantinga , Deblin Bagchi , Eric Fosler-Lussier

This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2024) for Irish-to-English speech translation. We built end-to-end systems based on Whisper, and employed a number of data…

计算与语言 · 计算机科学 2024-06-28 Yasmin Moslem

The URGENT 2024 Challenge aims to foster speech enhancement (SE) techniques with great universality, robustness, and generalizability, featuring a broader task definition, large-scale multi-domain data, and comprehensive evaluation metrics.…

In real scenarios, it is often necessary and significant to control the inference speed of speech enhancement systems under different conditions. To this end, we propose a stage-wise adaptive inference approach with early exit mechanism for…

声音 · 计算机科学 2021-06-23 Andong Li , Chengshi Zheng , Lu Zhang , Xiaodong Li

The integration of artificial intelligence into hearing assistance marks a paradigm shift from traditional amplification-based systems to intelligent, context-aware audio processing. This systematic literature review evaluates advances in…

声音 · 计算机科学 2025-08-05 Haris Khan , Shumaila Asif , Hassan Nasir , Kamran Aziz Bhatti , Shahzad Amin Sheikh

ADD in real-world scenarios has evolved from speech-only spoofing to more challenging component-level settings, where speech and environmental sounds may be independently manipulated. To tackle this, we propose EnvTriCascade, an…

声音 · 计算机科学 2026-05-19 Hengyan Huang , Xiaoxuan Guo , Jiayi Zhou , Yuankun Xie , Jian Liu , Haonan Cheng , Long Ye , Qin Zhang

As the cornerstone of other important technologies, such as speech recognition and speech synthesis, speech enhancement is a critical area in audio signal processing. In this paper, a new deep learning structure for speech enhancement is…

声音 · 计算机科学 2021-08-30 Yuzi Yan , Wei-Qiang Zhang , Michael T. Johnson

We present a system for non-intrusive prediction of speech quality in noisy and enhanced speech, developed for Track 3 of the VoiceMOS 2024 Challenge. The task required estimating the ITU-T P.835 metrics SIG, BAK, and OVRL without reference…

音频与语音处理 · 电气工程与系统科学 2026-04-28 Marie Kunešová , Aleš Pražák , Jan Lehečka

This paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable…

音频与语音处理 · 电气工程与系统科学 2025-02-14 Gongping Huang , Jesper R. Jensen , Jingdong Chen , Jacob Benesty , Mads G. Christensen , Akihiko Sugiyama , Gary Elko , Tomas Gaensler

We present a work on low-complexity acoustic scene classification (ASC) with multiple devices, namely the subtask A of Task 1 of the DCASE2021 challenge. This subtask focuses on classifying audio samples of multiple devices with a…

音频与语音处理 · 电气工程与系统科学 2023-06-06 Yanxiong Li , Wenchang Cao , Wei Xie , Qisheng Huang , Wenfeng Pang , Qianhua He

The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the…

音频与语音处理 · 电气工程与系统科学 2022-12-16 Eric Guizzo , Christian Marinoni , Marco Pennese , Xinlei Ren , Xiguang Zheng , Chen Zhang , Bruno Masiero , Aurelio Uncini , Danilo Comminiello

With an increasing demand for assistive technologies that promote the independence and mobility of visually impaired people, this study suggests an innovative real-time system that gives audio descriptions of a user's surroundings to…

人机交互 · 计算机科学 2025-03-26 Kunal Chavan , Keertan Balaji , Spoorti Barigidad , Samba Raju Chiluveru

In this letter, we derive a new super Gaussian Joint Maximum a Posteriori based single microphone speech enhancement gain function. The developed Speech Enhancement method is implemented on a smartphone, and this arrangement functions as an…

声音 · 计算机科学 2019-07-04 Chandan K A Reddy , Nikhil Shankar , Gautam Bhat , Ram Charan , Issa Panahi