中文
相关论文

相关论文: A Multi-stage Low-latency Enhancement System for H…

200 篇论文

This paper describes a new baseline system for automatic speech recognition (ASR) in the CHiME-4 challenge to promote the development of noisy ASR in speech processing communities by providing 1) state-of-the-art system with a simplified…

声音 · 计算机科学 2018-03-28 Szu-Jui Chen , Aswin Shanmugam Subramanian , Hainan Xu , Shinji Watanabe

In recent years, the introduction of neural networks (NNs) into the field of speech enhancement has brought significant improvements. However, many of the proposed methods are quite demanding in terms of computational complexity and memory…

音频与语音处理 · 电气工程与系统科学 2024-04-19 Ernst Seidel , Pejman Mowlaee , Tim Fingscheidt

In this paper, we design a first of its kind transceiver (PHY layer) prototype for cloud-based audio-visual (AV) speech enhancement (SE) complying with high data rate and low latency requirements of future multimodal hearing assistive…

音频与语音处理 · 电气工程与系统科学 2022-10-25 Abhijeet Bishnu , Ankit Gupta , Mandar Gogate , Kia Dashtipour , Ahsan Adeel , Amir Hussain , Mathini Sellathurai , Tharmalingam Ratnarajah

A central problem in building effective sound event detection systems is the lack of high-quality, strongly annotated sound event datasets. For this reason, Task 4 of the DCASE 2024 challenge proposes learning from two heterogeneous…

音频与语音处理 · 电气工程与系统科学 2024-07-19 Florian Schmid , Paul Primus , Tobias Morocutti , Jonathan Greif , Gerhard Widmer

Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates the actual audio. High-quality spectrum models usually…

声音 · 计算机科学 2021-04-05 Qing He , Zhiping Xiu , Thilo Koehler , Jilong Wu

This paper presents the speech restoration and enhancement system created by the 1024K team for the ICASSP 2024 Speech Signal Improvement (SSI) Challenge. Our system consists of a generative adversarial network (GAN) in complex-domain for…

声音 · 计算机科学 2024-02-06 Guochen Yu , Runqiang Han , Chenglin Xu , Haoran Zhao , Nan Li , Chen Zhang , Xiguang Zheng , Chao Zhou , Qi Huang , Bing Yu

Verbatim transcription for automatic speaking assessment demands accurate capture of disfluencies, crucial for downstream tasks like error analysis and feedback. However, many ASR systems discard or generalize hesitations, losing important…

计算与语言 · 计算机科学 2025-07-28 Jhen-Ke Lin , Hao-Chien Lu , Chung-Chun Wang , Hong-Yun Lin , Berlin Chen

Full-duplex speech interaction, as the most natural and intuitive mode of human communication, is driving artificial intelligence toward more human-like conversational systems. Traditional cascaded speech processing pipelines suffer from…

人工智能 · 计算机科学 2026-05-01 Yadong Li , Guoxin Wu , Haiping Hou , Biye Li

This paper presents our MSXF TTS system for Task 3.1 of the Audio Deep Synthesis Detection (ADD) Challenge 2022. We use an end to end text to speech system, and add a constraint loss to the system when training stage. The end to end TTS…

声音 · 计算机科学 2022-01-28 Chunyong Yang , Pengfei Liu , Yanli Chen , Hongbin Wang , Min Liu

The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder…

声音 · 计算机科学 2017-03-16 Tsubasa Ochiai , Shinji Watanabe , Takaaki Hori , John R. Hershey

This paper presents the details of Task 1A Acoustic Scene Classification in the DCASE 2021 Challenge. The task targeted development of low-complexity solutions with good generalization properties. The provided baseline system is based on a…

音频与语音处理 · 电气工程与系统科学 2021-07-21 Irene Martín-Morató , Toni Heittola , Annamaria Mesaros , Tuomas Virtanen

The primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing, with a specific emphasis on 3D speech enhancement and 3D…

音频与语音处理 · 电气工程与系统科学 2024-02-15 Christian Marinoni , Riccardo Fosco Gramaccioni , Changan Chen , Aurelio Uncini , Danilo Comminiello

Recent years have seen an increasing number of studies around the design of computer-assisted interpreting tools with integrated automatic speech processing and their use by trainees and professional interpreters. This paper discusses the…

计算与语言 · 计算机科学 2022-01-11 Claudio Fantinuoli , Maddalena Montecchio

This paper describes a real-time General Speech Reconstruction (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This novel proposed system is a two-stage architecture, in which the speech restoration…

声音 · 计算机科学 2023-06-16 Wenzhe Liu , Yupeng Shi , Jun Chen , Wei Rao , Shulin He , Andong Li , Yannan Wang , Zhiyong Wu

We propose a novel formulation for phase synchronization -- the statistical problem of jointly estimating alignment angles from noisy pairwise comparisons -- as a nonconvex optimization problem that enforces consistency among the pairwise…

信息论 · 计算机科学 2019-05-15 Tingran Gao , Zhizhen Zhao

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Chenda Li , Wei Wang , Marvin Sach , Wangyou Zhang , Kohei Saijo , Samuele Cornell , Yihui Fu , Zhaoheng Ni , Tim Fingscheidt , Shinji Watanabe , Yanmin Qian

Usually, hearing impaired people use hearing aids which are implemented with speech enhancement algorithms. Estimation of speech and estimation of nose are the components in single channel speech enhancement system. The main objective of…

声音 · 计算机科学 2014-11-10 M. Ravichandra Kumar , B. Ravi Teja

Modern neural speech enhancement models usually include various forms of phase information in their training loss terms, either explicitly or implicitly. However, these loss terms are typically designed to reduce the distortion of phase…

声音 · 计算机科学 2022-02-25 Doyeon Kim , Hyewon Han , Hyeon-Kyeong Shin , Soo-Whan Chung , Hong-Goo Kang

Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Dejun Li

The conventional wisdom has been that designing ultra-compact, battery-constrained wireless hearables with on-device speech AI models is challenging due to the high computational demands of streaming deep learning models. Speech AI models…

声音 · 计算机科学 2025-10-23 Malek Itani , Tuochao Chen , Arun Raghavan , Gavriel Kohlberg , Shyamnath Gollakota