中文
相关论文

相关论文: MAGE: A Coarse-to-Fine Speech Enhancer with Masked…

200 篇论文

In this paper, we propose the coarse-to-fine optimization for the task of speech enhancement. Cosine similarity loss [1] has proven to be an effective metric to measure similarity of speech signals. However, due to the large variance of the…

声音 · 计算机科学 2019-08-23 Jian Yao , Ahmad Al-Dahle

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

声音 · 计算机科学 2025-06-03 Nabarun Goswami , Tatsuya Harada

Generative modeling and representation learning are two key tasks in computer vision. However, these models are typically trained independently, which ignores the potential for each task to help the other, and leads to training and model…

计算机视觉与模式识别 · 计算机科学 2023-07-03 Tianhong Li , Huiwen Chang , Shlok Kumar Mishra , Han Zhang , Dina Katabi , Dilip Krishnan

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Shaojun E , Yuchen Yang , Jiaheng Wu , Yan Zhang , Tiejun Zhao , Ziyan Chen

In challenging environments with significant noise and reverberation, traditional speech enhancement (SE) methods often lead to over-suppressed speech, creating artifacts during listening and harming downstream tasks performance. To…

音频与语音处理 · 电气工程与系统科学 2024-10-03 Hsin-Tien Chiang , Hao Zhang , Yong Xu , Meng Yu , Dong Yu

Multimodal music creation requires models that can both generate audio from high-level cues and edit existing mixtures in a targeted manner. Yet most multimodal music systems are built for a single task and a fixed prompting interface,…

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spectrogram mask for the…

音频与语音处理 · 电气工程与系统科学 2025-05-22 Hao Ma , Rujin Chen , Xiao-Lei Zhang , Ju Liu , Xuelong Li

Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We present GenTSE, a two-stage decoder-only generative LM approach for TSE:…

音频与语音处理 · 电气工程与系统科学 2025-12-25 Haoyang Li , Xuyi Zhuang , Azmat Adnan , Ye Ni , Wei Rao , Shreyas Gopal , Eng Siong Chng

Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-art speech processing models. Although end-to-end approaches…

声音 · 计算机科学 2026-01-27 Kohei Asai , Wataru Nakata , Yuki Saito , Hiroshi Saruwatari

We introduce AnyEnhance, a unified generative model for voice enhancement that processes both speech and singing voices. Based on a masked generative model, AnyEnhance is capable of handling both speech and singing voices, supporting a wide…

声音 · 计算机科学 2025-11-04 Junan Zhang , Jing Yang , Zihao Fang , Yuancheng Wang , Zehua Zhang , Zhuo Wang , Fan Fan , Zhizheng Wu

Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder framework that augments…

声音 · 计算机科学 2026-01-28 Samir Sadok , Stéphane Lathuilière , Xavier Alameda-Pineda

We introduce MAGNeT, a masked generative sequence modeling method that operates directly over several streams of audio tokens. Unlike prior work, MAGNeT is comprised of a single-stage, non-autoregressive transformer. During training, we…

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix…

机器学习 · 计算机科学 2020-02-11 Simon Leglaive , Xavier Alameda-Pineda , Laurent Girin , Radu Horaud

Deep learning has become a de facto method of choice for speech enhancement tasks with significant improvements in speech quality. However, real-time processing with reduced size and computations for low-power edge devices drastically…

音频与语音处理 · 电气工程与系统科学 2024-05-28 Monisankha Pal , Arvind Ramanathan , Ted Wada , Ashutosh Pandey

Existing deep learning based speech enhancement mainly employ a data-driven approach, which leverage large amounts of data with a variety of noise types to achieve noise removal from noisy signal. However, the high dependence on the data…

声音 · 计算机科学 2024-01-24 Huaying Xue , Xiulian Peng , Yan Lu

Masked Autoencoders (MAEs) trained on audio spectrogram patches have emerged as a prominent approach for learning self-supervised audio representations. While several recent papers have evaluated key aspects of training MAEs on audio data,…

声音 · 计算机科学 2025-07-15 Sarthak Yadav , Sergios Theodoridis , Zheng-Hua Tan

Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio…

音频与语音处理 · 电气工程与系统科学 2025-07-15 Shu-wen Yang , Byeonggeun Kim , Kuan-Po Huang , Qingming Tang , Huy Phan , Bo-Ru Lu , Harsha Sundar , Shalini Ghosh , Hung-yi Lee , Chieh-Chi Kao , Chao Wang

Speech enhancement (SE) improves degraded speech's quality, with generative models like flow matching gaining attention for their outstanding perceptual quality. However, the flow-based model requires multiple numbers of function…

音频与语音处理 · 电气工程与系统科学 2025-09-26 Jiahe Wang , Hongyu Wang , Wei Wang , Lei Yang , Chenda Li , Wangyou Zhang , Lufen Tan , Yanmin Qian

Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal prediction techniques. This paper introduces an efficient and…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Xingwei Sun , Heinrich Dinkel , Yadong Niu , Linzhang Wang , Junbo Zhang , Jian Luan

Semantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns and contextual…

音频与语音处理 · 电气工程与系统科学 2025-02-06 Jixun Yao , Hexin Liu , Chen Chen , Yuchen Hu , EngSiong Chng , Lei Xie
‹ 上一页 1 2 3 10 下一页 ›