中文
相关论文

相关论文: VocBench: A Neural Vocoder Benchmark for Speech Sy…

200 篇论文

We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks:…

Noise suppression systems generally produce output speech with compromised quality. We propose to utilize the high quality speech generation capability of neural vocoders for noise suppression. We use a neural network to predict clean…

声音 · 计算机科学 2019-11-15 Soumi Maiti , Michael I Mandel

Recent advancements in end-to-end neural speech codecs enable compressing audio at extremely low bitrates while maintaining high-fidelity reconstruction. Meanwhile, low computational complexity and low latency are crucial for real-time…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Leyan Yang , Ronghui Hu , Yang Xu , Jing Lu

While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be…

计算与语言 · 计算机科学 2018-08-08 Yu-Hsuan Wang , Hung-yi Lee , Lin-shan Lee

Conventional vocoders are commonly used as analysis tools to provide interpretable features for downstream tasks such as speech synthesis and voice conversion. They are built under certain assumptions about the signals following signal…

音频与语音处理 · 电气工程与系统科学 2021-10-14 Sergey Nikonorov , Berrak Sisman , Mingyang Zhang , Haizhou Li

Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance…

声音 · 计算机科学 2026-01-22 Mohammed Salah Al-Radhi , Riad Larbi , Mátyás Bartalis , Géza Németh

Evaluating code generation models for 3D spatial reasoning requires executing generated code in realistic environments and assessing outputs beyond surface-level correctness. We introduce a platform VoxelCode, for analyzing code generation…

机器学习 · 计算机科学 2026-04-06 Yan Zheng , Florian Bordes

Vocoders received renewed attention as main components in statistical parametric text-to-speech (TTS) synthesis and speech transformation systems. Even though there are vocoding techniques give almost accepted synthesized speech, their high…

声音 · 计算机科学 2021-06-22 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Géza Németh

Recent advancements in deep learning led to human-level performance in single-speaker speech synthesis. However, there are still limitations in terms of speech quality when generalizing those systems into multiple-speaker models especially…

音频与语音处理 · 电气工程与系统科学 2020-08-13 Dipjyoti Paul , Yannis Pantazis , Yannis Stylianou

This paper presents a refinement framework of WaveNet vocoders for variational autoencoder (VAE) based voice conversion (VC), which reduces the quality distortion caused by the mismatch between the training data and testing data.…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Wen-Chin Huang , Yi-Chiao Wu , Hsin-Te Hwang , Patrick Lumban Tobing , Tomoki Hayashi , Kazuhiro Kobayashi , Tomoki Toda , Yu Tsao , Hsin-Min Wang

This paper presents a novel method for extracting the vocal track from a musical mixture. The musical mixture consists of a singing voice and a backing track which may comprise of various instruments. We use a convolutional network with…

声音 · 计算机科学 2020-02-13 Pritish Chandna , Merlijn Blaauw , Jordi Bonada , Emilia Gomez

Recently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains. In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring…

音频与语音处理 · 电气工程与系统科学 2019-07-18 Jee-weon Jung , Hee-Soo Heo , Ju-ho Kim , Hye-jin Shim , Ha-Jin Yu

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Silin Gao , Sheryl Mathew , Li Mi , Sepideh Mamooler , Mengjie Zhao , Hiromi Wakaki , Yuki Mitsufuji , Syrielle Montariol , Antoine Bosselut

Emotional voice conversion (EVC) is one way to generate expressive synthetic speech. Previous approaches mainly focused on modeling one-to-one mapping, i.e., conversion from one emotional state to another emotional state, with Mel-cepstral…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

Speech enhancement (SE) and neural vocoding are traditionally viewed as separate tasks. In this work, we observe them under a common thread: the rank behavior of these processes. This observation prompts two key questions: \textit{Can a…

声音 · 计算机科学 2025-01-24 Andong Li , Zhihang Sun , Fengyuan Hao , Xiaodong Li , Chengshi Zheng

Neural network (NN) verification aims to formally verify properties of NNs, which is crucial for ensuring the behavior of NN-based models in safety-critical applications. In recent years, the community has developed many NN verifiers and…

机器学习 · 计算机科学 2026-01-01 Xingjian Zhou , Keyi Shen , Andy Xu , Hongji Xu , Cho-Jui Hsieh , Huan Zhang , Zhouxing Shi

Cross-modal associations between voice and face from a person can be learnt algorithmically, which can benefit a lot of applications. The problem can be defined as voice-face matching and retrieval tasks. Much research attention has been…

计算机视觉与模式识别 · 计算机科学 2020-01-01 Chuyuan Xiong , Deyuan Zhang , Tao Liu , Xiaoyong Du

The increasing realism of synthetic speech, driven by advancements in text-to-speech models, raises ethical concerns regarding impersonation and disinformation. Audio watermarking offers a promising solution via embedding…

机器学习 · 计算机科学 2024-11-14 Hongbin Liu , Moyang Guo , Zhengyuan Jiang , Lun Wang , Neil Zhenqiang Gong

We describe a new convolutional framework for waveform evaluation, WEnets, and build a Narrowband Audio Waveform Evaluation Network, or NAWEnet, using this framework. NAWEnet is single-ended (or no-reference) and was trained three separate…

音频与语音处理 · 电气工程与系统科学 2019-09-20 Andrew A. Catellier , Stephen D. Voran

Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research…