English
Related papers

Related papers: Prosody-Guided Harmonic Attention for Phase-Cohere…

200 papers

Autoregressive neural vocoders have achieved outstanding performance in speech synthesis tasks such as text-to-speech and voice conversion. An autoregressive vocoder predicts a sample at some time step conditioned on those at previous time…

Sound · Computer Science 2024-06-06 Po-chun Hsu , Da-rong Liu , Andy T. Liu , Hung-yi Lee

Neural transducers have achieved human level performance on standard speech recognition benchmarks. However, their performance significantly degrades in the presence of cross-talk, especially when the primary speaker has a low…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-31 Desh Raj , Junteng Jia , Jay Mahadeokar , Chunyang Wu , Niko Moritz , Xiaohui Zhang , Ozlem Kalinli

Deploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Heitor R. Guimarães , Ke Tan , Juan Azcarreta , Jesus Alvarez , Prabhav Agrawal , Ashutosh Pandey , Buye Xu

The research presents a voice conversion model using coefficient mapping and neural network. Most previous works on parametric speech synthesis did not account for losses in spectral details causing over smoothing and invariably, an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-12 Olaide Ayodeji Agbolade , Samson A. Oyetunji

Speech brain--computer interfaces require decoders that translate intracortical activity into linguistic output while remaining robust to limited data and day-to-day variability. While prior high-performing systems have largely relied on…

Computation and Language · Computer Science 2026-03-24 Michal Olak , Tommaso Boccato , Matteo Ferrante

In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, speech rate and energy…

Sound · Computer Science 2022-03-03 Pengyu Cheng , Zhenhua Ling

Complex-valued signals encode both amplitude and phase, yet most deep models treat attention as real-valued correlation, overlooking interference effects. We introduce the Holographic Transformer, a physics-inspired architecture that…

Signal Processing · Electrical Eng. & Systems 2025-10-31 Enhao Huang , Zhiyu Zhang , Tianxiang Xu , Chunshu Xia , Kaichun Hu , Yuchen Yang , Tongtong Pan , Dong Dong , Zhan Qin

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvements over audio-only…

In this paper, we propose an online speaker adaptation method for WaveNet-based neural vocoders in order to improve their performance on speaker-independent waveform generation. In this method, a speaker encoder is first constructed using a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Qiuchen Huang , Yang Ai , Zhenhua Ling

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle…

Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with robustness and generalization due to data quality constraints,…

Sound · Computer Science 2025-06-27 Rui Niu , Weihao Wu , Jie Chen , Long Ma , Zhiyong Wu

This paper presents an expressive speech synthesis architecture for modeling and controlling the speaking style at a word level. It attempts to learn word-level stylistic and prosodic representations of the speech data, with the aid of two…

Sound · Computer Science 2021-11-22 Konstantinos Klapsas , Nikolaos Ellinas , June Sig Sung , Hyoungmin Park , Spyros Raptis

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Yi Ren , Ming Lei , Zhiying Huang , Shiliang Zhang , Qian Chen , Zhijie Yan , Zhou Zhao

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively predicting magnitude and phase targets jointly. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Lingling Dai , Andong Li , Tong Lei , Meng Yu , Xiaodong Li , Chengshi Zheng

This paper proposes a novel neural audio codec, named APCodec+, which is an improved version of APCodec. The APCodec+ takes the audio amplitude and phase spectra as the coding object, and employs an adversarial training strategy.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-31 Hui-Peng Du , Yang Ai , Rui-Chen Zheng , Zhen-Hua Ling

Vocoder models have recently achieved substantial progress in generating authentic audio comparable to human quality while significantly reducing memory requirement and inference time. However, these data-hungry generative models require…

Sound · Computer Science 2023-12-19 Haoming Guo , Seth Z. Zhao , Jiachen Lian , Gopala Anumanchipalli , Gerald Friedland

In this paper, we show that a simple self-supervised pre-trained audio model can achieve comparable inference efficiency to more complicated pre-trained models with speech transformer encoders. These speech transformers rely on mixing…

Sound · Computer Science 2024-02-09 Sungho Jeon , Ching-Feng Yeh , Hakan Inan , Wei-Ning Hsu , Rashi Rungta , Yashar Mehdad , Daniel Bikel

Despite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F)…

Sound · Computer Science 2025-07-29 Andong Li , Tong Lei , Zhihang Sun , Rilin Chen , Erwei Yin , Xiaodong Li , Chengshi Zheng

Recent development of neural vocoders based on the generative adversarial neural network (GAN) has shown obvious advantages of generating raw waveform conditioned on mel-spectrogram with fast inference speed and lightweight networks.…

Sound · Computer Science 2023-05-30 Kun Song , Yongmao Zhang , Yi Lei , Jian Cong , Hanzhao Li , Lei Xie , Gang He , Jinfeng Bai