English
Related papers

Related papers: Improving Prosody Modelling with Cross-Utterance B…

200 papers

This paper presents an end-to-end high-quality singing voice synthesis (SVS) system that uses bidirectional encoder representation from Transformers (BERT) derived semantic embeddings to improve the expressiveness of the synthesized singing…

Sound · Computer Science 2023-09-01 Shaohuan Zhou , Shun Lei , Weiya You , Deyi Tuo , Yuren You , Zhiyong Wu , Shiyin Kang , Helen Meng

The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for…

Sound · Computer Science 2026-01-21 Seymanur Akti , Alexander Waibel

We present a novel conversational-context aware end-to-end speech recognizer based on a gated neural network that incorporates conversational-context/word/speech embeddings. Unlike conventional speech recognition models, our model learns…

Computation and Language · Computer Science 2019-06-28 Suyoun Kim , Siddharth Dalmia , Florian Metze

Cross-lingual word sense disambiguation (WSD) tackles the challenge of disambiguating ambiguous words across languages given context. The pre-trained BERT embedding model has been proven to be effective in extracting contextual information…

Computation and Language · Computer Science 2020-12-11 Xingran Zhu

Although text-to-speech (TTS) systems have significantly improved, most TTS systems still have limitations in synthesizing speech with appropriate phrasing. For natural speech synthesis, it is important to synthesize the speech with a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-14 Ji-Sang Hwang , Sang-Hoon Lee , Seong-Whan Lee

Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from two limitations: 1) the acoustic model and vocoder are…

Sound · Computer Science 2022-07-12 Yanqing Liu , Ruiqing Xue , Lei He , Xu Tan , Sheng Zhao

This paper presents a simple yet effective method to achieve prosody transfer from a reference speech signal to synthesized speech. The main idea is to incorporate well-known acoustic correlates of prosody such as pitch and loudness…

Sound · Computer Science 2020-05-19 Siddharth Gururani , Kilol Gupta , Dhaval Shah , Zahra Shakeri , Jervis Pinto

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

Computation and Language · Computer Science 2025-08-19 Shumin Que , Anton Ragni

Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect corresponding recordings for…

Sound · Computer Science 2021-07-28 Shifeng Pan , Lei He

We propose a novel training strategy for Tacotron-based text-to-speech (TTS) system to improve the expressiveness of speech. One of the key challenges in prosody modeling is the lack of reference that makes explicit modeling difficult. The…

Sound · Computer Science 2021-04-13 Rui Liu , Berrak Sisman , Guanglai Gao , Haizhou Li

Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks when synthesizing…

Computation and Language · Computer Science 2024-12-20 Xiangheng He , Junjie Chen , Zixing Zhang , Björn W. Schuller

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the synthesized fine-grained prosody features often represent…

Sound · Computer Science 2023-03-15 Chunyu Qiang , Peng Yang , Hao Che , Ying Zhang , Xiaorui Wang , Zhongyuan Wang

This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrary-length speech insertion and even full sentence generation. In the proposed paradigm, global and local factors in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 Dacheng Yin , Chuanxin Tang , Yanqing Liu , Xiaoqiang Wang , Zhiyuan Zhao , Yucheng Zhao , Zhiwei Xiong , Sheng Zhao , Chong Luo

We present BERT-CTC-Transducer (BECTRA), a novel end-to-end automatic speech recognition (E2E-ASR) model formulated by the transducer with a BERT-enhanced encoder. Integrating a large-scale pre-trained language model (LM) into E2E-ASR has…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-20 Yosuke Higuchi , Tetsuji Ogawa , Tetsunori Kobayashi , Shinji Watanabe

Although end-to-end text-to-speech (TTS) models such as Tacotron have shown excellent results, they typically require a sizable set of high-quality <text, audio> pairs for training, which are expensive to collect. In this paper, we propose…

Computation and Language · Computer Science 2018-08-31 Yu-An Chung , Yuxuan Wang , Wei-Ning Hsu , Yu Zhang , RJ Skerry-Ryan

We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model…

Computation and Language · Computer Science 2024-06-28 Rendi Chevi , Alham Fikri Aji

This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC). Our formulation relaxes the conditional independence assumptions used in conventional CTC…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-21 Yosuke Higuchi , Brian Yan , Siddhant Arora , Tetsuji Ogawa , Tetsunori Kobayashi , Shinji Watanabe

Phrase representations derived from BERT often do not exhibit complex phrasal compositionality, as the model relies instead on lexical similarity to determine semantic relatedness. In this paper, we propose a contrastive fine-tuning…

Computation and Language · Computer Science 2021-10-15 Shufan Wang , Laure Thompson , Mohit Iyyer

Spoken language understanding is typically based on pipeline architectures including speech recognition and natural language understanding steps. These components are optimized independently to allow usage of available data, but the overall…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-13 Pavel Denisov , Ngoc Thang Vu

Recent advances in deep learning methods have elevated synthetic speech quality to human level, and the field is now moving towards addressing prosodic variation in synthetic speech.Despite successes in this effort, the state-of-the-art…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-30 Antti Suni , Sofoklis Kakouros , Martti Vainio , Juraj Šimko