English
Related papers

Related papers: Computational Narrative Understanding for Expressi…

200 papers

An end-to-end (e2e) text-to-speech (TTS) system is a deep architecture that learns to associate a text string with acoustic speech patterns from a curated dataset. It is expected that all aspects associated with speech production, such as…

Sound · Computer Science 2026-02-17 Parth Khadse , Sunil Kumar Kopparapu

We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting module, which processes…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-26 Joun Yeop Lee , Myeonghun Jeong , Minchan Kim , Ji-Hyun Lee , Hoon-Young Cho , Nam Soo Kim

Unlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of speech often end at…

Sound · Computer Science 2023-10-05 Robin Netzorg , Bohan Yu , Andrea Guzman , Peter Wu , Luna McNulty , Gopala Anumanchipalli

We present a new neural text to speech (TTS) method that is able to transform text to speech in voices that are sampled in the wild. Unlike other systems, our solution is able to deal with unconstrained voice samples and without requiring…

Machine Learning · Computer Science 2018-02-02 Yaniv Taigman , Lior Wolf , Adam Polyak , Eliya Nachmani

We propose a new paradigm to help Large Language Models (LLMs) generate more accurate factual knowledge without retrieving from an external corpus, called RECITation-augmented gEneration (RECITE). Different from retrieval-augmented language…

Computation and Language · Computer Science 2023-02-17 Zhiqing Sun , Xuezhi Wang , Yi Tay , Yiming Yang , Denny Zhou

Emotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity. In this paper, we propose a novel 2-stage training strategy for sequence-to-sequence emotional…

Computation and Language · Computer Science 2021-06-10 Kun Zhou , Berrak Sisman , Haizhou Li

Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a reference speech sample. There have been considerable advances in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Alexandra Torresquintero , Tian Huey Teh , Christopher G. R. Wallis , Marlene Staib , Devang S Ram Mohan , Vivian Hu , Lorenzo Foglianti , Jiameng Gao , Simon King

State of the art (SOTA) neural text to speech (TTS) models can generate natural-sounding synthetic voices. These models are characterized by large memory footprints and substantial number of operations due to the long-standing focus on…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-24 Rowel Atienza

The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustrate pho- netic concepts that are otherwise difficult to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-10 Frederik Rautenberg , Fritz Seebauer , Jana Wiechmann , Michael Kuhlmann , Petra Wagner , Reinhold Haeb-Umbach

We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual training objectives to…

Sound · Computer Science 2026-03-13 Suvendu Sekhar Mohanty

An unsupervised text-to-speech synthesis (TTS) system learns to generate speech waveforms corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that language; 2) a collection…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-17 Junrui Ni , Liming Wang , Heting Gao , Kaizhi Qian , Yang Zhang , Shiyu Chang , Mark Hasegawa-Johnson

This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional…

Sound · Computer Science 2025-09-22 Zhiwen Qian , Jinhua Liang , Huan Zhang

Although diffusion-based, non-autoregressive text-to-speech (TTS) systems have demonstrated impressive zero-shot synthesis capabilities, their efficacy is still hindered by two key challenges: the difficulty of text-speech alignment…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Chunyat Wu , Jiajun Deng , Zhengxi Liu , Zheqi Dai , Haolin He , Qiuqiang Kong

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose Spotlight-TTS,…

Sound · Computer Science 2025-07-01 Nam-Gyu Kim , Deok-Hyeon Cho , Seung-Bin Kim , Seong-Whan Lee

This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguistic and paralinguistic information. At its core is a speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-06 Jiyoung Lee , Song Park , Sanghyuk Chun , Soo-Whan Chung

Recent technological advances have made it possible to build real-time, interactive spoken dialogue systems for a wide variety of applications. However, when users do not respect the limitations of such systems, performance typically…

Computation and Language · Computer Science 2007-05-23 Diane J. Litman , Shimei Pan

We present enhancements to a speech-to-speech translation pipeline in order to perform automatic dubbing. Our architecture features neural machine translation generating output of preferred length, prosodic alignment of the translation with…

Computation and Language · Computer Science 2020-02-04 Marcello Federico , Robert Enyedi , Roberto Barra-Chicote , Ritwik Giri , Umut Isik , Arvindh Krishnaswamy , Hassan Sawaf

There has been limited evaluation of advanced Text-to-Speech (TTS) models with Mathematical eXpressions (MX) as inputs. In this work, we design experiments to evaluate quality and intelligibility of five TTS models through listening and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-16 Sujoy Roychowdhury , H. G. Ranjani , Sumit Soman , Nishtha Paul , Subhadip Bandyopadhyay , Siddhanth Iyengar

Speech-to-Text Translation (S2TT) systems built from Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) modules face two major limitations: error propagation and the inability to exploit prosodic or other acoustic cues.…

Computation and Language · Computer Science 2025-10-06 Jacobo Romero-Díaz , Gerard I. Gállego , Oriol Pareras , Federico Costa , Javier Hernando , Cristina España-Bonet

Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete symbols. Audio tokenization often poses a necessary…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Zhijun Liu , Shuai Wang , Sho Inoue , Qibing Bai , Haizhou Li
‹ Prev 1 8 9 10 Next ›