中文
相关论文

相关论文: A Holistic Cascade System, benchmark, and Human Ev…

200 篇论文

We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting module, which processes…

音频与语音处理 · 电气工程与系统科学 2024-06-26 Joun Yeop Lee , Myeonghun Jeong , Minchan Kim , Ji-Hyun Lee , Hoon-Young Cho , Nam Soo Kim

Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate, failing to distinguish SOTA systems. To address this, we…

声音 · 计算机科学 2026-03-26 Shengfan Shen , Di Wu , Xingchen Song , Dinghao Zhou , Liumeng Xue , Meng Meng , Jian Luan , Shuai Wang

It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with…

声音 · 计算机科学 2024-05-16 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Neural transducers have been widely used in automatic speech recognition (ASR). In this paper, we introduce it to streaming end-to-end speech translation (ST), which aims to convert audio signals to texts in other languages directly.…

计算与语言 · 计算机科学 2022-07-05 Jian Xue , Peidong Wang , Jinyu Li , Matt Post , Yashesh Gaur

Expressive speech synthesis requires vibrant prosody and well-timed pauses. We propose an effective strategy to augment a small dataset to train an expressive end-to-end Text-to-Speech model. We merge audios of emotionally congruent text…

声音 · 计算机科学 2026-02-12 Raymond Chung

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

音频与语音处理 · 电气工程与系统科学 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

End-to-end models for speech translation (ST) more tightly couple speech recognition (ASR) and machine translation (MT) than a traditional cascade of separate ASR and MT models, with simpler model architectures and the potential for reduced…

计算与语言 · 计算机科学 2020-05-29 Elizabeth Salesky , Alan W Black

We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and frame-level loss…

计算与语言 · 计算机科学 2023-12-27 Ziyang Ma , Zhisheng Zheng , Jiaxin Ye , Jinchao Li , Zhifu Gao , Shiliang Zhang , Xie Chen

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the style embeddings at one single scale from the information…

声音 · 计算机科学 2023-08-01 Shun Lei , Yixuan Zhou , Liyang Chen , Zhiyong Wu , Xixin Wu , Shiyin Kang , Helen Meng

Prosody modeling is important, but still challenging in expressive voice conversion. As prosody is difficult to model, and other factors, e.g., speaker, environment and content, which are entangled with prosody in speech, should be removed…

音频与语音处理 · 电气工程与系统科学 2022-01-04 Wendong Gan , Bolong Wen , Ying Yan , Haitao Chen , Zhichao Wang , Hongqiang Du , Lei Xie , Kaixuan Guo , Hai Li

In this paper, we describe our end-to-end multilingual speech translation system submitted to the IWSLT 2021 evaluation campaign on the Multilingual Speech Translation shared task. Our system is built by leveraging transfer learning across…

计算与语言 · 计算机科学 2021-08-17 Yun Tang , Hongyu Gong , Xian Li , Changhan Wang , Juan Pino , Holger Schwenk , Naman Goyal

Holistic perception of affective attributes is an important human perceptual ability. However, this ability is far from being realized in current affective computing, as not all of the attributes are well studied and their…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Yuanchao Li , Peter Bell , Catherine Lai

There has been a growing demand for automated spoken language assessment systems in recent years. A standard pipeline for this process is to start with a speech recognition system and derive features, either hand-crafted or based on…

音频与语音处理 · 电气工程与系统科学 2022-11-17 Stefano Bannò , Kate M. Knill , Marco Matassoni , Vyas Raina , Mark J. F. Gales

Attention-based seq2seq text-to-speech systems, especially those use self-attention networks (SAN), have achieved state-of-art performance. But an expressive corpus with rich prosody is still challenging to model as 1) prosodic aspects,…

音频与语音处理 · 电气工程与系统科学 2020-08-04 Fengyu Yang , Shan Yang , Qinghua Wu , Yujun Wang , Lei Xie

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED,…

声音 · 计算机科学 2021-02-04 Weiquan Fan , Xiangmin Xu , Xiaofen Xing , Weidong Chen , Dongyan Huang

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose…

Traditional voice conversion(VC) has been focused on speaker identity conversion for speech with a neutral expression. We note that emotional expression plays an essential role in daily communication, and the emotional style of speech can…

音频与语音处理 · 电气工程与系统科学 2021-10-22 Zongyang Du , Berrak Sisman , Kun Zhou , Haizhou Li

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale…

多媒体 · 计算机科学 2025-01-03 Yuan Zhao , Rui Liu , Gaoxiang Cong

Adding an emotions using prosody manipulation method for Indonesian text to speech system. Text To Speech (TTS) is a system that can convert text in one language into speech, accordance with the reading of the text in the language used. The…

声音 · 计算机科学 2016-06-30 Salita Ulitia Prini , Ary Setijadi Prihatmanto

We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous works, we explore joint training with both modalities, rather…

计算与语言 · 计算机科学 2022-10-17 Tara N. Sainath , Rohit Prabhavalkar , Ankur Bapna , Yu Zhang , Zhouyuan Huo , Zhehuai Chen , Bo Li , Weiran Wang , Trevor Strohman
‹ 上一页 1 8 9 10 下一页 ›