中文
相关论文

相关论文: Marco-Voice Technical Report

200 篇论文

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthesis models currently rely on annotated audio data, but it is…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Rongjie Huang , Chunlei Zhang , Yongqi Wang , Dongchao Yang , Luping Liu , Zhenhui Ye , Ziyue Jiang , Chao Weng , Zhou Zhao , Dong Yu

Recently, there has been an increasing interest in neural speech synthesis. While the deep neural network achieves the state-of-the-art result in text-to-speech (TTS) tasks, how to generate a more emotional and more expressive speech is…

计算与语言 · 计算机科学 2021-06-24 Chenye Cui , Yi Ren , Jinglin Liu , Feiyang Chen , Rongjie Huang , Ming Lei , Zhou Zhao

We investigate a novel cross-lingual multi-speaker text-to-speech synthesis approach for generating high-quality native or accented speech for native/foreign seen/unseen speakers in English and Mandarin. The system consists of three…

音频与语音处理 · 电气工程与系统科学 2019-11-27 Zhaoyu Liu , Brian Mak

This paper presents CAMEO -- a curated collection of multilingual emotional speech datasets designed to facilitate research in emotion recognition and other speech-related tasks. The main objectives were to ensure easy access to the data,…

计算与语言 · 计算机科学 2026-01-28 Iwona Christop , Maciej Czajka

Emotional voice conversion (VC) aims to convert a neutral voice to an emotional (e.g. happy) one while retaining the linguistic information and speaker identity. We note that the decoupling of emotional features from other speech…

音频与语音处理 · 电气工程与系统科学 2021-10-05 Zhaojie Luo , Shoufeng Lin , Rui Liu , Jun Baba , Yuichiro Yoshikawa , Ishiguro Hiroshi

Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these…

声音 · 计算机科学 2025-10-13 Huu Tuong Tu , Huan Vu , cuong tien nguyen , Dien Hy Ngo , Nguyen Thi Thu Trang

Text-to-Speech (TTS) synthesis plays an important role in human-computer interaction. Currently, most TTS technologies focus on the naturalness of speech, namely,making the speeches sound like humans. However, the key tasks of the…

声音 · 计算机科学 2021-05-11 Jinyin Chen , Linhui Ye , Zhaoyan Ming

The goal of this work is to generate natural speech in multiple languages while maintaining the same speaker identity, a task known as cross-lingual speech synthesis. A key challenge of cross-lingual speech synthesis is the language-speaker…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Ji-Hoon Kim , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive…

声音 · 计算机科学 2022-01-19 Yi Lei , Shan Yang , Xinsheng Wang , Lei Xie

In this work, we propose a zero-shot voice conversion method using speech representations trained with self-supervised learning. First, we develop a multi-task model to decompose a speech utterance into features such as linguistic content,…

声音 · 计算机科学 2023-02-17 Shehzeen Hussain , Paarth Neekhara , Jocelyn Huang , Jason Li , Boris Ginsburg

When people try to influence others to do something, they subconsciously adjust their speech to include appropriate emotional information. In order for a robot to influence people in the same way, the robot should be able to imitate the…

Emotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we seek to generate speech…

计算与语言 · 计算机科学 2023-01-02 Kun Zhou , Berrak Sisman , Rajib Rana , B. W. Schuller , Haizhou Li

In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International…

The cross-speaker emotion transfer task in text-to-speech (TTS) synthesis particularly aims to synthesize speech for a target speaker with the emotion transferred from reference speech recorded by another (source) speaker. During the…

声音 · 计算机科学 2022-04-11 Tao Li , Xinsheng Wang , Qicong Xie , Zhichao Wang , Lei Xie

Speech-based digital biomarkers represent a scalable, non-invasive frontier for the early identification of Mild Cognitive Impairment (MCI). However, the development of robust diagnostic models remains impeded by acute clinical data…

计算与语言 · 计算机科学 2026-02-10 Rui Feng , Zhiyao Luo , Liuyu Wu , Wei Wang , Yuting Song , Yong Liu , Kok Pin Ng , Jianqing Li , Xingyao Wang

Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges when collecting sensitive emotional states. We introduce…

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack…

声音 · 计算机科学 2021-02-02 Paarth Neekhara , Shehzeen Hussain , Shlomo Dubnov , Farinaz Koushanfar , Julian McAuley

Preserving a speaker's voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, particularly in specialized domains such as scientific communication. In this paper, we…

音频与语音处理 · 电气工程与系统科学 2026-04-30 Amanuel Gizachew Abebe , Yasmin Moslem

We introduce a novel speech synthesis system, called NAUTILUS, that can generate speech with a target voice either from a text input or a reference utterance of an arbitrary source speaker. By using a multi-speaker speech corpus to train…

音频与语音处理 · 电气工程与系统科学 2020-10-08 Hieu-Thi Luong , Junichi Yamagishi

Emotional voice conversion (EVC) traditionally targets the transformation of spoken utterances from one emotional state to another, with previous research mainly focusing on discrete emotion categories. This paper departs from the norm by…

音频与语音处理 · 电气工程与系统科学 2023-09-19 Kun Zhou , Berrak Sisman , Carlos Busso , Bin Ma , Haizhou Li
‹ 上一页 1 2 3 10 下一页 ›