中文
相关论文

相关论文: Marco-Voice Technical Report

200 篇论文

We present Marco-MoE, a suite of fully open multilingual sparse Mixture-of-Experts (MoE) models. Marco-MoE features a highly sparse design in which only around 5\% of the total parameters are activated per input token. This extreme…

计算与语言 · 计算机科学 2026-04-29 Fan Jiang , Yu Zhao , Chenyang Lyu , Tianqi Shi , Yichao Du , Feihu Jiang , Longyue Wang , Weihua Luo

With recent advancements in voice cloning, the performance of speech synthesis for a target speaker has been rendered similar to the human level. However, autoregressive voice cloning systems still suffer from text alignment failures,…

音频与语音处理 · 电气工程与系统科学 2022-01-27 Artem Gorodetskii , Ivan Ozhiganov

In this paper, we construct a Japanese audiobook speech corpus called "J-MAC" for speech synthesis research. With the success of reading-style speech synthesis, the research target is shifting to tasks that use complicated contexts.…

声音 · 计算机科学 2022-01-27 Shinnosuke Takamichi , Wataru Nakata , Naoko Tanji , Hiroshi Saruwatari

Existing Voice Cloning (VC) tasks aim to convert a paragraph text to a speech with desired voice specified by a reference audio. This has significantly boosted the development of artificial speech applications. However, there also exist…

计算机视觉与模式识别 · 计算机科学 2021-11-29 Qi Chen , Yuanqing Li , Yuankai Qi , Jiaqiu Zhou , Mingkui Tan , Qi Wu

We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective information in a…

音频与语音处理 · 电气工程与系统科学 2023-07-03 Jarod Duret , Titouan Parcollet , Yannick Estève

High-fidelity speech can be synthesized by end-to-end text-to-speech models in recent years. However, accessing and controlling speech attributes such as speaker identity, prosody, and emotion in a text-to-speech system remains a challenge.…

音频与语音处理 · 电气工程与系统科学 2020-08-05 Zexin Cai , Chuxiong Zhang , Ming Li

In this work, we introduce a framework for cross-lingual speech synthesis, which involves an upstream Voice Conversion (VC) model and a downstream Text-To-Speech (TTS) model. The proposed framework consists of 4 stages. In the first two…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Dariusz Piotrowski , Renard Korzeniowski , Alessio Falai , Sebastian Cygert , Kamil Pokora , Georgi Tinchev , Ziyao Zhang , Kayoko Yanagisawa

We present the first edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthetic speech. This challenge drew 22 participating teams from academia…

声音 · 计算机科学 2022-07-05 Wen-Chin Huang , Erica Cooper , Yu Tsao , Hsin-Min Wang , Tomoki Toda , Junichi Yamagishi

High-quality speech dialogue datasets are crucial for Speech-LLM development, yet existing acquisition methods face significant limitations. Human recordings incur high costs and privacy concerns, while synthetic approaches often lack…

计算与语言 · 计算机科学 2025-04-01 Minghan Wang , Ye Bai , Yuxia Wang , Thuy-Trang Vu , Ehsan Shareghi , Gholamreza Haffari

Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This…

音频与语音处理 · 电气工程与系统科学 2024-02-19 Vinotha R , Hepsiba D , L. D. Vijay Anand , Deepak John Reji

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Ziqi Zhang , Cheng Deng

The goal of our research is to automatically retrieve the satisfaction and the frustration in real-life call-center conversations. This study focuses an industrial application in which the customer satisfaction is continuously tracked down…

音频与语音处理 · 电气工程与系统科学 2023-10-10 Manon Macary , Marie Tahon , Yannick Estève , Daniel Luzzati

With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy in countering these…

声音 · 计算机科学 2024-05-16 Yang Hou , Haitao Fu , Chuankai Chen , Zida Li , Haoyu Zhang , Jianjun Zhao

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatically. In this work…

声音 · 计算机科学 2022-11-01 Luigi Attorresi , Davide Salvi , Clara Borrelli , Paolo Bestagini , Stefano Tubaro

Different from the emotion recognition in individual utterances, we propose a multimodal learning framework using relation and dependencies among the utterances for conversational emotion analysis. The attention mechanism is applied to the…

计算与语言 · 计算机科学 2019-10-25 Zheng Lian , Jianhua Tao , Bin Liu , Jian Huang

Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecified factor values,…

音频与语音处理 · 电气工程与系统科学 2025-05-28 Tianhua Qi , Shiyan Wang , Cheng Lu , Tengfei Song , Hao Yang , Zhanglin Wu , Wenming Zheng

Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality…

音频与语音处理 · 电气工程与系统科学 2025-04-01 Jinming Chen , Jingyi Fang , Yuanzhong Zheng , Yaoxuan Wang , Haojun Fei

Speech emotion recognition is a challenging task because the emotion expression is complex, multimodal and fine-grained. In this paper, we propose a novel multimodal deep learning approach to perform fine-grained emotion recognition from…

声音 · 计算机科学 2021-07-16 Hang Li , Wenbiao Ding , Zhongqin Wu , Zitao Liu

Speech representation learning has improved both speech understanding and speech synthesis tasks for single language. However, its ability in cross-lingual scenarios has not been explored. In this paper, we extend the pretraining method for…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Xiaoran Fan , Chao Pang , Tian Yuan , He Bai , Renjie Zheng , Pengfei Zhu , Shuohuan Wang , Junkun Chen , Zeyu Chen , Liang Huang , Yu Sun , Hua Wu

Zero-shot emotion transfer in cross-lingual speech synthesis refers to generating speech in a target language, where the emotion is expressed based on reference speech from a different source language. However, this task remains challenging…

音频与语音处理 · 电气工程与系统科学 2025-08-13 Tianlun Zuo , Jingbin Hu , Yuke Li , Xinfa Zhu , Hai Li , Ying Yan , Junhui Liu , Danming Xie , Lei Xie