English
Related papers

Related papers: Towards Naturalistic Voice Conversion: NaturalVoic…

200 papers

The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic…

Sound · Computer Science 2026-03-17 Vamshi Nallaguntla , Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, we introduce MIT, a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Zeyu Zhu , Weijia Wu , Mike Zheng Shou

Generating spoken dialogue is inherently more complex than monologue text-to-speech (TTS), as it demands both realistic turn-taking and the maintenance of distinct speaker timbres. While existing autoregressive (AR) models have made…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Han Zhu , Wei Kang , Liyong Guo , Zengwei Yao , Fangjun Kuang , Weiji Zhuang , Zhaoqing Li , Zhifeng Han , Dong Zhang , Xin Zhang , Xingchen Song , Lingxuan Ye , Long Lin , Daniel Povey

Conversation is a subject of increasing interest in the social, cognitive, and computational sciences. Yet as conversational datasets continue to increase in size and complexity, researchers lack scalable methods to segment speech-to-text…

Computation and Language · Computer Science 2025-11-13 Gus Cooney , Andrew Reece

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional…

Sound · Computer Science 2024-12-13 Weizhen Bian , Yubo Zhou , Kaitai Zhang , Xiaohan Gu

Whether it is in the form of transcribed conversations, blog posts, or tweets, qualitative data provides a reader with rich insight into both the overarching trends as well as the diversity of human ideas expressed through text. Handling…

Human-Computer Interaction · Computer Science 2022-09-27 Huyen N. Nguyen , Tommy Dang , Kathleen A. Bowe

We present the latest iteration of the voice conversion challenge (VCC) series, a bi-annual scientific event aiming to compare and understand different voice conversion (VC) systems based on a common dataset. This year we shifted our focus…

Sound · Computer Science 2023-07-07 Wen-Chin Huang , Lester Phillip Violeta , Songxiang Liu , Jiatong Shi , Tomoki Toda

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-14 Yuke Lin , Xiaoyi Qin , Guoqing Zhao , Ming Cheng , Ning Jiang , Haiyang Wu , Ming Li

We propose noise-robust voice conversion (VC) which takes into account the recording quality and environment of noisy source speech. Conventional denoising training improves the noise robustness of a VC model by learning noisy-to-clean VC…

In recent times, voice assistants have become a part of our day-to-day lives, allowing information retrieval by voice synthesis, voice recognition, and natural language processing. These voice assistants can be found in many modern-day…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-03 Kashav Piya , Srijal Shrestha , Cameran Frank , Estephanos Jebessa , Tauheed Khan Mohd

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker…

Sound · Computer Science 2021-08-17 Joon Son Chung , Jaesung Huh , Arsha Nagrani , Triantafyllos Afouras , Andrew Zisserman

Data visualization (DV) has become the prevailing tool in the market due to its effectiveness into illustrating insights in vast amounts of data. To lower the barrier of using DVs, automatic DV tasks, such as natural language question (NLQ)…

Artificial Intelligence · Computer Science 2023-08-01 Yuanfeng Song , Xuefang Zhao , Raymond Chi-Wing Wong

News podcasts are a popular medium to stay informed and dive deep into news topics. Today, most podcasts are handcrafted by professionals. In this work, we advance the state-of-the-art in automatically generated podcasts, making use of…

Human-Computer Interaction · Computer Science 2022-02-16 Philippe Laban , Elicia Ye , Srujay Korlakunta , John Canny , Marti A. Hearst

Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in…

Sound · Computer Science 2024-05-02 Yimin Deng , Jianzong Wang , Xulong Zhang , Ning Cheng , Jing Xiao

Recently, voice conversion (VC) without parallel data has been successfully adapted to multi-target scenario in which a single model is trained to convert the input voice to many different speakers. However, such model suffers from the…

Machine Learning · Computer Science 2019-08-23 Ju-chieh Chou , Cheng-chieh Yeh , Hung-yi Lee

Recent advancements in conversational systems have significantly enhanced human-machine interactions across various domains. However, training these systems is challenging due to the scarcity of specialized dialogue data. Traditionally,…

Computation and Language · Computer Science 2026-05-29 Heydar Soudani , Roxana Petcu , Evangelos Kanoulas , Faegheh Hasibi

The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released…

Artificial Intelligence · Computer Science 2024-08-26 Irina-Elena Veliche , Zhuangqun Huang , Vineeth Ayyat Kochaniyan , Fuchun Peng , Ozlem Kalinli , Michael L. Seltzer

Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challenging due to the lack…

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural…

Sound · Computer Science 2023-12-08 Hira Dhamyal , Benjamin Elizalde , Soham Deshmukh , Huaming Wang , Bhiksha Raj , Rita Singh