English
Related papers

Related papers: PAVITS: Exploring Prosody-aware VITS for End-to-En…

200 papers

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to…

In the latest social networks, more and more people prefer to express their emotions in videos through text, speech, and rich facial expressions. Multimodal video emotion analysis techniques can help understand users' inner world…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Qinglan Wei , Xuling Huang , Yuan Zhang

We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective information in a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-03 Jarod Duret , Titouan Parcollet , Yannick Estève

Neural evaluation metrics derived for numerous speech generation tasks have recently attracted great attention. In this paper, we propose SVSNet, the first end-to-end neural network model to assess the speaker voice similarity between…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-28 Cheng-Hung Hu , Yu-Huai Peng , Junichi Yamagishi , Yu Tsao , Hsin-Min Wang

Emotional Voice Conversion (EVC) aims to convert the emotional style of a source speech signal to a target style while preserving its content and speaker identity information. Previous emotional conversion studies do not disentangle…

Sound · Computer Science 2021-07-20 Xiangheng He , Junjie Chen , Georgios Rizos , Björn W. Schuller

Short-form videos (SVs) have become a vital part of our online routine for acquiring and sharing information. Their multimodal complexity poses new challenges for video analysis, highlighting the need for video emotion analysis (VEA) within…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xuecheng Wu , Dingkang Yang , Danlei Huang , Xinyi Yin , Yifan Wang , Jia Zhang , Jiayu Nie , Liangyu Fu , Yang Liu , Junxiao Xue , Hadi Amirpour , Wei Zhou

Emotion Prediction in Conversation (EPC) aims to forecast the emotions of forthcoming utterances by utilizing preceding dialogues. Previous EPC approaches relied on simple context modeling for emotion extraction, overlooking fine-grained…

Multimedia · Computer Science 2024-08-09 Haoxiang Shi , Ziqi Liang , Jun Yu

In affective neuroscience and emotion-aware AI, understanding how complex auditory stimuli drive emotion arousal dynamics remains unresolved. This study introduces a computational framework to model the brain's encoding of naturalistic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-29 Guandong Pan , Yaqian Yang , Shi Chen , Xin Wang , Longzhao Liu , Hongwei Zheng , Shaoting Tang

Data augmentation via voice conversion (VC) has been successfully applied to low-resource expressive text-to-speech (TTS) when only neutral data for the target speaker are available. Although the quality of VC is crucial for this approach,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-06 Ryo Terashima , Ryuichi Yamamoto , Eunwoo Song , Yuma Shirahata , Hyun-Wook Yoon , Jae-Min Kim , Kentaro Tachibana

Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which makes it challenging when converting a paragraph of texts into…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-11 Guanghui Xu , Wei Song , Zhengchen Zhang , Chao Zhang , Xiaodong He , Bowen Zhou

Neural Machine Translation (NMT) is the task of translating a text from one language to another with the use of a trained neural network. Several existing works aim at incorporating external information into NMT models to improve or control…

Computation and Language · Computer Science 2024-04-30 Charles Brazier , Jean-Luc Rouas

Singing voice correction (SVC) is an appealing application for amateur singers. Commercial products automate SVC by snapping pitch contours to equal-tempered scales, which could lead to deadpan modifications. Together with the neglect of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Yin-Jyun Luo , Yuen-Jen Lin , Li Su

We propose a lightweight end-to-end text-to-speech model using multi-band generation and inverse short-time Fourier transform. Our model is based on VITS, a high-quality end-to-end text-to-speech model, but adopts two changes for more…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-22 Masaya Kawamura , Yuma Shirahata , Ryuichi Yamamoto , Kentaro Tachibana

This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with…

Sound · Computer Science 2022-07-20 Juan F. Montesinos , Venkatesh S. Kadandale , Gloria Haro

Vision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Chen Zhu , Wangbo Zhao , Huiwen Zhang , Samir Khaki , Yuhao Zhou , Weidong Tang , Shuo Wang , Zhihang Yuan , Yuzhang Shang , Xiaojiang Peng , Kai Wang , Dawei Yang

Voice conversion (VC) is a task to transform a person's voice to different style while conserving linguistic contents. Previous state-of-the-art on VC is based on sequence-to-sequence (seq2seq) model, which could mislead linguistic…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-28 Tae-Ho Kim , Sungjae Cho , Shinkook Choi , Sejik Park , Soo-Young Lee

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Wayne Wu , Chen Change Loy , Xun Cao , Feng Xu

This paper introduces the T23 team's system submitted to the Singing Voice Conversion Challenge 2023. Following the recognition-synthesis framework, our singing conversion model is based on VITS, incorporating four key modules: a prior…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-05 Ziqian Ning , Yuepeng Jiang , Zhichao Wang , Bin Zhang , Lei Xie

Vision transformers (ViTs) have emerged as a significant area of focus, particularly for their capacity to be jointly trained with large language models and to serve as robust vision foundation models. Yet, the development of trustworthy…

Machine Learning · Computer Science 2024-11-04 Hengyi Wang , Shiwei Tan , Hao Wang

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-16 Aurosweta Mahapatra , Ismail Rasim Ulgen , Kong Aik Lee , Nicholas Andrews , Berrak Sisman