中文
相关论文

相关论文: CycleTransGAN-EVC: A CycleGAN-based Emotional Voic…

200 篇论文

We propose a way to use a transformer-based language model in conversational speech recognition. Specifically, we focus on decoding efficiently in a weighted finite-state transducer framework. We showcase an approach to lattice re-scoring…

计算与语言 · 计算机科学 2020-01-07 Kareem Nassar

Emotion recognition in conversations is essential for ensuring advanced human-machine interactions. However, creating robust and accurate emotion recognition systems in real life is challenging, mainly due to the scarcity of emotion…

计算与语言 · 计算机科学 2023-08-30 Théo Deschamps-Berger , Lori Lamel , Laurence Devillers

Integrating prior knowledge of neurophysiology into neural network architecture enhances the performance of emotion decoding. While numerous techniques emphasize learning spatial and short-term temporal patterns, there has been limited…

机器学习 · 计算机科学 2025-03-18 Yi Ding , Chengxuan Tong , Shuailei Zhang , Muyun Jiang , Yong Li , Kevin Lim Jun Liang , Cuntai Guan

Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using…

声音 · 计算机科学 2025-08-12 Jinsung Yoon , Wooyeol Jeong , Jio Gim , Young-Joo Suh

In this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior understanding. The model aims to capture and extract useful…

多媒体 · 计算机科学 2022-01-25 Minh Tran , Mohammad Soleymani

We propose Mask CycleGAN, a novel architecture for unpaired image domain translation built based on CycleGAN, with an aim to address two issues: 1) unimodality in image translation and 2) lack of interpretability of latent variables. Our…

机器学习 · 计算机科学 2022-05-17 Minfa Wang

Cycle-consistent generative adversarial networks (CycleGAN) were successfully applied to speech enhancement (SE) tasks with unpaired noisy-clean training data. The CycleGAN SE system adopted two generators and two discriminators trained…

音频与语音处理 · 电气工程与系统科学 2022-12-07 Wen-Yuan Ting , Syu-Siang Wang , Hsin-Li Chang , Borching Su , Yu Tsao

We propose a new Generative Adversarial Network for Compressed Video quality Enhancement (CVEGAN). The CVEGAN generator benefits from the use of a novel Mul2Res block (with multiple levels of residual learning branches), an enhanced…

图像与视频处理 · 电气工程与系统科学 2025-06-10 Di Ma , Fan Zhang , David R. Bull

In recent years, several high-performance conversational systems have been proposed based on the Transformer encoder-decoder model. Although previous studies analyzed the effects of the model parameters and the decoding method on subjective…

Polarimetric imaging, along with deep learning, has shown improved performances on different tasks including scene analysis. However, its robustness may be questioned because of the small size of the training datasets. Though the issue…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Cyprien Ruffino , Rachel Blin , Samia Ainouz , Gilles Gasso , Romain Hérault , Fabrice Meriaudeau , Stéphane Canu

Emotion recognition in conversation (ERC) aims to detect the emotion label for each utterance. Motivated by recent studies which have proven that feeding training examples in a meaningful order rather than considering them randomly can…

计算与语言 · 计算机科学 2022-04-22 Lin Yang , Yi Shen , Yue Mao , Longjun Cai

Different from the emotion recognition in individual utterances, we propose a multimodal learning framework using relation and dependencies among the utterances for conversational emotion analysis. The attention mechanism is applied to the…

计算与语言 · 计算机科学 2019-10-25 Zheng Lian , Jianhua Tao , Bin Liu , Jian Huang

Electrolarynx is a commonly used assistive device to help patients with removed vocal cords regain their ability to speak. Although the electrolarynx can generate excitation signals like the vocal cords, the naturalness and intelligibility…

声音 · 计算机科学 2023-06-13 Yung-Lun Chien , Hsin-Hao Chen , Ming-Chi Yen , Shu-Wei Tsai , Hsin-Min Wang , Yu Tsao , Tai-Shih Chi

This paper presents Fd-CycleGAN, an image-to-image (I2I) translation framework that enhances latent representation learning to approximate real data distributions. Building upon the foundation of CycleGAN, our approach integrates Local…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Shivangi Nigam , Adarsh Prasad Behera , Shekhar Verma , P. Nagabhushan

Emotion Recognition in Conversations (ERC) facilitates a deeper understanding of the emotions conveyed by speakers in each utterance within a conversation. Recently, Graph Neural Networks (GNNs) have demonstrated their strengths in…

计算与语言 · 计算机科学 2024-12-24 Cuong Tran Van , Thanh V. T. Tran , Van Nguyen , Truong Son Hy

This work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal target speech into the training criterion of…

Speech emotion conversion is the task of modifying the perceived emotion of a speech utterance while preserving the lexical content and speaker identity. In this study, we cast the problem of emotion conversion as a spoken language…

We present a novel learning-based framework for face reenactment. The proposed method, known as ReenactGAN, is capable of transferring facial movements and expressions from monocular video input of an arbitrary person to a target person.…

计算机视觉与模式识别 · 计算机科学 2018-07-31 Wayne Wu , Yunxuan Zhang , Cheng Li , Chen Qian , Chen Change Loy

General embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general tasks such as skip-gram models and natural language…

计算与语言 · 计算机科学 2020-11-03 Aparna Khare , Srinivas Parthasarathy , Shiva Sundaram

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. This paper proposes a hierarchical network with transformer encoders and memory mechanism to address this problem. The proposed…

声音 · 计算机科学 2020-11-02 Yanpei Shi , Mingjie Chen , Qiang Huang , Thomas Hain