中文
相关论文

相关论文: EAD-VC: Enhancing Speech Auto-Disentanglement for …

200 篇论文

Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both generate and understand audio. However, preserving key…

Dialogue disentanglement aims to detach the chronologically ordered utterances into several independent sessions. Conversation utterances are essentially organized and described by the underlying discourse, and thus dialogue disentanglement…

计算与语言 · 计算机科学 2023-06-13 Bobo Li , Hao Fei , Fei Li , Shengqiong Wu , Lizi Liao , Yinwei Wei , Tat-Seng Chua , Donghong Ji

This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong translation quality of…

计算与语言 · 计算机科学 2025-12-30 Jeongsoo Choi , Jaehun Kim , Joon Son Chung

The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP…

音频与语音处理 · 电气工程与系统科学 2025-09-22 Ziqi Dai , Yiting Chen , Jiacheng Xu , Liufei Xie , Yuchen Wang , Zhenchuan Yang , Bingsong Bai , Yangsheng Gao , Wenjiang Zhou , Weifeng Zhao , Ruohua Zhou

In this paper, we study the disentanglement of speaker and language representations in non-autoregressive cross-lingual TTS models from various aspects. We propose a phoneme length regulator that solves the length mismatch problem between…

音频与语音处理 · 电气工程与系统科学 2022-09-01 Haoyue Zhan , Xinyuan Yu , Haitong Zhang , Yang Zhang , Yue Lin

In this study, Disentanglement in Difference(DiD) is proposed to address the inherent inconsistency between the statistical independence of latent variables and the goal of semantic disentanglement in disentanglement representation…

机器学习 · 计算机科学 2025-04-04 Xingshen Zhang , Lin Wang , Shuangrong Liu , Xintao Lu , Chaoran Pang , Bo Yang

Ad-hoc distributed microphone environments, where microphone locations and numbers are unpredictable, present a challenge to traditional deep learning models, which typically require fixed architectures. To tailor deep learning models to…

音频与语音处理 · 电气工程与系统科学 2024-06-17 Jihyun Kim , Stijn Kindt , Nilesh Madhu , Hong-Goo Kang

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped…

音频与语音处理 · 电气工程与系统科学 2025-01-03 Shota Horiguchi , Takafumi Moriya , Atsushi Ando , Takanori Ashihara , Hiroshi Sato , Naohiro Tawara , Marc Delcroix

Learning good representations is of crucial importance in deep learning. Mutual Information (MI) or similar measures of statistical dependence are promising tools for learning these representations in an unsupervised way. Even though the…

音频与语音处理 · 电气工程与系统科学 2019-04-09 Mirco Ravanelli , Yoshua Bengio

Emotional voice conversion (EVC) traditionally targets the transformation of spoken utterances from one emotional state to another, with previous research mainly focusing on discrete emotion categories. This paper departs from the norm by…

音频与语音处理 · 电气工程与系统科学 2023-09-19 Kun Zhou , Berrak Sisman , Carlos Busso , Bin Ma , Haizhou Li

Automatic accent identification (AID) remains a challenging task due to the complex variability of accents, the entanglement of accent cues with speaker traits, and the scarcity of reliable accentlabelled data. To address these challenges,…

信号处理 · 电气工程与系统科学 2026-04-29 Rayane Bakari , Olivier Le Blouch , Nicolas Gengembre , Nicholas Evans

Decoding the attended speaker in a multi-speaker environment from electroencephalography (EEG) has attracted growing interest in recent years, with neuro-steered hearing devices as a driver application. Current approaches typically rely on…

信号处理 · 电气工程与系统科学 2026-02-05 Yuanyuan Yao , Simon Geirnaert , Tinne Tuytelaars , Alexander Bertrand

Disentanglement is a highly desirable property of representation owing to its similarity to human understanding and reasoning. Many works achieve disentanglement upon information bottlenecks (IB). Despite their elegant mathematical…

机器学习 · 计算机科学 2022-04-26 Jiantao Wu , Lin Wang , Bo Yang , Fanqi Li , Chunxiuzi Liu , Jin Zhou

Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for each target speaker.…

音频与语音处理 · 电气工程与系统科学 2018-06-26 Ju-chieh Chou , Cheng-chieh Yeh , Hung-yi Lee , Lin-shan Lee

End-to-end speech translation relies on data that pair source-language speech inputs with corresponding translations into a target language. Such data are notoriously scarce, making synthetic data augmentation by back-translation or…

计算与语言 · 计算机科学 2023-06-12 Tsz Kin Lam , Shigehiko Schamoni , Stefan Riezler

Recent advances in sophisticated synthetic speech generated from text-to-speech (TTS) or voice conversion (VC) systems cause threats to the existing automatic speaker verification (ASV) systems. Since such synthetic speech is generated from…

音频与语音处理 · 电气工程与系统科学 2022-12-15 Youngsik Eom , Yeonghyeon Lee , Ji Sub Um , Hoirin Kim

Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings remain underexplored.…

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disentangle the emotional…

音频与语音处理 · 电气工程与系统科学 2022-07-22 Zongyang Du , Berrak Sisman , Kun Zhou , Haizhou Li

We present an approach for unsupervised learning of speech representation disentangling contents and styles. Our model consists of: (1) a local encoder that captures per-frame information; (2) a global encoder that captures per-utterance…

计算与语言 · 计算机科学 2021-06-22 Andros Tjandra , Ruoming Pang , Yu Zhang , Shigeki Karita

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives:…