中文
相关论文

相关论文: Hierarchical disentangled representation learning …

200 篇论文

Recently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned,…

声音 · 计算机科学 2023-02-28 Xue Jiang , Xiulian Peng , Yuan Zhang , Yan Lu

Any-to-any singing voice conversion (SVC) is an interesting audio editing technique, aiming to convert the singing voice of one singer into that of another, given only a few seconds of singing data. However, during the conversion process,…

音频与语音处理 · 电气工程与系统科学 2024-06-11 Shihao Chen , Yu Gu , Jie Zhang , Na Li , Rilin Chen , Liping Chen , Lirong Dai

Vector quantization (VQ) is a technique to deterministically learn features with discrete codebook representations. It is commonly performed with a variational autoencoding model, VQ-VAE, which can be further extended to hierarchical…

Multimodal representation learning seeks to relate and decompose information inherent in multiple modalities. By disentangling modality-specific information from information that is shared across modalities, we can improve interpretability…

机器学习 · 计算机科学 2025-03-18 Chenyu Wang , Sharut Gupta , Xinyi Zhang , Sana Tonekaboni , Stefanie Jegelka , Tommi Jaakkola , Caroline Uhler

The popular frameworks for self-supervised learning of speech representations have largely focused on frame-level masked prediction of speech regions. While this has shown promising downstream task performance for speech recognition and…

计算与语言 · 计算机科学 2025-07-22 Varun Krishna , Sriram Ganapathy

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of…

音频与语音处理 · 电气工程与系统科学 2020-04-23 Kazuhiro Nakamura , Shinji Takaki , Kei Hashimoto , Keiichiro Oura , Yoshihiko Nankaku , Keiichi Tokuda

Audio-visual representation learning is crucial for advancing multimodal speech processing tasks, such as lipreading and audio-visual speech recognition. Recently, speech foundation models (SFMs) have shown remarkable generalization…

音频与语音处理 · 电气工程与系统科学 2025-02-11 Jing-Xuan Zhang , Genshun Wan , Jianqing Gao , Zhen-Hua Ling

The aim of latent variable disentanglement is to infer the multiple informative latent representations that lie behind a data generation process and is a key factor in controllable data generation. In this paper, we propose a deep neural…

声音 · 计算机科学 2023-09-07 Yiming Wu

Singing voice detection (SVD), to recognize vocal parts in the song, is an essential task in music information retrieval (MIR). The task remains challenging since singing voice varies and intertwines with the accompaniment music, especially…

音频与语音处理 · 电气工程与系统科学 2022-05-09 Yifu Sun , Xulong Zhang , Yi Yu , Xi Chen , Wei Li

The careful construction of audio representations has become a dominant feature in the design of approaches to many speech tasks. Increasingly, such approaches have emphasized "disentanglement", where a representation contains only parts of…

Multi-speaker singing voice synthesis is to generate the singing voice sung by different speakers. To generalize to new speakers, previous zero-shot singing adaptation methods obtain the timbre of the target speaker with a fixed-size…

音频与语音处理 · 电气工程与系统科学 2022-01-12 Shoutong Wang , Jinglin Liu , Yi Ren , Zhen Wang , Changliang Xu , Zhou Zhao

The goal of this paper is to learn robust speaker representation for bilingual speaking scenario. The majority of the world's population speak at least two languages; however, most speaker recognition systems fail to recognise the same…

音频与语音处理 · 电气工程与系统科学 2023-06-08 Kihyun Nam , Youkyum Kim , Jaesung Huh , Hee Soo Heo , Jee-weon Jung , Joon Son Chung

In this paper, we propose an effective training strategy to ex-tract robust speaker representations from a speech signal. Oneof the key challenges in speaker recognition tasks is to learnlatent representations or embeddings containing…

音频与语音处理 · 电气工程与系统科学 2020-08-05 Yoohwan Kwon , Soo-Whan Chung , Hong-Goo Kang

This paper proposes singing voice synthesis (SVS) based on frame-level sequence-to-sequence models considering vocal timing deviation. In SVS, it is essential to synchronize the timing of singing with temporal structures represented by…

音频与语音处理 · 电气工程与系统科学 2023-02-23 Miku Nishihara , Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

This paper proposes a new task called spatial voice conversion, which aims to convert a target voice while preserving spatial information and non-target signals. Traditional voice conversion methods focus on single-channel waveforms,…

Vocal education in the music field is difficult to quantify due to the individual differences in singers' voices and the different quantitative criteria of singing techniques. Deep learning has great potential to be applied in music…

音频与语音处理 · 电气工程与系统科学 2024-11-01 Zhenyi Hou , Xu Zhao , Kejie Ye , Xinyu Sheng , Shanggerile Jiang , Jiajing Xia , Yitao Zhang , Chenxi Ban , Daijun Luo , Jiaxing Chen , Yan Zou , Yuchao Feng , Guangyu Fan , Xin Yuan

We present a large-scale comparative study of self-supervised speech representation (S3R)-based voice conversion (VC). In the context of recognition-synthesis VC, S3Rs are attractive owing to their potential to replace expensive supervised…

声音 · 计算机科学 2022-11-23 Wen-Chin Huang , Shu-Wen Yang , Tomoki Hayashi , Tomoki Toda

In this paper, we propose an invertible deep learning framework called INVVC for voice conversion. It is designed against the possible threats that inherently come along with voice conversion systems. Specifically, we develop an invertible…

音频与语音处理 · 电气工程与系统科学 2022-01-27 Zexin Cai , Ming Li

Voice conversion (VC) consists of digitally altering the voice of an individual to manipulate part of its content, primarily its identity, while maintaining the rest unchanged. Research in neural VC has accomplished considerable…

声音 · 计算机科学 2021-07-28 Laurent Benaroya , Nicolas Obin , Axel Roebel

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

声音 · 计算机科学 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin