English
Related papers

Related papers: Audiovisual Singing Voice Separation

200 papers

Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Ke Gu , Zhicong Wu , Peng Bai , Sitong Qiao , Zhiqi Jiang , Junchen Lu , Xiaodong Shi , Xinyuan Qian

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the…

Multimedia · Computer Science 2025-02-11 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly correlated nature of singing voices mixture. To address…

Sound · Computer Science 2026-01-21 Jihoo Jung , Ji-Hoon Kim , Doyeop Kwak , Junwon Lee , Juhan Nam , Joon Son Chung

Music demixing is the task of separating different tracks from the given single audio signal into components, such as drums, bass, and vocals from the rest of the accompaniment. Separation of sources is useful for a range of areas,…

Sound · Computer Science 2024-05-08 Roman Solovyev , Alexander Stempkovskiy , Tatiana Habruseva

We explore active audio-visual separation for dynamic sound sources, where an embodied agent moves intelligently in a 3D environment to continuously isolate the time-varying audio stream being emitted by an object of interest. The agent…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Sagnik Majumder , Kristen Grauman

In this paper, we propose a source separation method that is trained by observing the mixtures and the class labels of the sources present in the mixture without any access to isolated sources. Since our method does not require source class…

Sound · Computer Science 2019-08-06 Ertuğ Karamatlı , Ali Taylan Cemgil , Serap Kırbız

With the rapid development of neural network architectures and speech processing models, singing voice synthesis with neural networks is becoming the cutting-edge technique of digital music production. In this work, in order to explore how…

Sound · Computer Science 2021-08-29 Dengfeng Ke , Yuxing Lu , Xudong Liu , Yanyan Xu , Jing Sun , Cheng-Hao Cai

Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-22 Changhao Pan , Dongyu Yao , Yu Zhang , Wenxiang Guo , Jingyu Lu , Zhiyuan Zhu , Zhou Zhao

We present a deep learning method for singing voice conversion. The proposed network is not conditioned on the text or on the notes, and it directly converts the audio of one singer to the voice of another. Training is performed without any…

Machine Learning · Computer Science 2019-09-26 Eliya Nachmani , Lior Wolf

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Di Hu , Yake Wei , Rui Qian , Weiyao Lin , Ruihua Song , Ji-Rong Wen

There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised temporal…

Computer Vision and Pattern Recognition · Computer Science 2018-11-13 Bruno Korbar , Du Tran , Lorenzo Torresani

This paper presents a benchmark for singing voice enhancement. The development of singing voice enhancement is limited by the lack of realistic evaluation data. To address this gap, this paper introduces SingVERSE, the first real-world…

Sound · Computer Science 2025-09-26 Shaohan Jiang , Junan Zhang , Yunjia Zhang , Jing Yang , Fan Fan , Zhizheng Wu

Associating sound and its producer in complex audiovisual scene is a challenging task, especially when we are lack of annotated training data. In this paper, we present a flexible audiovisual model that introduces a soft-clustering module…

Computer Vision and Pattern Recognition · Computer Science 2020-01-28 Di Hu , Zheng Wang , Haoyi Xiong , Dong Wang , Feiping Nie , Dejing Dou

Automatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Chitralekha Gupta , Emre Yılmaz , Haizhou Li

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Qiushi Zhu , Jie Zhang , Yu Gu , Yuchen Hu , Lirong Dai

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and acoustic scene…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Abhinav Shukla , Stavros Petridis , Maja Pantic

Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the advantage that the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Tae-Woo Kim , Min-Su Kang , Gyeong-Hoon Lee

Vocal education in the music field is difficult to quantify due to the individual differences in singers' voices and the different quantitative criteria of singing techniques. Deep learning has great potential to be applied in music…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-01 Zhenyi Hou , Xu Zhao , Kejie Ye , Xinyu Sheng , Shanggerile Jiang , Jiajing Xia , Yitao Zhang , Chenxi Ban , Daijun Luo , Jiaxing Chen , Yan Zou , Yuchao Feng , Guangyu Fan , Xin Yuan

There has been a growing interest in using end-to-end acoustic models for singing voice synthesis (SVS). Typically, these models require an additional vocoder to transform the generated acoustic features into the final waveform. However,…

Sound · Computer Science 2023-08-08 Yuning Wu , Yifeng Yu , Jiatong Shi , Tao Qian , Qin Jin