English
Related papers

Related papers: SingNet: A Real-time Singing Voice Beat and Downbe…

200 papers

This paper proposes singing voice synthesis (SVS) based on frame-level sequence-to-sequence models considering vocal timing deviation. In SVS, it is essential to synchronize the timing of singing with temporal structures represented by…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Miku Nishihara , Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unified framework, and…

Sound · Computer Science 2019-12-24 Liqiang Zhang , Chengzhu Yu , Heng Lu , Chao Weng , Yusong Wu , Xiang Xie , Zijin Li , Dong Yu

Note-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications. Current AST methods, however,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-04 Ruiqi Li , Yu Zhang , Yongqi Wang , Zhiqing Hong , Rongjie Huang , Zhou Zhao

One of the key points in music recommendation is authoring engaging playlists according to sentiment and emotions. While previous works were mostly based on audio for music discovery and playlists generation, we take advantage of our…

Computation and Language · Computer Science 2019-01-16 Loreto Parisi , Simone Francia , Silvio Olivastri , Maria Stella Tavella

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio…

Sound · Computer Science 2023-09-07 Yuankun Xie , Jingjing Zhou , Xiaolin Lu , Zhenghao Jiang , Yuxin Yang , Haonan Cheng , Long Ye

Music is a form of expression that often requires interaction between players. If one wishes to interact in such a musical way with a computer, it is necessary for the machine to be able to interpret the input given by the human to find its…

Sound · Computer Science 2022-09-01 Filippo Carnovalini , Antonio Rodà

Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two…

Dialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We…

Sound · Computer Science 2023-12-05 Jihyun Lee , Yejin Jeon , Wonjun Lee , Yunsu Kim , Gary Geunbae Lee

Machine learning based singing voice models require large datasets and lengthy training times. In this work we present a lightweight architecture, based on the Differentiable Digital Signal Processing (DDSP) library, that is able to output…

Sound · Computer Science 2021-03-15 Juan Alonso , Cumhur Erkut

The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing systems have demonstrated promising performance on…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Tong Zhang , Yihuan Huang , Yanzhen Ren

This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-04 Xiang Hao , Xiangdong Su , Radu Horaud , Xiaofei Li

Recent progress in deep generative models has improved the quality of voice conversion in the speech domain. However, high-quality singing voice conversion (SVC) of unseen singers remains challenging due to the wider variety of musical…

Sound · Computer Science 2023-10-09 Naoya Takahashi , Mayank Kumar Singh , Yuki Mitsufuji

As the GAN-based face image and video generation techniques, widely known as DeepFakes, have become more and more matured and realistic, there comes a pressing and urgent demand for effective DeepFakes detectors. Motivated by the fact that…

Computer Vision and Pattern Recognition · Computer Science 2020-08-27 Hua Qi , Qing Guo , Felix Juefei-Xu , Xiaofei Xie , Lei Ma , Wei Feng , Yang Liu , Jianjun Zhao

Pitch manipulation is the process of producers adjusting the pitch of an audio segment to a specific key and intonation, which is essential in music production. Neural-network-based pitch-manipulation systems have been popular in recent…

Sound · Computer Science 2025-08-18 Yicheng Gu , Chaoren Wang , Zhizheng Wu , Lauri Juvela

Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal quality of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Ruiqi Li , Rongjie Huang , Yongqi Wang , Zhiqing Hong , Zhou Zhao

This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-22 Anmol Guragain , Tianchi Liu , Zihan Pan , Hardik B. Sailor , Qiongqiong Wang

Automatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Chitralekha Gupta , Emre Yılmaz , Haizhou Li

We propose a novel unsupervised singing voice detection method which use single-channel Blind Audio Source Separation (BASS) algorithm as a preliminary step. To reach this goal, we investigate three promising BASS approaches which operate…

Sound · Computer Science 2018-05-04 Dominique Fourer , Geoffroy Peeters

Reverb plays a critical role in music production, where it provides listeners with spatial realization, timbre, and texture of the music. Yet, it is challenging to reproduce the musical reverb of a reference music track even by skilled…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-04 Junghyun Koo , Seungryeol Paik , Kyogu Lee

Time-series forecasting is crucial for numerous real-world applications including weather prediction and financial market modeling. While temporal-domain methods remain prevalent, frequency-domain approaches can effectively capture…

Machine Learning · Computer Science 2025-08-05 Zhixuan Li , Naipeng Chen , Seonghwa Choi , Sanghoon Lee , Weisi Lin