English
Related papers

Related papers: A Framework for Multi-f0 Modeling in SATB Choir Re…

200 papers

Speaker identification typically involves three stages. First, a front-end speaker embedding model is trained to embed utterance and speaker profiles. Second, a scoring function is applied between a runtime utterance and each speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Zhenning Tan , Yuguang Yang , Eunjung Han , Andreas Stolcke

More music foundation models are recently being released, promising a general, mostly task independent encoding of musical information. Common ways of adapting music foundation models to downstream tasks are probing and fine-tuning. These…

Sound · Computer Science 2024-12-02 Yiwei Ding , Alexander Lerch

This paper introduces a new score-informed method for the segmentation of jingju a cappella singing phrase into syllables. The proposed method estimates the most likely sequence of syllable boundaries given the estimated syllable onset…

Sound · Computer Science 2017-07-13 Jordi Pons , Rong Gong , Xavier Serra

In this paper, we present a machine-learning approach to pitch correction for voice in a karaoke setting, where the vocals and accompaniment are on separate tracks and time-aligned. The network takes as input the time-frequency…

Sound · Computer Science 2018-05-08 Sanna Wager , Lijiang Guo , Aswin Sivaraman , Minje Kim

Signal processing in the time-frequency plane has a long history and remains a field of methodological innovation. For instance, detection and denoising based on the zeros of the spectrogram have been proposed since 2015, contrasting with a…

Signal Processing · Electrical Eng. & Systems 2024-02-14 Juan M. Miramont , Rémi Bardenet , Pierre Chainais , Francois Auger

With the emergence of GAN-based vocoders, the discriminator, as a crucial component, has been developed recently. In our work, we focus on improving the time-frequency based discriminator. Particularly, Short-Time Fourier Transform (STFT)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-04 Nan Xu , Zhaolong Huang , Xiao Zeng

Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-28 Daisuke Niizumi , Daiki Takeuchi , Masahiro Yasuda , Binh Thien Nguyen , Yasunori Ohishi , Noboru Harada

Channel estimation is a critical task in multiple-input multiple-output (MIMO) digital communications that substantially effects end-to-end system performance. In this work, we introduce a novel approach for channel estimation using deep…

Signal Processing · Electrical Eng. & Systems 2022-11-09 Marius Arvinte , Jonathan I Tamir

Cavity-based noise detection schemes are combined with ultrafast pulse shaping as a means to diagnose the spectral correlations of both the amplitude and phase noise of an ultrafast frequency comb. The comb is divided into ten spectral…

Quantum Physics · Physics 2015-06-23 Roman Schmeissner , Jonathan Roslund , Claude Fabre , Nicolas Treps

High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not meet requirements for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-21 Rongjie Huang , Feiyang Chen , Yi Ren , Jinglin Liu , Chenye Cui , Zhou Zhao

We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single…

Sound · Computer Science 2019-05-07 Michael Michelashvili , Sagie Benaim , Lior Wolf

The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time…

Sound · Computer Science 2024-01-15 Haohe Liu , Xubo Liu , Qiuqiang Kong , Wenwu Wang , Mark D. Plumbley

While there has been much recent progress using deep learning techniques to separate speech and music audio signals, these systems typically require large collections of isolated sources during the training process. When extending audio…

Sound · Computer Science 2020-09-01 Fatemeh Pishdadian , Gordon Wichern , Jonathan Le Roux

In this paper, we tackle the singing voice phoneme segmentation problem in the singing training scenario by using language-independent information -- onset and prior coarse duration. We propose a two-step method. In the first step, we…

Sound · Computer Science 2018-06-06 Rong Gong , Xavier Serra

In this paper, robust detection, tracking and geometry estimation methods are developed and combined into a system for estimating time-difference estimates, microphone localization and sound source movement. No assumptions on the 3D…

Singing voice generation progresses rapidly, yet evaluating singing quality remains a critical challenge. Human subjective assessment, typically in the form of listening tests, is costly and time consuming, while existing objective metrics…

Sound · Computer Science 2026-01-28 Yuxun Tang , Lan Liu , Wenhao Feng , Yiwen Zhao , Jionghao Han , Yifeng Yu , Jiatong Shi , Qin Jin

This paper presents a high quality singing synthesizer that is able to model a voice with limited available recordings. Based on the sequence-to-sequence singing model, we design a multi-singer framework to leverage all the existing singing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-19 Jie Wu , Jian Luan

Singing voice separation based on deep learning relies on the usage of time-frequency masking. In many cases the masking process is not a learnable function or is not encapsulated into the deep learning optimization. Consequently, most of…

Uncertainty modeling in speaker representation aims to learn the variability present in speech utterances. While the conventional cosine-scoring is computationally efficient and prevalent in speaker recognition, it lacks the capability to…

Sound · Computer Science 2024-03-12 Qiongqiong Wang , Kong Aik Lee

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

Sound · Computer Science 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin