English
Related papers

Related papers: JVS-MuSiC: Japanese multispeaker singing-voice cor…

200 papers

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness,…

Sound · Computer Science 2022-10-19 Naoya Takahashi , Mayank Kumar , Singh , Yuki Mitsufuji

Recently, denoising diffusion models have demonstrated remarkable performance among generative models in various domains. However, in the speech domain, the application of diffusion models for synthesizing time-varying audio faces…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Ji-Sang Hwang , Sang-Hoon Lee , Seong-Whan Lee

Separating a song into vocal and accompaniment components is an active research topic, and recent years witnessed an increased performance from supervised training using deep learning techniques. We propose to apply the visual information…

Sound · Computer Science 2021-07-02 Bochen Li , Yuxuan Wang , Zhiyao Duan

Deep learning-based works for singing voice separation have performed exceptionally well in the recent past. However, most of these works do not focus on allowing users to interact with the model to improve performance. This can be crucial…

Sound · Computer Science 2025-12-03 Ankur Gupta , Anshul Rai , Archit Bansal , Vipul Arora

Evaluating 'anime-like' voices currently relies on costly subjective judgments, yet no standardized objective metric exists. A key challenge is that anime-likeness, unlike naturalness, lacks a shared absolute scale, making conventional Mean…

Sound · Computer Science 2026-03-13 Joonyong Park , Jerry Li

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of publicly available training data, hindering fair comparison and…

While existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios.…

Sound · Computer Science 2024-07-31 Tianrui Pan , Jie Liu , Bohan Wang , Jie Tang , Gangshan Wu

Singing Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by a vocoder to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-16 Jianwei Cui , Yu Gu , Shihao Chen , Jie Zhang , Liping Chen , Lirong Dai

High-quality datasets for learning-based modelling of polyphonic symbolic music remain less readily-accessible at scale than in other domains, such as language modelling or image classification. Deep learning algorithms show great potential…

Sound · Computer Science 2022-04-04 Omar Peracha

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Muhammad Saad Saeed , Shah Nawaz , Pietro Morerio , Arif Mahmood , Ignazio Gallo , Muhammad Haroon Yousaf , Alessio Del Bue

This paper introduces FLEURS-R, a speech restoration applied version of the Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) corpus. FLEURS-R maintains an N-way parallel speech corpus in 102 languages as FLEURS,…

Computation and Language · Computer Science 2024-08-13 Min Ma , Yuma Koizumi , Shigeki Karita , Heiga Zen , Jason Riesa , Haruko Ishikawa , Michiel Bacchiani

Singing voice separation (SVS) is a task that separates singing voice audio from its mixture with instrumental audio. Previous SVS studies have mainly employed the spectrogram masking method which requires a large dimensionality in…

Sound · Computer Science 2022-11-30 Jaekwon Im , Soonbeom Choi , Sangeon Yong , Juhan Nam

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associated body gestures.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-11 Shivam Mehta , Ruibo Tu , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

This paper presents our systems (denoted as T13) for the singing voice conversion challenge (SVCC) 2023. For both in-domain and cross-domain English singing voice conversion (SVC) tasks (Task 1 and Task 2), we adopt a recognition-synthesis…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Ryuichi Yamamoto , Reo Yoneyama , Lester Phillip Violeta , Wen-Chin Huang , Tomoki Toda

Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-13 Sewade Ogun , Vincent Colotte , Emmanuel Vincent

In the recent years, singing voice separation systems showed increased performance due to the use of supervised training. The design of training datasets is known as a crucial factor in the performance of such systems. We investigate on how…

Sound · Computer Science 2019-06-07 Laure Prétet , Romain Hennequin , Jimena Royo-Letelier , Andrea Vaglio

Speaker-specific anti-spoofing and synthesis-source tracing are central challenges in audio anti-spoofing. Progress has been hampered by the lack of datasets that systematically vary model architectures, synthesis pipelines, and generative…

Sound · Computer Science 2026-01-14 Surya Subramani , Hashim Ali , Hafiz Malik

Melody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key…

Sound · Computer Science 2025-02-10 Wei Chen , Binzhu Sha , Jing Yang , Zhuo Wang , Fan Fan , Zhiyong Wu

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain…

Computation and Language · Computer Science 2024-03-28 Injy Hamed , Fadhl Eryani , David Palfreyman , Nizar Habash

The goal of this paper is twofold. First, we introduce DALI, a large and rich multimodal dataset containing 5358 audio tracks with their time-aligned vocal melody notes and lyrics at four levels of granularity. The second goal is to explain…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Gabriel Meseguer-Brocal , Alice Cohen-Hadria , Geoffroy Peeters
‹ Prev 1 8 9 10 Next ›