中文
相关论文

相关论文: S2Cap: A Benchmark and a Baseline for Singing Styl…

200 篇论文

In recent years, advancements in representation learning and language models have propelled Automated Captioning (AC) to new heights, enabling the generation of human-level descriptions. Leveraging these advancements, we propose AVCap, an…

音频与语音处理 · 电气工程与系统科学 2024-07-12 Jongsuk Kim , Jiwon Shin , Junmo Kim

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

Automatic music captioning, which generates natural language descriptions for given music tracks, holds significant potential for enhancing the understanding and organization of large volumes of musical data. Despite its importance,…

声音 · 计算机科学 2023-08-01 SeungHeon Doh , Keunwoo Choi , Jongpil Lee , Juhan Nam

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

声音 · 计算机科学 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

Singing voice transcription converts recorded singing audio to musical notation. Sound contamination (such as accompaniment) and lack of annotated data make singing voice transcription an extremely difficult task. We take two approaches to…

声音 · 计算机科学 2023-04-25 Xiangming Gu , Wei Zeng , Jianan Zhang , Longshen Ou , Ye Wang

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the…

音频与语音处理 · 电气工程与系统科学 2026-03-26 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Junbo Zhang , Jian Luan

In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric.…

声音 · 计算机科学 2024-06-21 Yuxun Tang , Jiatong Shi , Yuning Wu , Qin Jin

This paper addresses the challenges and advancements in speech recognition for singing, a domain distinctly different from standard speech recognition. Singing encompasses unique challenges, including extensive pitch variations, diverse…

声音 · 计算机科学 2024-03-15 Anna Kruspe

In the recent years, singing voice separation systems showed increased performance due to the use of supervised training. The design of training datasets is known as a crucial factor in the performance of such systems. We investigate on how…

声音 · 计算机科学 2019-06-07 Laure Prétet , Romain Hennequin , Jimena Royo-Letelier , Andrea Vaglio

Significant strides have been made in creating voice identity representations using speech data. However, the same level of progress has not been achieved for singing voices. To bridge this gap, we suggest a framework for training singer…

声音 · 计算机科学 2024-01-11 Bernardo Torres , Stefan Lattner , Gaël Richard

Audio captioning is a multi-modal task, focusing on using natural language for describing the contents of general audio. Most audio captioning methods are based on deep neural networks, employing an encoder-decoder scheme and a dataset with…

声音 · 计算机科学 2020-07-10 Emre Çakır , Konstantinos Drossos , Tuomas Virtanen

We present S$^2$Voice, the winning system of the Singing Voice Conversion Challenge (SVCC) 2025 for both the in-domain and zero-shot singing style conversion tracks. Built on the strong two-stage Vevo baseline, S$^2$Voice advances style…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Ziqian Wang , Xianjun Xia , Chuanzeng Huang , Lei Xie

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric-based semantics,…

图形学 · 计算机科学 2025-09-03 Zikai Huang , Yihan Zhou , Xuemiao Xu , Cheng Xu , Xiaofen Xing , Jing Qin , Shengfeng He

Massive web datasets play a key role in the success of large vision-language models like CLIP and Flamingo. However, the raw web data is noisy, and existing filtering methods to reduce noise often come at the expense of data diversity. Our…

机器学习 · 计算机科学 2023-10-27 Thao Nguyen , Samir Yitzhak Gadre , Gabriel Ilharco , Sewoong Oh , Ludwig Schmidt

In traditional audio captioning methods, a model is usually trained in a fully supervised manner using a human-annotated dataset containing audio-text pairs and then evaluated on the test sets from the same dataset. Such methods have two…

声音 · 计算机科学 2024-06-11 Yiming Zhang , Xuenan Xu , Ruoyi Du , Haohe Liu , Yuan Dong , Zheng-Hua Tan , Wenwu Wang , Zhanyu Ma

This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025 (SVCC2025)-a novel singing style conversion system that advances fine-grained style conversion and control within in-domain settings. To…

声音 · 计算机科学 2026-04-08 Zhetao Hu , Yiquan Zhou , Wenyu Wang , Zhiyu Wu , Xin Gao , Jihua Zhu

The analysis, processing, and extraction of meaningful information from sounds all around us is the subject of the broader area of audio analytics. Audio captioning is a recent addition to the domain of audio analytics, a cross-modal…

音频与语音处理 · 电气工程与系统科学 2023-05-04 Sandeep Kothinti , Dimitra Emmanouilidou

The task of audio captioning is similar in essence to tasks such as image and video captioning. However, it has received much less attention. We propose three desiderata for captioning audio -- (i) fluency of the generated text, (ii)…

声音 · 计算机科学 2023-09-08 Tal Shaharabany , Ariel Shaulov , Lior Wolf

Singing voice synthesis is a generative task that involves multi-dimensional control of the singing model, including lyrics, pitch, and duration, and includes the timbre of the singer and singing skills such as vibrato. In this paper, we…

声音 · 计算机科学 2022-05-25 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Detecting singing voice deepfakes, or SingFake, involves determining the authenticity and copyright of a singing voice. Existing models for speech deepfake detection have struggled to adapt to unseen attacks in this unique singing voice…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Xuanjun Chen , Haibin Wu , Jyh-Shing Roger Jang , Hung-yi Lee