中文
相关论文

相关论文: S2Cap: A Benchmark and a Baseline for Singing Styl…

200 篇论文

Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing…

音频与语音处理 · 电气工程与系统科学 2026-01-22 Changhao Pan , Dongyu Yao , Yu Zhang , Wenxiang Guo , Jingyu Lu , Zhiyuan Zhu , Zhou Zhao

Audio captioning quality metrics which are typically borrowed from the machine translation and image captioning areas measure the degree of overlap between predicted tokens and gold reference tokens. In this work, we consider a metric…

多媒体 · 计算机科学 2023-03-06 Rehana Mahfuz , Yinyi Guo , Erik Visser

Generating visually grounded image captions with specific linguistic styles using unpaired stylistic corpora is a challenging task, especially since we expect stylized captions with a wide variety of stylistic patterns. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2023-08-03 Kanzhi Cheng , Zheng Ma , Shi Zong , Jianbing Zhang , Xinyu Dai , Jiajun Chen

Speaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech…

音频与语音处理 · 电气工程与系统科学 2020-11-18 Berrak Sisman , Junichi Yamagishi , Simon King , Haizhou Li

Voice Cloning has rapidly advanced in today's digital world, with many researchers and corporations working to improve these algorithms for various applications. This article aims to establish a standardized terminology for voice cloning…

声音 · 计算机科学 2025-05-02 Hussam Azzuni , Abdulmotaleb El Saddik

Image captioning models are usually evaluated on their ability to describe a held-out set of images, not on their ability to generalize to unseen concepts. We study the problem of compositional generalization, which measures how well a…

机器学习 · 计算机科学 2019-11-12 Mitja Nikolaus , Mostafa Abdou , Matthew Lamm , Rahul Aralikatte , Desmond Elliott

Recent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input of aggregated features, which hardly deviates from image…

计算机视觉与模式识别 · 计算机科学 2016-05-19 Andrew Shin , Katsunori Ohnishi , Tatsuya Harada

In conventional studies on environmental sound separation and synthesis using captions, datasets consisting of multiple-source sounds with their captions were used for model training. However, when we collect the captions for…

声音 · 计算机科学 2023-05-30 Yuki Okamoto , Kanta Shimonishi , Keisuke Imoto , Kota Dohi , Shota Horiguchi , Yohei Kawaguchi

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

声音 · 计算机科学 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

This work presents a text-to-audio-retrieval system based on pre-trained text and spectrogram transformers. Our method projects recordings and textual descriptions into a shared audio-caption space in which related examples from different…

音频与语音处理 · 电气工程与系统科学 2023-08-09 Paul Primus , Khaled Koutini , Gerhard Widmer

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness,…

声音 · 计算机科学 2022-10-19 Naoya Takahashi , Mayank Kumar , Singh , Yuki Mitsufuji

The task of isolating a target singing voice in music videos has useful applications. In this work, we explore the single-channel singing voice separation problem from a multimodal perspective, by jointly learning from audio and visual…

声音 · 计算机科学 2021-10-20 Juan F. Montesinos , Venkatesh S. Kadandale , Gloria Haro

In scholarly documents, figures provide a straightforward way of communicating scientific findings to readers. Automating figure caption generation helps move model understandings of scientific documents beyond text and will help authors…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Zhishen Yang , Raj Dabre , Hideki Tanaka , Naoaki Okazaki

Image Captioning is a task that requires models to acquire a multi-modal understanding of the world and to express this understanding in natural language text. While the state-of-the-art for this task has rapidly improved in terms of n-gram…

计算机视觉与模式识别 · 计算机科学 2018-12-20 Annika Lindh , Robert J. Ross , Abhijit Mahalunkar , Giancarlo Salton , John D. Kelleher

The present paper describes a singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of…

音频与语音处理 · 电气工程与系统科学 2019-06-26 Kazuhiro Nakamura , Kei Hashimoto , Keiichiro Oura , Yoshihiko Nankaku , Keiichi Tokuda

Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre-trained models. However, the direct application of speech…

声音 · 计算机科学 2024-06-21 Yuxun Tang , Yuning Wu , Jiatong Shi , Qin Jin

We discuss a novel task, Chorus Recognition, which could potentially benefit downstream tasks such as song search and music summarization. Different from the existing tasks such as music summarization or lyrics summarization relying on…

信息检索 · 计算机科学 2021-07-01 Jiaan Wang , Zhixu Li , Binbin Gu , Tingyi Zhang , Qingsheng Liu , Zhigang Chen

Concept-based interpretability methods like TCAV require clean, well-separated positive and negative examples for each concept. Existing music datasets lack this structure: tags are sparse, noisy, or ill-defined. We introduce ConceptCaps, a…

声音 · 计算机科学 2026-02-05 Bruno Sienkiewicz , Łukasz Neumann , Mateusz Modrzejewski

Singing voice beat tracking is a challenging task, due to the lack of musical accompaniment that often contains robust rhythmic and harmonic patterns, something most existing beat tracking systems utilize and can be essential for estimating…

声音 · 计算机科学 2025-03-14 Jiajun Deng , Yaolong Ju , Jing Yang , Simon Lui , Xunying Liu

Singing voice conversion is converting the timbre in the source singing to the target speaker's voice while keeping singing content the same. However, singing data for target speaker is much more difficult to collect compared with normal…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Liqiang Zhang , Chengzhu Yu , Heng Lu , Chao Weng , Chunlei Zhang , Yusong Wu , Xiang Xie , Zijin Li , Dong Yu