中文
相关论文

相关论文: Creating an A Cappella Singing Audio Dataset for A…

200 篇论文

Ornamentations, embellishments, or microtonal inflections are essential to melodic expression across many musical traditions, adding depth, nuance, and emotional impact to performances. Recognizing ornamentations in singing voices is key to…

音频与语音处理 · 电气工程与系统科学 2025-05-09 Sumit Kumar , Parampreet Singh , Vipul Arora

The evaluation of music understanding in Large Audio-Language Models (LALMs) requires a rigorously defined benchmark that truly tests whether models can perceive and interpret music, a standard that current data methodologies frequently…

计算与语言 · 计算机科学 2026-03-31 Benno Weck , Pablo Puentes , Andrea Poltronieri , Satyajeet Prabhu , Dmitry Bogdanov

Since the vocal component plays a crucial role in popular music, singing voice detection has been an active research topic in music information retrieval. Although several proposed algorithms have shown high performances, we argue that…

声音 · 计算机科学 2018-06-05 Kyungyun Lee , Keunwoo Choi , Juhan Nam

We present a Chinese judicial reading comprehension (CJRC) dataset which contains approximately 10K documents and almost 50K questions with answers. The documents come from judgment documents and the questions are annotated by law experts.…

计算与语言 · 计算机科学 2019-12-21 Xingyi Duan , Baoxin Wang , Ziyue Wang , Wentao Ma , Yiming Cui , Dayong Wu , Shijin Wang , Ting Liu , Tianxiang Huo , Zhen Hu , Heng Wang , Zhiyuan Liu

This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the…

声音 · 计算机科学 2024-10-07 Gabriel Bibbó , Thomas Deacon , Arshdeep Singh , Mark D. Plumbley

Separation of multiple singing voices into each voice is a rarely studied area in music source separation research. The absence of a benchmark dataset has hindered its progress. In this paper, we present an evaluation dataset and provide…

声音 · 计算机科学 2023-05-05 Chang-Bin Jeon , Hyeongi Moon , Keunwoo Choi , Ben Sangbae Chon , Kyogu Lee

Mandarin Chinese is characterized by being a tonal language; the pitch (or $F_0$) of its utterances carries considerable linguistic information. However, speech samples from different individuals are subject to changes in amplitude and…

Music information retrieval is currently an active research area that addresses the extraction of musically important information from audio signals, and the applications of such information. The extracted information can be used for search…

音频与语音处理 · 电气工程与系统科学 2022-04-08 Preeti Rao

Lyric translation plays a pivotal role in amplifying the global resonance of music, bridging cultural divides, and fostering universal connections. Translating lyrics, unlike conventional translation tasks, requires a delicate balance…

计算与语言 · 计算机科学 2023-08-29 Haven Kim , Kento Watanabe , Masataka Goto , Juhan Nam

The field of sound healing includes ancient practices coming from a broad range of cultures. Across such practices there is a variety of acoustic instrumentation utilised. Practitioners suggest that sound has the ability to target both…

声音 · 计算机科学 2019-10-23 Alice Baird , Bjoern Schuller

The present paper investigated automatic melody construction for Persian lyrics as an input. It was assumed that there is a phonological correlation between the lyric syllables and the melody in a song. A seq2seq neural network was…

声音 · 计算机科学 2024-10-29 Farshad Jafari , Farzad Didehvar , Amin Gheibi

Great research interests have been attracted to devise AI services that are able to provide mental health support. However, the lack of corpora is a main obstacle to this research, particularly in Chinese language. In this paper, we propose…

计算与语言 · 计算机科学 2021-06-04 Hao Sun , Zhenru Lin , Chujie Zheng , Siyang Liu , Minlie Huang

The Music Emotion Recognition (MER) field has seen steady developments in recent years, with contributions from feature engineering, machine learning, and deep learning. The landscape has also shifted from audio-centric systems to bimodal…

We present PKSpell: a data-driven approach for the joint estimation of pitch spelling and key signatures from MIDI files. Both elements are fundamental for the production of a full-fledged musical score and facilitate many MIR tasks such as…

声音 · 计算机科学 2021-07-30 Francesco Foscarin , Nicolas Audebert , Raphaël Fournier-S'Niehotta

With the rapid advancement of Multimodal Large Language Models (MLLMs), their potential has gained significant attention in Chinese Classical Studies (CCS). While existing research primarily focuses on text and visual modalities, the audio…

计算与语言 · 计算机科学 2026-04-14 Yexing Du , Kaiyuan Liu , Bihe Zhang , Youcheng Pan , Bo Yang , Liangyu Huo , Xiyuan Zhang , Jian Xie , Daojing He , Yang Xiang , Ming Liu , Bing Qin

This paper introduces The Spheres dataset, multitrack orchestral recordings designed to advance machine learning research in music source separation and related MIR tasks within the classical music domain. The dataset is composed of over…

Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Kavya Ranjan Saxena , Vipul Arora

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio…

声音 · 计算机科学 2023-09-07 Yuankun Xie , Jingjing Zhou , Xiaolin Lu , Zhenghao Jiang , Yuxin Yang , Haonan Cheng , Long Ye

In this work, we demonstrate a Chinese classical poetry generation system called Deep Poetry. Existing systems for Chinese classical poetry generation are mostly template-based and very few of them can accept multi-modal input. Unlike…

计算与语言 · 计算机科学 2019-11-20 Yusen Liu , Dayiheng Liu , Jiancheng Lv

Artificial Intelligence Generated Content (AIGC) is currently a popular research area. Among its various branches, song generation has attracted growing interest. Despite the abundance of available songs, effective data preparation remains…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Wei Tan , Shun Lei , Huaicheng Zhang , Guangzheng Li , Yixuan Zhang , Hangting Chen , Jianwei Yu , Rongzhi Gu , Dong Yu