English
Related papers

Related papers: Finding Tori: Self-supervised Learning for Analyzi…

200 papers

Despite progress in melody-to-lyric generation, a substantial singability gap remains between machine-generated lyrics and those written by human lyricists. In this work, we aim to narrow this gap by jointly learning both wording and…

Computation and Language · Computer Science 2025-12-15 Longshen Ou , Xichu Ma , Ye Wang

The constant improvement of astronomical instrumentation provides the foundation for scientific discoveries. In general, these improvements have only implications forward in time, while previous observations do not benefit from this trend.…

Solar and Stellar Astrophysics · Physics 2025-05-29 Robert Jarolim , Astrid M. Veronig , Werner Pötzi , Tatiana Podladchikova

We introduce an extensive new dataset of MIDI files, created by transcribing audio recordings of piano performances into their constituent notes. The data pipeline we use is multi-stage, employing a language model to autonomously crawl and…

Sound · Computer Science 2025-07-01 Louis Bradshaw , Simon Colton

Lyrics alignment in long music recordings can be memory exhaustive when performed in a single pass. In this study, we present a novel method that performs audio-to-lyrics alignment with a low memory consumption footprint regardless of the…

Sound · Computer Science 2021-02-19 Emir Demirel , Sven Ahlbäck , Simon Dixon

Intention identification is a core issue in dialog management. However, due to the non-canonicality of the spoken language, it is difficult to extract the content automatically from the conversation-style utterances. This is much more…

Computation and Language · Computer Science 2019-07-10 Won Ik Cho , Young Ki Moon , Woo Hyun Kang , Nam Soo Kim

Choral music separation refers to the task of extracting tracks of voice parts (e.g., soprano, alto, tenor, and bass) from mixed audio. The lack of datasets has impeded research on this topic as previous work has only been able to train and…

Code switching, particularly between Korean and English, has become a defining feature of modern K-pop, reflecting both aesthetic choices and global market strategies. This paper is a primary investigation into the linguistic strategies…

Computation and Language · Computer Science 2025-09-30 Aditya Narayan Sankaran , Reza Farahbakhsh , Noel Crespi

This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as the training data, which relies on techniques and is not able…

Computation and Language · Computer Science 2019-12-30 Yusong Wu , Shengchen Li , Chengzhu Yu , Heng Lu , Chao Weng , Liqiang Zhang , Dong Yu

This paper outlines the methodology for modeling tonal learning in fully unsupervised models of human language acquisition. Tonal patterns are among the computationally most complex learning objectives in language. We argue that a realistic…

Computation and Language · Computer Science 2025-09-23 Kai Schenck , Gašper Beguš

Professional vocalists modulate their voice timbre or pitch to make their vocal performance more expressive. Such fluctuations are called singing techniques. Automatic detection of singing techniques from audio tracks can be beneficial to…

Sound · Computer Science 2023-06-27 Yuya Yamamoto , Juhan Nam , Hiroko Terasawa

In this paper, a time delay neural network (TDNN) based acoustic model is proposed to implement a fast-converged acoustic modeling for Korean speech recognition. The TDNN has an advantage in fast-convergence where the amount of training…

Computation and Language · Computer Science 2018-07-17 Hosung Park , Donghyun Lee , Minkyu Lim , Yoseb Kang , Juneseok Oh , Ji-Hwan Kim

We present a melody based classification of musical styles by exploiting the pitch and energy based characteristics derived from the audio signal. Three prominent musical styles were chosen which have improvisation as integral part with…

Sound · Computer Science 2019-06-24 Amruta Vidwans , Prateek Verma , Preeti Rao

Structure is one of the most essential aspects of music, and music structure is commonly indicated through repetition. However, the nature of repetition and structure in music is still not well understood, especially in the context of music…

Sound · Computer Science 2022-09-02 Shuqi Dai , Huiran Yu , Roger B. Dannenberg

In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric.…

Sound · Computer Science 2024-06-21 Yuxun Tang , Jiatong Shi , Yuning Wu , Qin Jin

Spoken content processing (such as retrieval and browsing) is maturing, but the singing content is still almost completely left out. Songs are human voice carrying plenty of semantic information just as speech, and may be considered as a…

Sound · Computer Science 2018-04-17 Che-Ping Tsai , Yi-Lin Tuan , Lin-shan Lee

Automated music playlist generation is a specific form of music recommendation. Generally stated, the user receives a set of song suggestions defining a coherent listening session. We hypothesize that the best way to convey such playlist…

Information Retrieval · Computer Science 2017-09-08 Andreu Vall , Hamid Eghbal-zadeh , Matthias Dorfer , Markus Schedl , Gerhard Widmer

Whereas chord transcription has received considerable attention during the past couple of decades, far less work has been devoted to transcribing and encoding the rhythmic patterns that occur in a song. The topic is especially relevant for…

Sound · Computer Science 2025-10-08 Aleksandr Lukoianov , Anssi Klapuri

Being able to predict whether a song can be a hit has impor- tant applications in the music industry. Although it is true that the popularity of a song can be greatly affected by exter- nal factors such as social and commercial influences,…

Sound · Computer Science 2017-04-06 Li-Chia Yang , Szu-Yu Chou , Jen-Yu Liu , Yi-Hsuan Yang , Yi-An Chen

HuQin is a family of traditional Chinese bowed string instruments. Playing techniques(PTs) embodied in various playing styles add abundant emotional coloring and aesthetic feelings to HuQin performance. The complex applied techniques make…

Multimedia · Computer Science 2023-10-10 Yu Zhang , Ziya Zhou , Xiaobing Li , Feng Yu , Maosong Sun

This study aims to compare three methods for translating ancient texts with sparse corpora: (1) the traditional statistical translation method of phrase alignment, (2) in-context LLM learning, and (3) proposed inter methodological approach…

Computation and Language · Computer Science 2024-07-17 Sojung Lucia Kim , Taehong Jang , Joonmo Ahn
‹ Prev 1 4 5 6 7 8 10 Next ›