English
Related papers

Related papers: Finding Tori: Self-supervised Learning for Analyzi…

200 papers

In high-noise environments such as factories, subways, and busy streets, capturing clear speech is challenging. Throat microphones can offer a solution because of their inherent noise-suppression capabilities; however, the passage of sound…

Sound · Computer Science 2026-04-23 Yunsik Kim , Yonghun Song , Yoonyoung Chung

Interpretation of retrieved results is an important issue in music recommender systems, particularly from a user perspective. In this study, we investigate the methods for providing interpretability of content features using self-attention.…

Information Retrieval · Computer Science 2018-09-05 Seungjin Lee , Juheon Lee , Kyogu lee

Lyric translation plays a pivotal role in amplifying the global resonance of music, bridging cultural divides, and fostering universal connections. Translating lyrics, unlike conventional translation tasks, requires a delicate balance…

Computation and Language · Computer Science 2023-08-29 Haven Kim , Kento Watanabe , Masataka Goto , Juhan Nam

This study examines pitch contours as a unifying semantic construct prevalent across various audio domains including music, speech, bioacoustics, and everyday sounds. Analyzing pitch contours offers insights into the universal role of pitch…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-26 Jakob Abeßer , Simon Schwär , Meinard Müller

The problem of pitch tracking has been extensively studied in the speech research community. The goal of this paper is to investigate how these techniques should be adapted to singing voice analysis, and to provide a comparative evaluation…

Sound · Computer Science 2020-01-01 Onur Babacan , Thomas Drugman , Nicolas d'Alessandro , Nathalie Henrich , Thierry Dutoit

Tone is a prosodic feature used to distinguish words in many languages, some of which are endangered and scarcely documented. In this work, we use unsupervised representation learning to identify probable clusters of syllables that share…

Sound · Computer Science 2020-05-18 Bai Li , Jing Yi Xie , Frank Rudzicz

One of the main limitations in the field of audio signal processing is the lack of large public datasets with audio representations and high-quality annotations due to restrictions of copyrighted commercial music. We present Melon Playlist…

Natural language inference (NLI) and semantic textual similarity (STS) are key tasks in natural language understanding (NLU). Although several benchmark datasets for those tasks have been released in English and a few other languages, there…

Computation and Language · Computer Science 2020-10-06 Jiyeon Ham , Yo Joong Choe , Kyubyong Park , Ilji Choi , Hyungjoon Soh

Japanese idol groups, comprising performers known as "idols," are an indispensable part of Japanese pop culture. They frequently appear in live concerts and television programs, entertaining audiences with their singing and dancing. Similar…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Hitoshi Suda , Junya Koguchi , Shunsuke Yoshida , Tomohiko Nakamura , Satoru Fukayama , Jun Ogata

Classifying Sorani Kurdish subdialects poses a challenge due to the need for publicly available datasets or reliable resources like social media or websites for data collection. We conducted field visits to various cities and villages to…

Computation and Language · Computer Science 2024-04-02 Sana Isam , Hossein Hassani

Self-supervision methods learn representations by solving pretext tasks that do not require human-generated labels, alleviating the need for time-consuming annotations. These methods have been applied in computer vision, natural language…

Sound · Computer Science 2023-06-27 Giovana Morais , Matthew E. P. Davies , Marcelo Queiroz , Magdalena Fuentes

The ability of deep neural networks to learn complex data relations and representations is established nowadays, but it generally relies on large sets of training data. This work explores a "piece-specific" autoencoding scheme, in which a…

Sound · Computer Science 2022-03-09 Axel Marmoret , Jérémy E. Cohen , Frédéric Bimbot

The data-driven computational research on automatic jingju (also known as Beijing or Peking opera) singing evaluation lacks a suitable and comprehensive a cappella singing audio dataset. In this work, we present an a cappella singing audio…

Sound · Computer Science 2017-08-15 Rong Gong , Rafael Caro Repetto , Xavier Serra

Large-scale optical music recognition (OMR) research has focused mainly on Western staff notation, leaving Chinese Jianpu (numbered notation) and its rich lyric resources underexplored. We present a modular expert-system pipeline that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Fan Bu , Rongfeng Li , Zijin Li , Ya Li , Linfeng Fan , Pei Huang

In this paper, we explore the intersection of technology and cultural preservation by developing a self-supervised learning framework for the classification of musical symbols in historical manuscripts. Optical Music Recognition (OMR) plays…

Information Retrieval · Computer Science 2024-11-26 Elona Shatri , Daniel Raymond , George Fazekas

The task of classifying emotions within a musical track has received widespread attention within the Music Information Retrieval (MIR) community. Music emotion recognition has traditionally relied on the use of acoustic features, verbal…

Sound · Computer Science 2021-06-15 Nicholas Farris , Brian Model , Richard Savery , Gil Weinberg

Since the early twentieth century, intervals and tuning systems have been subjects of discussion among Iranian musicians and scholars. The process of Westernization and then a cultural back to roots movement are among the reasons that…

Sound · Computer Science 2021-08-04 Sepideh Shafiei

In the realm of digital music, using tags to efficiently organize and retrieve music from extensive databases is crucial for music catalog owners. Human tagging by experts is labor-intensive but mostly accurate, whereas automatic tagging…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-18 T. Aleksandra Ma , Alexander Lerch

Instrumental playing techniques such as vibratos, glissandos, and trills often denote musical expressivity, both in classical and folk contexts. However, most existing approaches to music similarity retrieval fail to describe timbre beyond…

This study presents FruitsMusic, a metadata corpus of Japanese idol-group songs in the real world, precisely annotated with who sings what and when. Japanese idol-group songs, vital to Japanese pop culture, feature a unique vocal…

Sound · Computer Science 2024-09-20 Hitoshi Suda , Shunsuke Yoshida , Tomohiko Nakamura , Satoru Fukayama , Jun Ogata