English
Related papers

Related papers: Hear: Hierarchically Enhanced Aesthetic Representa…

200 papers

The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing mistake detection in the context of music pedagogy,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-09 Sumit Kumar , Suraj Jaiswal , Parampreet Singh , Vipul Arora

In recent years, AI-generated music has made significant progress, with several models performing well in multimodal and complex musical genres and scenes. While objective metrics can be used to evaluate generative music, they often lack…

Sound · Computer Science 2023-08-29 Zeyu Xiong , Weitao Wang , Jing Yu , Yue Lin , Ziyan Wang

We present an empirical study on embedding the lyrics of a song into a fixed-dimensional feature for the purpose of music tagging. Five methods of computing token-level and four methods of computing document-level representations are…

Computation and Language · Computer Science 2021-12-22 Matt McVicar , Bruno Di Giorgi , Baris Dundar , Matthias Mauch

Generating full-length, high-quality songs is challenging, as it requires maintaining long-term coherence both across text and music modalities and within the music modality itself. Existing non-autoregressive (NAR) frameworks, while…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Yuepeng Jiang , Huakang Chen , Ziqian Ning , Jixun Yao , Zerui Han , Di Wu , Meng Meng , Jian Luan , Zhonghua Fu , Lei Xie

We present an end-to-end system for musical key estimation, based on a convolutional neural network. The proposed system not only out-performs existing key estimation methods proposed in the academic literature; it is also capable of…

Machine Learning · Computer Science 2017-06-12 Filip Korzeniowski , Gerhard Widmer

We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-04 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Hsin-Min Wang , Yu Tsao

Many music AI models learn a map between music content and human-defined labels. However, many annotations, such as chords, can be naturally expressed within the music modality itself, e.g., as sequences of symbolic notes. This observation…

Sound · Computer Science 2025-09-30 Junyan Jiang , Daniel Chin , Liwei Lin , Xuanjie Liu , Gus Xia

This paper explores a new natural language processing task, review-driven multi-label music style classification. This task requires the system to identify multiple styles of music based on its reviews on websites. The biggest challenge…

Computation and Language · Computer Science 2018-08-24 Guangxiang Zhao , Jingjing Xu , Qi Zeng , Xuancheng Ren

This work proposes a novel feature selection algorithm to classify Songs into different groups. Classification of musical content is often a non-trivial job and still relatively less explored area. The main idea conveyed in this article is…

Information Retrieval · Computer Science 2019-01-09 Anish Acharya

Symbolic Music Emotion Recognition(SMER) is to predict music emotion from symbolic data, such as MIDI and MusicXML. Previous work mainly focused on learning better representation via (mask) language model pre-training but ignored the…

Sound · Computer Science 2022-01-19 Jibao Qiu , C. L. Philip Chen , Tong Zhang

The complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Yuanchao Li , Azalea Gui , Dimitra Emmanouilidou , Hannes Gamper

While efficient architectures and a plethora of augmentations for end-to-end image classification tasks have been suggested and heavily investigated, state-of-the-art techniques for audio classifications still rely on numerous…

Sound · Computer Science 2022-07-06 Avi Gazneli , Gadi Zimerman , Tal Ridnik , Gilad Sharir , Asaf Noy

Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Arthur N. dos Santos , Bruno S. Masiero

Previous approaches in singer identification have used one of monophonic vocal tracks or mixed tracks containing multiple instruments, leaving a semantic gap between these two domains of audio. In this paper, we present a system to learn a…

Sound · Computer Science 2019-06-27 Kyungyun Lee , Juhan Nam

During music listening, cortical activity encodes both acoustic and expectation-related information. Prior work has shown that ANN representations resemble cortical representations and can serve as supervisory signals for EEG recognition.…

Artificial Intelligence · Computer Science 2026-05-19 Shogo Noguchi , Taketo Akama , Tai Nakamura , Shun Minamikawa , Natalia Polouliakh

Audio-to-score alignment is a long-standing challenge in music information retrieval and arguably the most widely applicable alignment task for music research. Alignment algorithms match two versions of a piece of music, and for this to…

Sound · Computer Science 2026-05-20 Silvan Peter , Patricia Hu , Gerhard Widmer

Constructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as multi-instrument synthesis and timbre transfer. The…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-28 Xuan Shi , Erica Cooper , Junichi Yamagishi

General-purpose embedding is highly desirable for few-shot even zero-shot learning in many application scenarios, including audio tasks. In order to understand representations better, we conducted a thorough error analysis and visualization…

Sound · Computer Science 2023-03-08 Ankit Shah , Shuyi Chen , Kejun Zhou , Yue Chen , Bhiksha Raj

Hearables are becoming ubiquitous, yet their sound controls remain blunt: users can either enable global noise suppression or focus on a single target sound. Real-world acoustic scenes, however, contain many simultaneous sources that users…

Sound · Computer Science 2026-03-06 Seunghyun Oh , Malek Itani , Aseem Gauri , Shyamnath Gollakota

Modeling various aspects that make a music piece unique is a challenging task, requiring the combination of multiple sources of information. Deep learning is commonly used to obtain representations using various sources of information, such…

Sound · Computer Science 2021-04-05 Andres Ferraro , Xavier Favory , Konstantinos Drossos , Yuntae Kim , Dmitry Bogdanov