English
Related papers

Related papers: CMI-RewardBench: Evaluating Music Reward Models wi…

200 papers

Music Recommendation Systems (MRSs) are a cornerstone of modern streaming platforms. Existing recommendation models, spanning both recall and ranking stages, predominantly rely on collaborative filtering, which fails to exploit the…

Information Retrieval · Computer Science 2026-04-24 Yizhi Zhou , Jia-Qi Yang , De-Chuan Zhan , Da-Wei Zhou

Analysing music in the field of machine learning is a very difficult problem with numerous constraints to consider. The nature of audio data, with its very high dimensionality and widely varying scales of structure, is one of the primary…

Sound · Computer Science 2022-05-17 Tracy Qian , Jackson Kaunismaa , Tony Chung

Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine-tuning generative models. However, the reward models…

Computation and Language · Computer Science 2026-02-04 Brian Christian , Hannah Rose Kirk , Jessica A. F. Thompson , Christopher Summerfield , Tsvetomira Dumbalska

Reinforcement Learning from Human Feedback has become the standard paradigm for language model alignment, where reward models directly determine alignment effectiveness. In this work, we focus on how to evaluate the generalizability of…

Computation and Language · Computer Science 2026-05-05 Yangyang Zhou , Yi-Chen Li

Commercial adoption of automatic music composition requires the capability of generating diverse and high-quality music suitable for the desired context (e.g., music for romantic movies, action games, restaurants, etc.). In this paper, we…

Sound · Computer Science 2022-11-18 Lee Hyun , Taehyun Kim , Hyolim Kang , Minjoo Ki , Hyeonchan Hwang , Kwanho Park , Sharang Han , Seon Joo Kim

With the rise of artificial intelligence (AI), there has been increasing interest in human-AI co-creation in a variety of artistic domains including music as AI-driven systems are frequently able to generate human-competitive artifacts.…

Human-Computer Interaction · Computer Science 2025-04-22 Renaud Bougueng Tchemeube , Jeff Ens , Cale Plut , Philippe Pasquier , Maryam Safi , Yvan Grabit , Jean-Baptiste Rolland

Music captioning has gained significant attention in the wake of the rising prominence of streaming media platforms. Traditional approaches often prioritize either the audio or lyrics aspect of the music, inadvertently ignoring the…

Sound · Computer Science 2023-10-24 Zihao He , Weituo Hao , Wei-Tsung Lu , Changyou Chen , Kristina Lerman , Xuchen Song

In recent years, remarkable advancements in artificial intelligence-generated content (AIGC) have been achieved in the fields of image synthesis and text generation, generating content comparable to that produced by humans. However, the…

Sound · Computer Science 2025-01-16 Sida Tian , Can Zhang , Wei Yuan , Wei Tan , Wenjie Zhu

We present an end-to-end system for musical key estimation, based on a convolutional neural network. The proposed system not only out-performs existing key estimation methods proposed in the academic literature; it is also capable of…

Machine Learning · Computer Science 2017-06-12 Filip Korzeniowski , Gerhard Widmer

Human annotations of mood in music are essential for music generation and recommender systems. However, existing datasets predominantly focus on Western songs with terms derived from English, which may limit generalizability across diverse…

Information Retrieval · Computer Science 2025-10-29 Harin Lee , Elif Çelen , Peter Harrison , Manuel Anglada-Tort , Pol van Rijn , Minsu Park , Marc Schönwiesner , Nori Jacoby

Multimodal emotion recognition plays a crucial role in enhancing user experience in human-computer interaction. Over the past few decades, researchers have proposed a series of algorithms and achieved impressive progress. Although each…

Human-Computer Interaction · Computer Science 2024-04-23 Zheng Lian , Licai Sun , Yong Ren , Hao Gu , Haiyang Sun , Lan Chen , Bin Liu , Jianhua Tao

The present methodology is aimed at cross-modal machine learning and uses multidisciplinary tools and methods drawn from a broad range of areas and disciplines, including music, systematic musicology, dance, motion capture, human-computer…

Human-Computer Interaction · Computer Science 2017-12-04 Fabio Paolizzo

Generating a complex work of art such as a musical composition requires exhibiting true creativity that depends on a variety of factors that are related to the hierarchy of musical language. Music generation have been faced with Algorithmic…

Sound · Computer Science 2021-09-08 Carlos Hernandez-Olivan , Jose R. Beltran

Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-27 Frank Cwitkowitz , Zhiyao Duan

Generating music has a few notable differences from generating images and videos. First, music is an art of time, necessitating a temporal model. Second, music is usually composed of multiple instruments/tracks with their own temporal…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Hao-Wen Dong , Wen-Yi Hsiao , Li-Chia Yang , Yi-Hsuan Yang

Recent work on music question answering (Music-QA) has primarily focused on single-track understanding, where models answer questions about an individual audio clip using its tags, captions, or metadata. However, listeners often describe…

Information Retrieval · Computer Science 2026-04-14 Junyoung Koh , Jaeyun Lee , Soo Yong Kim , Gyu Hyeong Choi , Jung In Koh , Jordan Phillips , Yeonjin Lee , Min Song

Mood recognition is an important problem in music informatics and has key applications in music discovery and recommendation. These applications have become even more relevant with the rise of music streaming. Our work investigates the…

Sound · Computer Science 2021-10-12 Rajnish Kumar , Manjeet Dahiya

Background music (BGM) can enhance the video's emotion. However, selecting an appropriate BGM often requires domain knowledge. This has led to the development of video-music retrieval techniques. Most existing approaches utilize pretrained…

Multimedia · Computer Science 2023-09-19 Tianjun Mao , Shansong Liu , Yunxuan Zhang , Dian Li , Ying Shan

Frontier AI models have achieved remarkable progress, yet recent studies suggest they struggle with compositional reasoning, often performing at or below random chance on established benchmarks. We revisit this problem and show that widely…

Artificial Intelligence · Computer Science 2026-04-27 Yinglun Zhu , Jiancheng Zhang , Fuzhi Tang

This paper introduces effective design choices for text-to-music retrieval systems. An ideal text-based retrieval system would support various input queries such as pre-defined tags, unseen tags, and sentence-level descriptions. In reality,…

Information Retrieval · Computer Science 2022-11-29 SeungHeon Doh , Minz Won , Keunwoo Choi , Juhan Nam