中文
相关论文

相关论文: Multimodal Dataset Normalization and Perceptual Va…

200 篇论文

In musical compositions that include vocals, lyrics significantly contribute to artistic expression. Consequently, previous studies have introduced the concept of a recommendation system that suggests lyrics similar to a user's favorites or…

计算与语言 · 计算机科学 2024-08-28 Haven Kim , Taketo Akama

Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio for tasks such as cross-modal music retrieval, text-to-music…

Automatic Music Transcription (AMT) -- the task of converting music audio into note representations -- has seen rapid progress, driven largely by deep learning systems. Due to the limited availability of richly annotated music datasets,…

声音 · 计算机科学 2026-01-27 Lukáš Samuel Marták , Patricia Hu , Gerhard Widmer

While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory training data collected in natural…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Ege Ozguroglu , Junbang Liang , Ruoshi Liu , Mia Chiquier , Michael DeTienne , Wesley Wei Qian , Alexandra Horowitz , Andrew Owens , Carl Vondrick

Music Emotion Recogniser (MER) research faces challenges due to limited high-quality annotated datasets and difficulties in addressing cross-track feature drift. This work presents two primary contributions to address these issues.…

声音 · 计算机科学 2025-12-18 Qilin Li , C. L. Philip Chen , Tong Zhang

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Zitang Zhou , Ke Mei , Yu Lu , Tianyi Wang , Fengyun Rao

Deep learning has successfully shown excellent performance in learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal structures of…

多媒体 · 计算机科学 2019-08-13 Donghuo Zeng , Yi Yu , Keizo Oyama

Despite advances in deep algorithmic music generation, evaluation of generated samples often relies on human evaluation, which is subjective and costly. We focus on designing a homogeneous, objective framework for evaluating samples of…

Recommendation systems increasingly depend on massive human-labeled datasets; however, the human annotators hired to generate these labels increasingly come from homogeneous backgrounds. This poses an issue when downstream predictive models…

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Wangyuan Zhu , Jun Yu

Finding the music of the moment can often be a challenging problem, even for well-versed music listeners. Musical tastes are constantly in flux, and the problem of developing computational models for musical taste dynamics presents a rich…

信息检索 · 计算机科学 2018-06-19 Massimo Quadrana , Marta Reznakova , Tao Ye , Erik Schmidt , Hossein Vahabi

Computational food analysis (CFA) naturally requires multi-modal evidence of a particular food, e.g., images, recipe text, etc. A key to making CFA possible is multi-modal shared representation learning, which aims to create a joint…

计算机视觉与模式识别 · 计算机科学 2021-10-01 Ricardo Guerrero , Hai Xuan Pham , Vladimir Pavlovic

Nowadays, humans are constantly exposed to music, whether through voluntary streaming services or incidental encounters during commercial breaks. Despite the abundance of music, certain pieces remain more memorable and often gain greater…

信息检索 · 计算机科学 2024-05-22 Li-Yang Tseng , Tzu-Ling Lin , Hong-Han Shuai , Jen-Wei Huang , Wen-Whei Chang

Large language models (LLMs) make it easy to rewrite a text in any style -- e.g. to make it more polite, persuasive, or more positive -- but evaluation thereof is not straightforward. A challenge lies in measuring content preservation: that…

计算与语言 · 计算机科学 2025-09-18 Amalie Brogaard Pauli , Isabelle Augenstein , Ira Assent

This paper addresses the problem of cross-modal musical piece identification and retrieval: finding the appropriate recording(s) from a database given a sheet music query, and vice versa, working directly with audio and scanned sheet music…

音频与语音处理 · 电气工程与系统科学 2021-05-27 Luis Carvalho , Gerhard Widmer

The way we perceive a sound depends on many aspects-- its ecological frequency, acoustic features, typicality, and most notably, its identified source. In this paper, we present the HCU400: a dataset of 402 sounds ranging from easily…

音频与语音处理 · 电气工程与系统科学 2019-11-14 Ishwarya Ananthabhotla , David B. Ramsay , Joseph A. Paradiso

The study of Music Cognition and neural responses to music has been invaluable in understanding human emotions. Brain signals, though, manifest a highly complex structure that makes processing and retrieving meaningful features challenging,…

声音 · 计算机科学 2022-02-22 Kleanthis Avramidis , Christos Garoufis , Athanasia Zlatintsi , Petros Maragos

This paper introduces a new large-scale music dataset, MusicNet, to serve as a source of supervision and evaluation of machine learning methods for music research. MusicNet consists of hundreds of freely-licensed classical music recordings…

机器学习 · 统计学 2017-04-07 John Thickstun , Zaid Harchaoui , Sham Kakade

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that…

声音 · 计算机科学 2024-09-13 Tanisha Hisariya , Huan Zhang , Jinhua Liang

Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in…

多媒体 · 计算机科学 2025-04-29 Rashini Liyanarachchi , Aditya Joshi , Erik Meijering