中文
相关论文

相关论文: MOSA: Music Motion with Semantic Annotation Datase…

200 篇论文

People are sharing their opinions, stories and reviews through online video sharing websites every day. Studying sentiment and subjectivity in these opinion videos is experiencing a growing attention from academia and industry. While…

计算与语言 · 计算机科学 2016-11-18 Amir Zadeh , Rowan Zellers , Eli Pincus , Louis-Philippe Morency

We introduce an extensive new dataset of MIDI files, created by transcribing audio recordings of piano performances into their constituent notes. The data pipeline we use is multi-stage, employing a language model to autonomously crawl and…

声音 · 计算机科学 2025-07-01 Louis Bradshaw , Simon Colton

Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be…

We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand the…

声音 · 计算机科学 2025-11-21 Suhwan Choi , Kyu Won Kim , Myungjoo Kang

Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design, virtual prototyping, and emerging data-driven engine sound synthesis methods. These applications require large volumes of…

声音 · 计算机科学 2026-03-10 Robin Doerfler , Lonce Wyse

We introduce MMIS, a novel dataset designed to advance MultiModal Interior Scene generation and recognition. MMIS consists of nearly 160,000 images. Each image within the dataset is accompanied by its corresponding textual description and…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Hozaifa Kassab , Ahmed Mahmoud , Mohamed Bahaa , Ammar Mohamed , Ali Hamdi

The Music Emotion Recognition (MER) field has seen steady developments in recent years, with contributions from feature engineering, machine learning, and deep learning. The landscape has also shifted from audio-centric systems to bimodal…

In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamlined encoder-decoder design of LLaVA, enabling effective…

声音 · 计算机科学 2026-05-04 Zuyao You , Zhesong Yu , Mingyu Liu , Bilei Zhu , Yuan Wan , Zuxuan Wu

MEx: Multi-modal Exercises Dataset is a multi-sensor, multi-modal dataset, implemented to benchmark Human Activity Recognition(HAR) and Multi-modal Fusion algorithms. Collection of this dataset was inspired by the need for recognising and…

计算机视觉与模式识别 · 计算机科学 2019-08-27 Anjana Wijekoon , Nirmalie Wiratunga , Kay Cooper

This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for downstream tasks such as humming melody based music generation.…

声音 · 计算机科学 2023-10-18 Shansong Liu , Xu Li , Dian Li , Ying Shan

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within audio and video…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Sheng Wu , Xiaobao Wang , Longbiao Wang , Dongxiao He , Jianwu Dang

At the core of many important machine learning problems faced by online streaming services is a need to model how users interact with the content they are served. Unfortunately, there are no public datasets currently available that enable…

信息检索 · 计算机科学 2020-10-16 Brian Brost , Rishabh Mehrotra , Tristan Jehan

Music structure analysis (MSA) systems aim to segment a song recording into non-overlapping sections with useful labels. Previous MSA systems typically predict abstract labels in a post-processing step and require the full context of the…

声音 · 计算机科学 2022-11-30 Ju-Chiang Wang , Jordan B. L. Smith , Yun-Ning Hung

In this paper, we propose an efficient and reproducible deep learning model for musical onset detection (MOD). We first review the state-of-the-art deep learning models for MOD, and identify their shortcomings and challenges: (i) the lack…

声音 · 计算机科学 2018-06-20 Rong Gong , Xavier Serra

Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, which is essential…

多媒体 · 计算机科学 2024-07-19 Arda Senocak , Hyeonggon Ryu , Junsik Kim , Tae-Hyun Oh , Hanspeter Pfister , Joon Son Chung

Incomplete multi-modal emotion recognition (IMER) aims at understanding human intentions and sentiments by comprehensively exploring the partially observed multi-source data. Although the multi-modal data is expected to provide more…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Wen-Jue He , Xiaofeng Zhu , Zheng Zhang

This work presents MAD (Multimodal Affection Dataset), a multimodal emotion dataset designed for affective computing and neurophysiological modeling. MAD is built upon synchronous collection of diverse physiological signals (EEG, ECG, EOG,…

信号处理 · 电气工程与系统科学 2026-03-09 Shengwei Guo , Yunqing Qiao , Wenzhan Zhang , Bo Liu , Yong Wang , Guobing Sun

Representation learning focused on disentangling the underlying factors of variation in given data has become an important area of research in machine learning. However, most of the studies in this area have relied on datasets from the…

机器学习 · 计算机科学 2020-07-31 Ashis Pati , Siddharth Gururani , Alexander Lerch

In this paper, we propose Emotionally paired Music and Image Dataset (EMID), a novel dataset designed for the emotional matching of music and images, to facilitate auditory-visual cross-modal tasks such as generation and retrieval. Unlike…

多媒体 · 计算机科学 2024-08-12 Jialing Zou , Jiahao Mei , Guangze Ye , Tianyu Huai , Qiwei Shen , Daoguo Dong

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

声音 · 计算机科学 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot