English
Related papers

Related papers: MOSA: Music Motion with Semantic Annotation Datase…

200 papers

Many motion-centric video analysis tasks, such as atomic actions, detecting atypical motor behavior in individuals with autism, or analyzing articulatory motion in real-time MRI of human speech, require efficient and interpretable temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Hong Nguyen , Dung Tran , Hieu Hoang , Phong Nguyen , Shrikanth Narayanan

Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music…

Sound · Computer Science 2026-03-09 Shuyu Li , Shulei Ji , Zihao Wang , Songruoyao Wu , Jiaxing Yu , Kejun Zhang

Multimodal sentiment analysis (MSA), which supposes to improve text-based sentiment analysis with associated acoustic and visual modalities, is an emerging research area due to its potential applications in Human-Computer Interaction (HCI).…

Multimedia · Computer Science 2022-09-07 Yihe Liu , Ziqi Yuan , Huisheng Mao , Zhiyun Liang , Wanqiuyue Yang , Yuanzhe Qiu , Tie Cheng , Xiaoteng Li , Hua Xu , Kai Gao

This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two existing tracking datasets using a semi-automatic annotation…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Paul Voigtlaender , Michael Krause , Aljosa Osep , Jonathon Luiten , Berin Balachandar Gnana Sekar , Andreas Geiger , Bastian Leibe

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Arda Senocak , Hyeonggon Ryu , Junsik Kim , Tae-Hyun Oh , Hanspeter Pfister , Joon Son Chung

Music scores are written representations of music and contain rich information about musical components. The visual information on music scores includes notes, rests, staff lines, clefs, dynamics, and articulations. This visual information…

Multimedia · Computer Science 2024-06-18 Yuheng Lin , Zheqi Dai , Qiuqiang Kong

Multimodal fine-grained sentiment analysis has recently attracted increasing attention due to its broad applications. However, the existing multimodal fine-grained sentiment datasets most focus on annotating the fine-grained elements in…

Computation and Language · Computer Science 2022-06-29 Hao Yang , Yanyan Zhao , Jianwei Liu , Yang Wu , Bing Qin

Singing voice generation progresses rapidly, yet evaluating singing quality remains a critical challenge. Human subjective assessment, typically in the form of listening tests, is costly and time consuming, while existing objective metrics…

Sound · Computer Science 2026-01-28 Yuxun Tang , Lan Liu , Wenhao Feng , Yiwen Zhao , Jionghao Han , Yifeng Yu , Jiatong Shi , Qin Jin

Music Emotion Recogniser (MER) research faces challenges due to limited high-quality annotated datasets and difficulties in addressing cross-track feature drift. This work presents two primary contributions to address these issues.…

Sound · Computer Science 2025-12-18 Qilin Li , C. L. Philip Chen , Tong Zhang

Musical dynamics form a core part of expressive singing voice performances. However, automatic analysis of musical dynamics for singing voice has received limited attention partly due to the scarcity of suitable datasets and a lack of clear…

Sound · Computer Science 2024-10-29 Jyoti Narang , Nazif Can Tamer , Viviana De La Vega , Xavier Serra

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

Music information is often conveyed or recorded across multiple data modalities including but not limited to audio, images, text and scores. However, music information retrieval research has almost exclusively focused on single modality…

Sound · Computer Science 2021-06-03 Ho-Hsiang Wu , Magdalena Fuentes , Juan P. Bello

Achieving level-5 driving automation in autonomous vehicles necessitates a robust semantic visual perception system capable of parsing data from different sensors across diverse conditions. However, existing semantic perception datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Tim Brödermann , David Bruggemann , Christos Sakaridis , Kevin Ta , Odysseas Liagouris , Jason Corkill , Luc Van Gool

Achieving realistic, vivid, and human-like synthesized conversational gestures conditioned on multi-modal data is still an unsolved problem due to the lack of available datasets, models and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Haiyang Liu , Zihao Zhu , Naoya Iwamoto , Yichen Peng , Zhengqing Li , You Zhou , Elif Bozkurt , Bo Zheng

Ornamentations, embellishments, or microtonal inflections are essential to melodic expression across many musical traditions, adding depth, nuance, and emotional impact to performances. Recognizing ornamentations in singing voices is key to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-09 Sumit Kumar , Parampreet Singh , Vipul Arora

Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance,…

Graphics · Computer Science 2025-11-04 Shuyang Xu , Zhiyang Dou , Mingyi Shi , Liang Pan , Leo Ho , Jingbo Wang , Yuan Liu , Cheng Lin , Yuexin Ma , Wenping Wang , Taku Komura

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data…

Graphics · Computer Science 2025-08-26 Prerit Gupta , Jason Alexander Fotso-Puepi , Zhengyuan Li , Jay Mehta , Aniket Bera

Multi-modal deep learning techniques for matching free-form text with music have shown promising results in the field of Music Information Retrieval (MIR). Prior work is often based on large proprietary data while publicly available…

Computation and Language · Computer Science 2024-04-18 Benno Weck , Holger Kirchhoff , Peter Grosche , Xavier Serra

Data movement between main memory and the CPU is a major bottleneck in parallel data-intensive applications. In response, researchers have proposed using compilers and intermediate representations (IRs) that apply optimizations such as loop…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-09-20 Shoumik Palkar , Matei Zaharia

This thesis combines audio-analysis with computer vision to approach Music Information Retrieval (MIR) tasks from a multi-modal perspective. This thesis focuses on the information provided by the visual layer of music videos and how it can…

Multimedia · Computer Science 2020-02-04 Alexander Schindler