中文
相关论文

相关论文: HarmonySet: A Comprehensive Dataset for Understand…

200 篇论文

Representation learning focused on disentangling the underlying factors of variation in given data has become an important area of research in machine learning. However, most of the studies in this area have relied on datasets from the…

机器学习 · 计算机科学 2020-07-31 Ashis Pati , Siddharth Gururani , Alexander Lerch

Story video-text alignment, a core task in computational story understanding, aims to align video clips with corresponding sentences in their descriptions. However, progress on the task has been held back by the scarcity of manually…

计算与语言 · 计算机科学 2024-10-04 Yidan Sun , Jianfei Yu , Boyang Li

Music scores are written representations of music and contain rich information about musical components. The visual information on music scores includes notes, rests, staff lines, clefs, dynamics, and articulations. This visual information…

多媒体 · 计算机科学 2024-06-18 Yuheng Lin , Zheqi Dai , Qiuqiang Kong

Thanks to the substantial and explosively inscreased instructional videos on the Internet, novices are able to acquire knowledge for completing various tasks. Over the past decade, growing efforts have been devoted to investigating the…

计算机视觉与模式识别 · 计算机科学 2020-03-23 Yansong Tang , Jiwen Lu , Jie Zhou

Music arrangement generation is a subtask of automatic music generation, which involves reconstructing and re-conceptualizing a piece with new compositional techniques. Such a generation process inevitably requires reference from the…

声音 · 计算机科学 2020-08-18 Ziyu Wang , Ke Chen , Junyan Jiang , Yiyi Zhang , Maoran Xu , Shuqi Dai , Xianbin Gu , Gus Xia

Understanding relations between objects is crucial for understanding the semantics of a visual scene. It is also an essential step in order to bridge visual and language models. However, current state-of-the-art computer vision models still…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Palaash Agrawal , Haidi Azaman , Cheston Tan

We present a new large-scale emotion-labeled symbolic music dataset consisting of 12k MIDI songs. To create this dataset, we first trained emotion classification models on the GoEmotions dataset, achieving state-of-the-art results with a…

音频与语音处理 · 电气工程与系统科学 2023-07-28 Serkan Sulun , Pedro Oliveira , Paula Viana

This work present a music dataset named MusicTM-Dataset, which is utilized in improving the representation learning ability of different types of cross-modal retrieval (CMR). Little large music dataset including three modalities is…

声音 · 计算机科学 2021-05-10 Donghuo Zeng , Yi Yu , Keizo Oyama

Harmony in visual compositions is a concept that cannot be defined or easily expressed mathematically, even by humans. The goal of the research described in this paper was to find a numerical representation of artistic compositions with…

计算机视觉与模式识别 · 计算机科学 2020-12-11 Adam Vandor , Marie van Vollenhoven , Gerhard Weiss , Gerasimos Spanakis

Understanding movies and their structural patterns is a crucial task in decoding the craft of video editing. While previous works have developed tools for general analysis, such as detecting characters or recognizing cinematography…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Alejandro Pardo , Fabian Caba Heilbron , Juan León Alcázar , Ali Thabet , Bernard Ghanem

Many social media users prefer consuming content in the form of videos rather than text. However, in order for content creators to produce videos with a high click-through rate, much editing is needed to match the footage to the music. This…

机器学习 · 计算机科学 2022-01-03 Chin-Tung Lin , Mu Yang

While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-rounded common-sense reasoning akin to human visual…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Kate Sanders , Benjamin Van Durme

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

We present HumanEdit, a high-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through open-form language instructions. Previous large-scale editing…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Jinbin Bai , Wei Chow , Ling Yang , Xiangtai Li , Juncheng Li , Hanwang Zhang , Shuicheng Yan

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Teng Hu , Zhentao Yu , Guozhen Zhang , Zihan Su , Zhengguang Zhou , Youliang Zhang , Yuan Zhou , Qinglin Lu , Ran Yi

This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio,…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Songtao Li , Hao Tang

Music can be represented in multiple forms, such as in the audio form as a recording of a performance, in the symbolic form as a computer readable score, or in the image form as a scan of the sheet music. Music synchronisation provides a…

声音 · 计算机科学 2022-06-02 Ruchit Agrawal

Synthesizing human motion with a global structure, such as a choreography, is a challenging task. Existing methods tend to concentrate on local smooth pose transitions and neglect the global context or the theme of the motion. In this work,…

The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing mistake detection in the context of music pedagogy,…

音频与语音处理 · 电气工程与系统科学 2026-02-09 Sumit Kumar , Suraj Jaiswal , Parampreet Singh , Vipul Arora

Video content comprehension is essential for various applications, ranging from video analysis to interactive systems. Despite advancements in large-scale vision-language models (VLMs), these models often struggle to capture the nuanced,…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Shuyi Zhang , Xiaoshuai Hao , Yingbo Tang , Lingfeng Zhang , Pengwei Wang , Zhongyuan Wang , Hongxuan Ma , Shanghang Zhang