中文
相关论文

相关论文: HarmonySet: A Comprehensive Dataset for Understand…

200 篇论文

Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evaluation framework that combines: (1) a multi-source…

声音 · 计算机科学 2026-01-01 Shuyang Liu , Yuan Jin , Rui Lin , Shizhe Chen , Junyu Dai , Tao Jiang

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Tao Feng , Yifan Xie , Xun Guan , Jiyuan Song , Zhou Liu , Fei Ma , Fei Yu

Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban…

计算机视觉与模式识别 · 计算机科学 2016-04-08 Marius Cordts , Mohamed Omran , Sebastian Ramos , Timo Rehfeld , Markus Enzweiler , Rodrigo Benenson , Uwe Franke , Stefan Roth , Bernt Schiele

Short-video platforms show an increasing impact on people's daily lives nowadays, with billions of active users spending plenty of time each day. The interactions between users and online platforms give rise to many scientific problems…

多媒体 · 计算机科学 2025-02-11 Yu Shang , Chen Gao , Nian Li , Yong Li

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-dance co-generation,…

Synthetic data has emerged as a promising source for 3D human research as it offers low-cost access to large-scale human datasets. To advance the diversity and annotation quality of human models, we introduce a new synthetic dataset,…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Zhitao Yang , Zhongang Cai , Haiyi Mei , Shuai Liu , Zhaoxi Chen , Weiye Xiao , Yukun Wei , Zhongfei Qing , Chen Wei , Bo Dai , Wayne Wu , Chen Qian , Dahua Lin , Ziwei Liu , Lei Yang

Quantitative analysis of commonalities and differences between recorded music performances is an increasingly common task in computational musicology. A typical scenario involves manual annotation of different recordings of the same piece…

多媒体 · 计算机科学 2020-09-28 Thassilo Gadermaier , Gerhard Widmer

In this document, we introduce a new dataset designed for training machine learning models of symbolic music data. Five datasets are provided, one of which is from a newly collected corpus of 20K midi files. We describe our preprocessing…

声音 · 计算机科学 2016-06-09 Christian Walder

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate…

计算与语言 · 计算机科学 2022-06-20 Josiah Wang , Pranava Madhyastha , Josiel Figueiredo , Chiraag Lala , Lucia Specia

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Thanadol Singkhornart , Olarik Surinta

With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap,…

Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be…

The main challenges of Optical Music Recognition (OMR) come from the nature of written music, its complexity and the difficulty of finding an appropriate data representation. This paper provides a first look at DoReMi, an OMR dataset that…

信息检索 · 计算机科学 2021-07-19 Elona Shatri , György Fazekas

We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Dahun Kim , AJ Piergiovanni , Ganesh Mallya , Anelia Angelova

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

声音 · 计算机科学 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang

Humans often experience not just a single basic emotion at a time, but rather a blend of several emotions with varying salience. Despite the importance of such blended emotions, most video-based emotion recognition approaches are designed…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Tim Lachmann , Alexandra Israelsson , Christina Tornberg , Teimuraz Saghinadze , Michal Balazia , Philipp Müller , Petri Laukka

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment…

计算机视觉与模式识别 · 计算机科学 2020-11-05 Soo-Whan Chung , Joon Son Chung , Hong-Goo Kang

Information Synchronization of semi-structured data across languages is challenging. For instance, Wikipedia tables in one language should be synchronized across languages. To address this problem, we introduce a new dataset InfoSyncC and a…

计算与语言 · 计算机科学 2023-07-10 Siddharth Khincha , Chelsi Jain , Vivek Gupta , Tushar Kataria , Shuo Zhang

Effective human-AI interaction relies on AI's ability to accurately perceive and interpret human emotions. Current benchmarks for vision and vision-language models are severely limited, offering a narrow emotional spectrum that overlooks…

‹ 上一页 1 8 9 10 下一页 ›