中文
相关论文

相关论文: Osu2MIR: Beat Tracking Dataset Derived From Osu! D…

200 篇论文

Connecting large libraries of digitized audio recordings to their corresponding sheet music images has long been a motivation for researchers to develop new cross-modal retrieval systems. In recent years, retrieval systems based on…

信息检索 · 计算机科学 2019-06-27 Stefan Balke , Matthias Dorfer , Luis Carvalho , Andreas Arzt , Gerhard Widmer

In this late-breaking abstract we propose a modified approach for beat tracking evaluation which poses the problem in terms of the effort required to transform a sequence of beat detections such that they maximise the well-known F-measure…

声音 · 计算机科学 2020-11-04 A. Sá Pinto , I. Domingues , M. E. P. Davies

We introduce a new high resolution, high frame rate stereo video dataset, which we call SPIN, for tracking and action recognition in the game of ping pong. The corpus consists of ping pong play with three main annotation streams that can be…

计算机视觉与模式识别 · 计算机科学 2019-12-16 Steven Schwarcz , Peng Xu , David D'Ambrosio , Juhana Kangaspunta , Anelia Angelova , Huong Phan , Navdeep Jaitly

We introduce Seshat, a new, simple and open-source software to efficiently manage annotations of speech corpora. The Seshat software allows users to easily customise and manage annotations of large audio corpora while ensuring compliance…

To extract information at scale, researchers increasingly apply semantic segmentation techniques to remotely-sensed imagery. While fully-supervised learning enables accurate pixel-wise segmentation, compiling the exhaustive datasets…

计算机视觉与模式识别 · 计算机科学 2020-05-28 Simone Fobi , Terence Conlon , Jayant Taneja , Vijay Modi

Objective evaluation (OE) is essential to artificial music, but it's often very hard to determine the quality of OEs. Hitherto, subjective evaluation (SE) remains reliable and prevailing but suffers inevitable disadvantages that OEs may…

声音 · 计算机科学 2021-08-31 Songhe Wang , Zheng Bao , Jingtong E

Onset detection is the process of identifying the start points of musical note events within an audio recording. While the detection of percussive onsets is often considered a solved problem, soft onsets-as found in string instrument…

音频与语音处理 · 电气工程与系统科学 2022-11-17 Maciej Tomczak , Min Susan Li , Adrian Bradbury , Mark Elliott , Ryan Stables , Maria Witek , Tom Goodman , Diar Abdlkarim , Massimiliano Di Luca , Alan Wing , Jason Hockman

We present a computational assessment system that promotes the learning of basic rhythmic patterns. The system is capable of generating multiple rhythmic patterns with increasing complexity within various cycle lengths. For a generated…

多媒体 · 计算机科学 2021-09-10 Noel Alben , Ranjani H. G

We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visualize rubato (temporally expressive playing), a learner diagnose ensemble precision and…

声音 · 计算机科学 2026-05-26 Nazif Can Tamer , Victoria Ebert , Guang Yang , Noah A. Smith

This paper provides an outline of the algorithms submitted for the WSDM Cup 2019 Spotify Sequential Skip Prediction Challenge (team name: mimbres). In the challenge, complete information including acoustic features and user interaction logs…

信息检索 · 计算机科学 2020-10-27 Sungkyun Chang , Seungjin Lee , Kyogu Lee

Understanding human behavior is key for robots and intelligent systems that share a space with people. Accordingly, research that enables such systems to perceive, track, learn and predict human behavior as well as to plan and interact with…

In this paper, we introduce score difficulty classification as a sub-task of music information retrieval (MIR), which may be used in music education technologies, for personalised curriculum generation, and score retrieval. We introduce a…

声音 · 计算机科学 2022-03-25 Pedro Ramoneda , Nazif Can Tamer , Vsevolod Eremenko , Xavier Serra , Marius Miron

Rhythm is a fundamental aspect of human behaviour, present from infancy and deeply embedded in cultural practices. Rhythm anticipation is a spontaneous cognitive process that typically occurs before the onset of actual beats. While most…

神经元与认知 · 定量生物学 2025-03-18 Zhongju Yuan , Geraint Wiggins , Dick Botteldooren

Existing manual labeling of micro-expressions is subject to errors in accuracy, especially in cross-cultural scenarios where deviation in labeling of key frames is more prominent. To address this issue, this paper presents a novel Global…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Feng Liu , Bingyu Nan , Xuezhong Qian , Xiaolan Fu

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Qiang Huang , Thomas Hain

Optical Music Recognition (OMR) aims to convert printed or handwritten music score images into editable symbolic representations. This paper presents an end-to-end OMR framework that combines residual bottleneck convolutions with…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Junwen Ma , Huhu Xue , Xingyuan Zhao , and Weicheng Fu

Multi-Object Tracking (MOT) aims to detect and associate all targets of given classes across frames. Current dominant solutions, e.g. ByteTrack and StrongSORT++, follow the hybrid pipeline, which first accomplish most of the associations in…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Yunhao Du , Zhicheng Zhao , Fei Su

Personalized recommendation on new track releases has always been a challenging problem in the music industry. To combat this problem, we first explore user listening history and demographics to construct a user embedding representing the…

声音 · 计算机科学 2021-03-31 Ke Chen , Beici Liang , Xiaoshuan Ma , Minwei Gu

Emotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct".…

Environmental sounds like footsteps, keyboard typing, or dog barking carry rich information and emotional context, making them valuable for designing haptics in user applications. Existing audio-to-vibration methods, however, rely on…

人机交互 · 计算机科学 2026-01-27 Yinan Li , Hasti Seifi