中文
相关论文

相关论文: Audio-Conditioned U-Net for Position Estimation in…

200 篇论文

We propose a method for the blind separation of sounds of musical instruments in audio signals. We describe the individual tones via a parametric model, training a dictionary to capture the relative amplitudes of the harmonics. The model…

音频与语音处理 · 电气工程与系统科学 2021-08-10 Sören Schulze , Johannes Leuschner , Emily J. King

This paper addresses the problem of cross-modal musical piece identification and retrieval: finding the appropriate recording(s) from a database given a sheet music query, and vice versa, working directly with audio and scanned sheet music…

音频与语音处理 · 电气工程与系统科学 2021-05-27 Luis Carvalho , Gerhard Widmer

Deep learning has recently been applied to optical music recognition (OMR). However, currently OMR processing from various sheet music images still lacks precision to be widely applicable. Here, we present an MMdA (Measure-based Multimodal…

计算机视觉与模式识别 · 计算机科学 2021-06-24 Tomoyuki Shishido , Fehmiju Fati , Daisuke Tokushige , Yasuhiro Ono

Music creation is typically composed of two parts: composing the musical score, and then performing the score with instruments to make sounds. While recent work has made much progress in automatic music generation in the symbolic domain,…

声音 · 计算机科学 2018-11-13 Bryan Wang , Yi-Hsuan Yang

Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Recently, several…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Yidi Li , Hong Liu , Bing Yang

Music structure analysis (MSA) methods traditionally search for musically meaningful patterns in audio: homogeneity, repetition, novelty, and segment-length regularity. Hand-crafted audio features such as MFCCs or chromagrams are often used…

音频与语音处理 · 电气工程与系统科学 2022-05-03 Ju-Chiang Wang , Jordan B. L. Smith , Wei-Tsung Lu , Xuchen Song

As artificial intelligence becomes more and more ingrained in daily life, we present a novel system that uses deep learning for music recommendation and emotion-based detection. Through the use of facial recognition and the DeepFace…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Swetha Kambham , Hubert Jhonson , Sai Prathap Reddy Kambham

Music mixing traditionally involves recording instruments in the form of clean, individual tracks and blending them into a final mixture using audio effects and expert knowledge (e.g., a mixing engineer). The automation of music production…

音频与语音处理 · 电气工程与系统科学 2022-08-30 Marco A. Martínez-Ramírez , Wei-Hsiang Liao , Giorgio Fabbro , Stefan Uhlich , Chihiro Nagashima , Yuki Mitsufuji

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Chuang Gan , Deng Huang , Hang Zhao , Joshua B. Tenenbaum , Antonio Torralba

Emotion is a complicated notion present in music that is hard to capture even with fine-tuned feature engineering. In this paper, we investigate the utility of state-of-the-art pre-trained deep audio embedding methods to be used in the…

声音 · 计算机科学 2021-04-15 Eunjeong Koh , Shlomo Dubnov

The digitization of musical scores plays a crucial role in their preservation and accessibility, yet information retrieval still depends mainly on metadata searches, such as by title or composer. Content based search in music score images…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Noelia Luna-Barahona , Antonio Ríos-Vila , Félix Fuentes-Hurtado , David Rizo , Jorge Calvo-Zaragoza

In this work, we present a deep learning-based automatic multitrack music mixing system catered towards live performances. In a live performance, channels are often corrupted with acoustic bleeds of co-located instruments. Moreover,…

音频与语音处理 · 电气工程与系统科学 2026-04-30 Devansh Zurale , Iris Lorente , Michael Lester , Alex Mitchell

The Music Emotion Recognition (MER) field has seen steady developments in recent years, with contributions from feature engineering, machine learning, and deep learning. The landscape has also shifted from audio-centric systems to bimodal…

Visual events are usually accompanied by sounds in our daily lives. We pose the question: Can the machine learn the correspondence between visual scene and the sound, and localize the sound source only by observing sound and visual scene…

计算机视觉与模式识别 · 计算机科学 2019-02-18 Arda Senocak , Tae-Hyun Oh , Junsik Kim , Ming-Hsuan Yang , In So Kweon

We propose an efficient workflow for high-quality offline alignment of in-the-wild performance audio and corresponding sheet music scans (images). Recent work on audio-to-score alignment extends dynamic time warping (DTW) to be…

声音 · 计算机科学 2024-11-13 Irmak Bukey , Michael Feffer , Chris Donahue

A key aspect of machine learning models lies in their ability to learn efficient intermediate features. However, the input representation plays a crucial role in this process, and polyphonic musical scores remain a particularly complex type…

机器学习 · 计算机科学 2021-09-09 Mathieu Prang , Philippe Esling

In this paper we present a new definition of the distortion matrix for a score following framework based on DTW. The proposal consists of arranging the score information in a sequence of note combinations and learning a spectral pattern for…

声音 · 计算机科学 2019-05-30 A. J. Muñoz-Montoro , P. Vera-Candeas , D. Suarez-Dou , R. Cortina

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

The original MV2H metric was designed to evaluate systems which transcribe from an input audio (or MIDI) piece to a complete musical score. However, it requires both the transcribed score and the ground truth score to be time-aligned with…

音频与语音处理 · 电气工程与系统科学 2019-07-09 Andrew McLeod

Music information is often conveyed or recorded across multiple data modalities including but not limited to audio, images, text and scores. However, music information retrieval research has almost exclusively focused on single modality…

声音 · 计算机科学 2021-06-03 Ho-Hsiang Wu , Magdalena Fuentes , Juan P. Bello