中文
相关论文

相关论文: Audio-Conditioned U-Net for Position Estimation in…

200 篇论文

The task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the…

音频与语音处理 · 电气工程与系统科学 2018-02-20 Jong Wook Kim , Justin Salamon , Peter Li , Juan Pablo Bello

Ultrasound imaging is challenging to interpret due to non-uniform intensities, low contrast, and inherent artifacts, necessitating extensive training for non-specialists. Advanced representation with clear tissue structure separation could…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Oleksandra Tmenova , Yordanka Velikova , Mahdi Saleh , Nassir Navab

As an important format of multimedia, music has filled almost everyone's life. Automatic analyzing music is a significant step to satisfy people's need for music retrieval and music recommendation in an effortless way. Thereinto, downbeat…

信息检索 · 计算机科学 2019-12-11 Bijue Jia , Jiancheng Lv , Dayiheng Liu

Current models for audio--sheet music retrieval via multimodal embedding space learning use convolutional neural networks with a fixed-size window for the input audio. Depending on the tempo of a query performance, this window captures more…

声音 · 计算机科学 2018-09-18 Matthias Dorfer , Jan Hajič , Gerhard Widmer

Modelling human perception of musical similarity is critical for the evaluation of generative music systems, musicological research, and many Music Information Retrieval tasks. Although human similarity judgments are the gold standard,…

音频与语音处理 · 电气工程与系统科学 2020-06-29 Jeff Ens , Philippe Pasquier

Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment. We extend the application of these models,…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Sooyoung Park , Arda Senocak , Joon Son Chung

We present the Score-based Autoencoder for Multiscale Inference (SAMI), a method for unsupervised representation learning that combines the theoretical frameworks of diffusion models and VAEs. By unifying their respective evidence lower…

机器学习 · 统计学 2025-12-23 Benjamin S. H. Lyo , Eero P. Simoncelli , Cristina Savin

This study introduces RUMAA, a transformer-based framework for music performance analysis that unifies score-to-performance alignment, score-informed transcription, and mistake detection in a near end-to-end manner. Unlike prior methods…

声音 · 计算机科学 2025-07-17 Sungkyun Chang , Simon Dixon , Emmanouil Benetos

In-bed pose estimation has shown value in fields such as hospital patient monitoring, sleep studies, and smart homes. In this paper, we explore different strategies for detecting body pose from highly ambiguous pressure data, with the aid…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Vandad Davoodnia , Saeed Ghorbani , Ali Etemad

Neural Posterior Estimation methods for simulation-based inference can be ill-suited for dealing with posterior distributions obtained by conditioning on multiple observations, as they tend to require a large number of simulator calls to…

机器学习 · 计算机科学 2023-07-11 Tomas Geffner , George Papamakarios , Andriy Mnih

This paper presents the first step in a research project situated within the field of musical agents. The objective is to achieve, through training, the tuning of the desired musical relationship between a live musical input and a real-time…

声音 · 计算机科学 2025-10-01 Balthazar Bujard , Jérôme Nika , Fédéric Bevilacqua , Nicolas Obin

Proposed in Hyv\"arinen (2005), score matching is a parameter estimation procedure that does not require computation of distributional normalizing constants. In this work we utilize the geometric median of means to develop a robust score…

机器学习 · 统计学 2025-06-23 Richard Schwank , Andrew McCormack , Mathias Drton

This paper proposes a novel approach to map-based navigation system for unmanned aircraft. The proposed system attempts label-to-label matching, not image-to-image matching, between aerial images and a map database. The ground objects can…

计算机视觉与模式识别 · 计算机科学 2022-05-31 Youngjoo Kim

Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, \ie synthesizing missing audio segments that correspond to their accompanying…

计算机视觉与模式识别 · 计算机科学 2019-10-25 Hang Zhou , Ziwei Liu , Xudong Xu , Ping Luo , Xiaogang Wang

The area of computer vision is one of the most discussed topics amongst many scholars, and stereo matching is its most important sub fields. After the parallax map is transformed into a depth map, it can be applied to many intelligent…

计算机视觉与模式识别 · 计算机科学 2021-05-25 Hewei Wang , Muhammad Salman Pathan , Soumyabrata Dev

Video stereo matching is the task of estimating consistent disparity maps from rectified stereo videos. There is considerable scope for improvement in both datasets and methods within this area. Recent learning-based methods often focus on…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Junpeng Jing , Ye Mao , Anlan Qiu , Krystian Mikolajczyk

Despite recent advancements in text-to-image diffusion models facilitating various image editing techniques, complex text prompts often lead to an oversight of some requests due to a bottleneck in processing text information. To tackle this…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Hangeol Chang , Jinho Chang , Jong Chul Ye

This paper offers a precise, formal definition of an audio-to-score alignment. While the concept of an alignment is intuitively grasped, this precision affords us new insight into the evaluation of audio-to-score alignment algorithms.…

声音 · 计算机科学 2020-10-01 John Thickstun , Jennifer Brennan , Harsh Verma

Stereo matching provides depth estimation from binocular images for downstream applications. These applications mostly take video streams as input and require temporally consistent depth maps. However, existing methods mainly focus on the…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Jiaxi Zeng , Chengtang Yao , Yuwei Wu , Yunde Jia

With the development of deep learning and artificial intelligence, audio synthesis has a pivotal role in the area of machine learning and shows strong applicability in the industry. Meanwhile, significant efforts have been dedicated by…

音频与语音处理 · 电气工程与系统科学 2021-08-03 Zhaofeng Shi