中文
相关论文

相关论文: Leveraging Category Information for Single-Frame V…

200 篇论文

Recent years have witnessed the success of deep learning on the visual sound separation task. However, existing works follow similar settings where the training and testing datasets share the same musical instrument categories, which to…

多媒体 · 计算机科学 2022-03-28 Xinchi Zhou , Dongzhan Zhou , Wanli Ouyang , Hang Zhou , Ziwei Liu , Di Hu

Music source separation is the task of separating a mixture of instruments into constituent tracks. Music source separation models are typically trained using only audio data, although additional information can be used to improve the…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Eetu Tunturi , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen

Typical methods for binaural source separation consider only the direct sound as the target signal in a mixture. However, in most scenarios, this assumption limits the source separation performance. It is well known that the early…

声音 · 计算机科学 2019-10-10 Luca Remaggi , Philip J. B. Jackson , Wenwu Wang

The identification of device brands and models plays a pivotal role in the realm of multimedia forensic applications. This paper presents a framework capable of identifying devices using audio, visual content, or a fusion of them. The…

机器学习 · 计算机科学 2024-06-27 Ioannis Tsingalis , Christos Korgialas , Constantine Kotropoulos

Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single speaker and eliminate…

计算机视觉与模式识别 · 计算机科学 2018-02-13 Aviv Gabbay , Ariel Ephrat , Tavi Halperin , Shmuel Peleg

Supervised multi-channel audio source separation requires extracting useful spectral, temporal, and spatial features from the mixed signals. The success of many existing systems is therefore largely dependent on the choice of features used…

声音 · 计算机科学 2018-03-05 Emad M. Grais , Dominic Ward , Mark D. Plumbley

The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independently for each task.…

声音 · 计算机科学 2023-06-01 Shentong Mo , Pedro Morgado

Music demixing is the task of separating different tracks from the given single audio signal into components, such as drums, bass, and vocals from the rest of the accompaniment. Separation of sources is useful for a range of areas,…

声音 · 计算机科学 2024-05-08 Roman Solovyev , Alexander Stempkovskiy , Tatiana Habruseva

Universal source separation (USS) is a fundamental research task for computational auditory scene analysis, which aims to separate mono recordings into individual source tracks. There are three potential challenges awaiting the solution to…

Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank…

音频与语音处理 · 电气工程与系统科学 2018-01-12 Tak-Shing T. Chan , Yi-Hsuan Yang

The goal of the multi-sound source localization task is to localize sound sources from the mixture individually. While recent multi-sound source localization methods have shown improved performance, they face challenges due to their…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Dongjin Kim , Sung Jin Um , Sangmin Lee , Jung Uk Kim

The identification of source cameras from videos, though it is a highly relevant forensic analysis topic, has been studied much less than its counterpart that uses images. In this work we propose a method to identify the source camera of a…

计算机视觉与模式识别 · 计算机科学 2020-12-14 Derrick Timmerman , Swaroop Bennabhaktula , Enrique Alegre , George Azzopardi

Audio source separation is the process of separating a mixture (e.g. a pop band recording) into isolated sounds from individual sources (e.g. just the lead vocals). Deep learning models are the state-of-the-art in source separation, given…

音频与语音处理 · 电气工程与系统科学 2020-07-28 Alisa Liu , Prem Seetharaman , Bryan Pardo

The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cost of fine-grained manual annotations. Recent works attempt…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Kyungbok Lee , You Zhang , Zhiyao Duan

Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach,…

声音 · 计算机科学 2025-11-26 Genís Plaja-Roglans , Yun-Ning Hung , Xavier Serra , Igor Pereira

We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source. This is achieved by combining two distinct models. The first model,…

音频与语音处理 · 电气工程与系统科学 2022-04-13 Kevin Kilgour , Beat Gfeller , Qingqing Huang , Aren Jansen , Scott Wisdom , Marco Tagliasacchi

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Haytham M. Fayek , Anurag Kumar

Associating sound and its producer in complex audiovisual scene is a challenging task, especially when we are lack of annotated training data. In this paper, we present a flexible audiovisual model that introduces a soft-clustering module…

计算机视觉与模式识别 · 计算机科学 2020-01-28 Di Hu , Zheng Wang , Haoyi Xiong , Dong Wang , Feiping Nie , Dejing Dou

The goal of video segmentation is to turn video data into a set of concrete motion clusters that can be easily interpreted as building blocks of the video. There are some works on similar topics like detecting scene cuts in a video, but…

计算机视觉与模式识别 · 计算机科学 2019-03-07 Hajar Sadeghi Sokeh , Vasileios Argyriou , Dorothy Monekosso , Paolo Remagnino

Unsupervised audio-visual source localization aims at localizing visible sound sources in a video without relying on ground-truth localization for training. Previous works often seek high audio-visual similarities for likely positive…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Shentong Mo , Pedro Morgado