中文
相关论文

相关论文: Leveraging Category Information for Single-Frame V…

200 篇论文

Source separation (SS) aims to separate individual sources from an audio recording. Sound event detection (SED) aims to detect sound events from an audio recording. We propose a joint separation-classification (JSC) model trained only on…

声音 · 计算机科学 2019-12-10 Qiuqiang Kong , Yong Xu , Wenwu Wang , Mark D. Plumbley

Separating the individual elements in a musical mixture is an essential process for music analysis and practice. While this is generally addressed using neural networks optimized to mask or transform the time-frequency representation of a…

声音 · 计算机科学 2025-11-27 Genís Plaja-Roglans , Yun-Ning Hung , Xavier Serra , Igor Pereira

Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel categories in real-world videos. To address this…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Haochen Wang , Cilin Yan , Shuai Wang , Xiaolong Jiang , XU Tang , Yao Hu , Weidi Xie , Efstratios Gavves

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

计算机视觉与模式识别 · 计算机科学 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

In recent studies, diffusion models have shown promise as priors for solving audio inverse problems. These models allow us to sample from the posterior distribution of a target signal given an observed signal by manipulating the diffusion…

音频与语音处理 · 电气工程与系统科学 2024-10-22 Chin-Yun Yu , Emilian Postolache , Emanuele Rodolà , György Fazekas

The integration of additional side information to improve music source separation has been investigated numerous times, e.g., by adding features to the input or by adding learning targets in a multi-task learning scenario. These approaches,…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Yun-Ning Hung , Alexander Lerch

Visual semantic segmentation aims at separating a visual sample into diverse blocks with specific semantic attributes and identifying the category for each block, and it plays a crucial role in environmental perception. Conventional…

计算机视觉与模式识别 · 计算机科学 2022-11-16 Wenqi Ren , Yang Tang , Qiyu Sun , Chaoqiang Zhao , Qing-Long Han

The state of the art in music source separation employs neural networks trained in a supervised fashion on multi-track databases to estimate the sources from a given mixture. With only few datasets available, often extensive data…

机器学习 · 计算机科学 2018-04-09 Daniel Stoller , Sebastian Ewert , Simon Dixon

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

Multimedia Forensics allows to determine whether videos or images have been captured with the same device, and thus, eventually, by the same person. Currently, the most promising technology to achieve this task, exploits the unique traces…

多媒体 · 计算机科学 2017-05-05 Massimo Iuliani , Marco Fontani , Dasara Shullani , Alessandro Piva

The aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Dawei Hao , Yuxin Mao , Bowen He , Xiaodong Han , Yuchao Dai , Yiran Zhong

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective.…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Arda Senocak , Hyeonggon Ryu , Junsik Kim , Tae-Hyun Oh , Hanspeter Pfister , Joon Son Chung

The objective of this work is to localize sound sources that are visible in a video without using manual annotations. Our key technical contribution is to show that, by training the network to explicitly discriminate challenging image…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Honglie Chen , Weidi Xie , Triantafyllos Afouras , Arsha Nagrani , Andrea Vedaldi , Andrew Zisserman

The performance of audio source separation from underdetermined convolutive mixture assuming known mixing filters can be significantly improved by using an analysis sparse prior optimized by a reweighting l1 scheme and a wideband…

声音 · 计算机科学 2015-06-18 Simon Arberet , Pierre Vandergheynst

Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research about VC, mainly focusing on clean voices, pay rare…

音频与语音处理 · 电气工程与系统科学 2022-11-08 Jixun Yao , Yi Lei , Qing Wang , Pengcheng Guo , Ziqian Ning , Lei Xie , Hai Li , Junhui Liu , Danming Xie

In this paper, we study whether music source separation can be used as a pre-training strategy for music representation learning, targeted at music classification tasks. To this end, we first pre-train U-Net networks under various music…

音频与语音处理 · 电气工程与系统科学 2024-04-24 Christos Garoufis , Athanasia Zlatintsi , Petros Maragos

Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Sung Jin Um , Dongjin Kim , Sangmin Lee , Jung Uk Kim

Estimation of the location of sound sources is usually done using microphone arrays. Such settings provide an environment where we know the difference between the received signals among different microphones in the terms of phase or…

声音 · 计算机科学 2026-01-08 Helena Peic Tukuljac , Herve Lissek , Pierre Vandergheynst

Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Jinxing Zhou , Liang Zheng , Yiran Zhong , Shijie Hao , Meng Wang