中文
相关论文

相关论文: Cross-modal Generative Model for Visual-Guided Bin…

200 篇论文

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

Existing machine learning research has achieved promising results in monaural audio-visual separation (MAVS). However, most MAVS methods purely consider what the sound source is, not where it is located. This can be a problem in VR/AR…

声音 · 计算机科学 2023-11-01 Yuxin Ye , Wenming Yang , Yapeng Tian

Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful consideration of specific objectives and biases that can significantly…

We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art…

声音 · 计算机科学 2024-01-23 Clara Borrelli , James Rae , Dogac Basaran , Matt McVicar , Mehrez Souden , Matthias Mauch

In this work, we build a simple but strong baseline for sounding video generation. Given base diffusion models for audio and video, we integrate them with additional modules into a single model and train it to make the model jointly…

机器学习 · 计算机科学 2025-04-10 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR,…

The paper presents results from a project aiming to create horizontally distributed surround sound sources and virtual sound images as auditory BCI (aBCI) stimuli. The purpose is to create evoked brain wave response patterns depending on…

人机交互 · 计算机科学 2012-10-11 Nozomu Nishikawa , Yoshihiro Matsumoto , Shoji Makino , Tomasz M. Rutkowski

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and…

多媒体 · 计算机科学 2024-10-01 Kun Su , Xiulong Liu , Eli Shlizerman

In most scenarios, conditional image generation can be thought of as an inversion of the image understanding process. Since generic image understanding involves solving multiple tasks, it is natural to aim at generating images via…

计算机视觉与模式识别 · 计算机科学 2022-07-15 Ritika Chakraborty , Nikola Popovic , Danda Pani Paudel , Thomas Probst , Luc Van Gool

It has been advocated that medical imaging systems and reconstruction algorithms should be assessed and optimized by use of objective measures of image quality that quantify the performance of an observer at specific diagnostic tasks. One…

图像与视频处理 · 电气工程与系统科学 2020-06-02 Weimin Zhou , Sayantan Bhadra , Frank J. Brooks , Hua Li , Mark A. Anastasio

We propose a new task named Audio-driven Per-formance Video Generation (APVG), which aims to synthesizethe video of a person playing a certain instrument guided bya given music audio clip. It is a challenging task to gener-ate the…

计算机视觉与模式识别 · 计算机科学 2020-11-06 Hao Zhu , Yi Li , Feixia Zhu , Aihua Zheng , Ran He

While many text-to-audio systems produce monophonic or fixed-stereo outputs, generating audio with user-defined spatial properties remains a challenge. Existing deep learning-based spatialization methods often rely on latent-space…

声音 · 计算机科学 2025-09-16 Tutti Chi , Letian Gao , Yixiao Zhang

Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment…

声音 · 计算机科学 2026-05-06 Jan Melechovsky , Ambuj Mehrish , Abhinaba Roy , Dorien Herremans

Language models pretrained on text-only corpora often struggle with tasks that require auditory commonsense knowledge. Previous work addresses this problem by augmenting the language model to retrieve knowledge from external audio…

计算与语言 · 计算机科学 2025-06-10 Suho Yoo , Hyunjong Ok , Jaeho Lee

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide…

Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lacking bidirectional modality interaction. This results in two key limitations: (i)…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Sen Liang , Cong Wang , Fengbin Guan , Zhentao Yu , Yiting Lu , Yuanzhi Wang , Yuan Zhou , Xin Li , Zhibo Chen

Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematically analyse the performance of existing multimodal fusion…

多媒体 · 计算机科学 2025-10-10 Han Hu , Dongheng Lin , Qiming Huang , Yuqi Hou , Hyung Jin Chang , Jianbo Jiao

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on template-like instruction formats and are restricted to…

声音 · 计算机科学 2025-09-29 Zitong Lan , Yiduo Hao , Mingmin Zhao

In audio-related creative tasks, sound designers often seek to extend and morph different sounds from their libraries. Generative audio models, capable of creating audio using examples as references, offer promising solutions. By masking…

声音 · 计算机科学 2026-02-20 Prem Seetharaman , Oriol Nieto , Justin Salamon

State-of-the-art under-determined audio source separation systems rely on supervised end-end training of carefully tailored neural network architectures operating either in the time or the spectral domain. However, these methods are…

音频与语音处理 · 电气工程与系统科学 2020-05-29 Vivek Narayanaswamy , Jayaraman J. Thiagarajan , Rushil Anirudh , Andreas Spanias