中文
相关论文

相关论文: Schrodinger Audio-Visual Editor: Object-Level Audi…

200 篇论文

Given an isolated garment image in a canonical product view and a separate image of a person, the virtual try-on task aims to generate a new image of the person wearing the target garment. Prior virtual try-on works face two major…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Nannan Li , Kevin J. Shih , Bryan A. Plummer

Large language model agents have made strong progress on software engineering, yet current systems suffer from a context coupling problem: the standard code editing interface conflates code inspection, modification planning, and edit…

Audio-Visual Segmentation (AVS) faces a fundamental challenge of effectively aligning audio and visual modalities. While recent approaches leverage foundation models to address data scarcity, they often rely on single-modality knowledge or…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Ziyang Luo , Nian Liu , Xuguang Yang , Salman Khan , Rao Muhammad Anwer , Hisham Cholakkal , Fahad Shahbaz Khan , Junwei Han

The aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Dawei Hao , Yuxin Mao , Bowen He , Xiaodong Han , Yuchao Dai , Yiran Zhong

Recognizing the sounding objects in scenes is a longstanding objective in embodied AI, with diverse applications in robotics and AR/VR/MR. To that end, Audio-Visual Segmentation (AVS), taking as condition an audio signal to identify the…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Artem Sokolov , Swapnil Bhosale , Xiatian Zhu

In this paper our objectives are, first, networks that can embed audio and visual inputs into a common space that is suitable for cross-modal retrieval; and second, a network that can localize the object that sounds in an image, given the…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Relja Arandjelović , Andrew Zisserman

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

多媒体 · 计算机科学 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

Despite the remarkable progress in text-driven video editing, generating coherent non-rigid deformations remains a critical challenge, often plagued by physical distortion and temporal flicker. To bridge this gap, we propose NRVBench, the…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Bingzheng Qu , Kehai Chen , Xuefeng Bai , Jun Yu , Min Zhang

In this paper, we propose two techniques, namely joint modeling and data augmentation, to improve system performances for audio-visual scene classification (AVSC). We employ pre-trained networks trained only on image data sets to extract…

Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address these problems, we…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Gnana Praveen Rajasekhar , Jahangir Alam

Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end, we introduce a model that uses attention to localize and group sound sources, and optical flow to aggregate…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Triantafyllos Afouras , Andrew Owens , Joon Son Chung , Andrew Zisserman

Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get started. Systems that could follow natural language…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Tsu-Jui Fu , Xin Eric Wang , Scott T. Grafton , Miguel P. Eckstein , William Yang Wang

The objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task, it involves a comprehensive consideration of both the data…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Jinxiang Liu , Yu Wang , Chen Ju , Chaofan Ma , Ya Zhang , Weidi Xie

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization.…

音频与语音处理 · 电气工程与系统科学 2025-03-04 Ruijie Tao , Xinyuan Qian , Yidi Jiang , Junjie Li , Jiadong Wang , Haizhou Li

Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Haojie Zheng , Yixin Yang , Siqi Yang , Shuchen Weng , Boxin Shi

The combination of audio and vision has long been a topic of interest in the multi-modal community. Recently, a new audio-visual segmentation (AVS) task has been introduced, aiming to locate and segment the sounding objects in a given…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Shengyi Gao , Zhe Chen , Guo Chen , Wenhai Wang , Tong Lu

Video object removal is a challenging task in video processing that often requires massive human efforts. Given the mask of the foreground object in each frame, the goal is to complete (inpaint) the object region and generate a video…

计算机视觉与模式识别 · 计算机科学 2019-04-16 Ya-Liang Chang , Zhe Yu Liu , Winston Hsu

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Juhyeong Seon , Woobin Im , Sebin Lee , Jumin Lee , Sung-Eui Yoon

Motivated by the superior performance of image diffusion models, more and more researchers strive to extend these models to the text-based video editing task. Nevertheless, current video editing tasks mainly suffer from the dilemma between…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Yutao Chen , Xingning Dong , Tian Gan , Chunluan Zhou , Ming Yang , Qingpei Guo

Visual dubbing, the synchronization of facial movements with new speech, is crucial for making content accessible across different languages, enabling broader global reach. However, current methods face significant limitations. Existing…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Binyamin Manela , Sharon Gannot , Ethan Fetyaya