English
Related papers

Related papers: FoleySpace: Vision-Aligned Binaural Spatial Audio …

200 papers

Audio-visual saliency prediction aims to mimic human visual attention by identifying salient regions in videos through the integration of both visual and auditory information. Although visual-only approaches have significantly advanced,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Kiana Hooshanfar , Alireza Hosseini , Ahmad Kalhor , Babak Nadjar Araabi

Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible motion-sound relations remains challenging. Existing methods often produce object motions…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Junchao Liao , Zhenghao Zhang , Xiangyu Meng , Litao Li , Ziying Zhang , Siyu Zhu , Long Qin , Weizhi Wang

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

Segment Anything Model 2 (SAM2) exhibits strong generalisation for promptable segmentation in video clips; however, its integration with the audio modality remains underexplored. Existing approaches either convert audio into visual prompts…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yuyuan Liu , Yuanhong Chen , Chong Wang , Junlin Han , Junde Wu , Can Peng , Jingkun Chen , Yu Tian , Gustavo Carneiro

Current open-source diffusion models struggle to generate stable and synchronized audio-visual content, particularly in scenarios demanding complex semantic reasoning. The root cause is that existing methods rely on coarse text embeddings…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Shuyuan Tu , Qi Tian , Zihan Yang , Yue Wu , Xintong Han , Weijie Kong , Jiangfeng Xiong , Jian-Wei Zhang , Zhao Zhong , Liefeng Bo , Zuxuan Wu , Yu-Gang Jiang

There has been a growing interest in the task of generating sound for silent videos, primarily because of its practicality in streamlining video post-production. However, existing methods for video-sound generation attempt to directly…

Multimedia · Computer Science 2024-04-04 Zhifeng Xie , Shengye Yu , Qile He , Mengtian Li

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

Generating high-quality stereo videos requires consistent depth perception and temporal coherence across frames. Despite advances in image and video synthesis using diffusion models, producing high-quality stereo videos remains a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Jian Shi , Qian Wang , Zhenyu Li , Wenqing Cui , Ramzi Idoughi , Peter Wonka

Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematically analyse the performance of existing multimodal fusion…

Multimedia · Computer Science 2025-10-10 Han Hu , Dongheng Lin , Qiming Huang , Yuqi Hou , Hyung Jin Chang , Jianbo Jiao

The widespread application of AIGC contents has brought not only unprecedented opportunities, but also potential security concerns, e.g., audio-visual deepfakes. Therefore, it is of great importance to develop an effective and generalizable…

Multimedia · Computer Science 2025-11-25 Fan Nie , Jiangqun Ni , Jian Zhang , Bin Zhang , Weizhe Zhang , Bin Li

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incoherent content that…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zhexiao Xiong , Yizhi Song , Liu He , Wei Xiong , Yu Yuan , Feng Qiao , Nathan Jacobs

Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-31 Nils L. Westhausen , Hendrik Kayser , Theresa Jansen , Bernd T. Meyer

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech…

Computation and Language · Computer Science 2025-04-29 Tuochao Chen , Qirui Wang , Runlin He , Shyam Gollakota

This paper presents a novel framework for converting 2D videos to immersive stereoscopic 3D, addressing the growing demand for 3D content in immersive experience. Leveraging foundation models as priors, our approach overcomes the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Sijie Zhao , Wenbo Hu , Xiaodong Cun , Yong Zhang , Xiaoyu Li , Zhe Kong , Xiangjun Gao , Muyao Niu , Ying Shan

Neural audio codecs have been widely studied for mono and stereo signals, but spatial audio remains largely unexplored. We present the first discrete neural spatial audio codec for first-order ambisonics (FOA). Building on the WavTokenizer…

Sound · Computer Science 2025-10-28 Parthasaarathy Sudarsanam , Sebastian Braun , Hannes Gamper

While 3D Gaussian representations (3DGS) have proven effective for modeling the geometry and appearance of objects, their potential for capturing other physical attributes-such as sound-remains largely unexplored. In this paper, we present…

Sound · Computer Science 2025-07-29 Chunshi Wang , Hongxing Li , Yawei Luo

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound. Inspired by the challenges of creating a soundtrack for a video that differs…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Yuexi Du , Ziyang Chen , Justin Salamon , Bryan Russell , Andrew Owens

Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visual actions that are aligned with them, otherwise unnatural…

Sound · Computer Science 2024-07-16 Santiago Pascual , Chunghsin Yeh , Ioannis Tsiamas , Joan Serrà

Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on…

Sound · Computer Science 2026-01-28 Congyi Fan , Jian Guan , Youtian Lin , Dongli Xu , Tong Ye , Qiaoxi Zhu , Pengming Feng , Wenwu Wang

We present Sat2Sound, a unified multimodal framework for geospatial soundscape understanding, designed to predict and map the distribution of sounds across the Earth's surface. Existing methods for this task rely on paired satellite images…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Subash Khanal , Srikumar Sastry , Aayush Dhakal , Adeel Ahmad , Abby Stylianou , Nathan Jacobs
‹ Prev 1 8 9 10 Next ›