English
Related papers

Related papers: SpatialV2A: Visual-Guided High-fidelity Spatial Au…

200 papers

While many text-to-audio systems produce monophonic or fixed-stereo outputs, generating audio with user-defined spatial properties remains a challenge. Existing deep learning-based spatialization methods often rely on latent-space…

Sound · Computer Science 2025-09-16 Tutti Chi , Letian Gao , Yixiao Zhang

We present Sat2Sound, a unified multimodal framework for geospatial soundscape understanding, designed to predict and map the distribution of sounds across the Earth's surface. Existing methods for this task rely on paired satellite images…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Subash Khanal , Srikumar Sastry , Aayush Dhakal , Adeel Ahmad , Abby Stylianou , Nathan Jacobs

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework, where the input and output of the system are multimodal (i.e., audio and visual speech). With the proposed AV2AV, two key…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Jeongsoo Choi , Se Jin Park , Minsu Kim , Yong Man Ro

In the development of spatial audio technologies, reliable and shared methods for evaluating audio quality are essential. Listening tests are currently the standard but remain costly in terms of time and resources. Several models predicting…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Adrien Llave , Emma Granier , Grégory Pallone

Video-to-Video synthesis (Vid2Vid) has achieved remarkable results in generating a photo-realistic video from a sequence of semantic maps. However, this pipeline suffers from high computational cost and long inference latency, which largely…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Long Zhuo , Guangcong Wang , Shikai Li , Wayne Wu , Ziwei Liu

Recent advances in visually-induced audio generation are based on sampling short, low-fidelity, and one-class sounds. Moreover, sampling 1 second of audio from the state-of-the-art model takes minutes on a high-end GPU. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Vladimir Iashin , Esa Rahtu

Spatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Xiaoyang Huang , Yanjun Wang , Yang Liu , Bingbing Ni , Wenjun Zhang , Jinxian Liu , Teng Li

We present an end-to-end binaural audio rendering approach (Listen2Scene) for virtual reality (VR) and augmented reality (AR) applications. We propose a novel neural-network-based binaural sound propagation method to generate acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-09 Anton Ratnarajah , Dinesh Manocha

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models often lack the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Qi Feng

Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. Previous benchmarks primarily focus on audio-video temporal…

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Zhaokai Wang , Chenxi Bao , Le Zhuo , Jingrui Han , Yang Yue , Yihong Tang , Victor Shea-Jay Huang , Yue Liao

Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision.…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Hang Zhou , Xudong Xu , Dahua Lin , Xiaogang Wang , Ziwei Liu

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yan-Bo Lin , Jonah Casebeer , Long Mai , Aniruddha Mahapatra , Gedas Bertasius , Nicholas J. Bryan

The study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-18 Shoichi Koyama , Enzo De Sena , Prasanga Samarasinghe , Mark R. P. Thomas , Fabio Antonacci

Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Shentong Mo , Haofan Wang

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Chao Huang , Ruohan Gao , J. M. F. Tsang , Jan Kurcius , Cagdas Bilen , Chenliang Xu , Anurag Kumar , Sanjeel Parekh

We propose Im2Wav, an image guided open-domain audio generation system. Given an input image or a sequence of images, Im2Wav generates a semantically relevant sound. Im2Wav is based on two Transformer language models, that operate over a…

Sound · Computer Science 2023-02-28 Roy Sheffer , Yossi Adi

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our…

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…

Multimedia · Computer Science 2023-08-21 Sung Jin Um , Dongjin Kim , Jung Uk Kim

Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial…

Sound · Computer Science 2025-04-28 Ayushi Mishra , Yang Bai , Priyadarshan Narayanasamy , Nakul Garg , Nirupam Roy