English
Related papers

Related papers: Sat2Sound: A Unified Framework for Zero-Shot Sound…

200 papers

Satellite image classification is a challenging problem that lies at the crossroads of remote sensing, computer vision, and machine learning. Due to the high variability inherent in satellite data, most of the current object classification…

Computer Vision and Pattern Recognition · Computer Science 2015-09-14 Saikat Basu , Sangram Ganguly , Supratik Mukhopadhyay , Robert DiBiano , Manohar Karki , Ramakrishna Nemani

We introduce Sky2Ground, a three-view dataset designed for varying altitude camera localization, correspondence learning, and reconstruction. The dataset combines structured synthetic imagery with real, in-the-wild images, providing both…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Zengyan Wang , Sirshapan Mitra , Rajat Modi , Grace Lim , Yogesh Rawat

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

Sound · Computer Science 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang

This paper presents a new approach for synthesizing a novel street-view panorama given an overhead satellite image. Taking a small satellite image patch as input, our method generates a Google's omnidirectional street-view type panorama, as…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Yujiao Shi , Dylan Campbell , Xin Yu , Hongdong Li

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

Multimedia · Computer Science 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo…

Sound · Computer Science 2025-02-26 Peiwen Sun , Sitong Cheng , Xiangtai Li , Zhen Ye , Huadai Liu , Honggang Zhang , Wei Xue , Yike Guo

Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Konstantin Klemmer , Esther Rolf , Caleb Robinson , Lester Mackey , Marc Rußwurm

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Rishabh Garg , Ruohan Gao , Kristen Grauman

Recent advances in generative modeling have substantially enhanced 3D urban generation, enabling applications in digital twins, virtual cities, and large-scale simulations. However, existing methods face two key challenges: (1) the need for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Yijie Kang , Xinliang Wang , Zhenyu Wu , Yifeng Shi , Hailong Zhu

Sound localization aims to find the source of the audio signal in the visual scene. However, it is labor-intensive to annotate the correlations between the signals sampled from the audio and visual modalities, thus making it difficult to…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Yan-Bo Lin , Hung-Yu Tseng , Hsin-Ying Lee , Yen-Yu Lin , Ming-Hsuan Yang

This paper presents a novel approach for cross-view synthesis aimed at generating plausible ground-level images from corresponding satellite imagery or vice versa. We refer to these tasks as satellite-to-ground (Sat2Grd) and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Tao Jun Lin , Wenqing Wang , Yujiao Shi , Akhil Perincherry , Ankit Vora , Hongdong Li

Radio frequency (RF) signal mapping, which is the process of analyzing and predicting the RF signal strength and distribution across specific areas, is crucial for cellular network planning and deployment. Traditional approaches to RF…

Signal Processing · Electrical Eng. & Systems 2024-01-08 Yiming Li , Zeyu Li , Zhihui Gao , Tingjun Chen

Recent advances in deep-learning based methods for image matching have demonstrated their superiority over traditional algorithms, enabling correspondence estimation in challenging scenes with significant differences in viewing angles,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Rahul Deshmukh , Avinash Kak

In geographical image segmentation, performance is often constrained by the limited availability of training data and a lack of generalizability, particularly for segmenting mobility infrastructure such as roads, sidewalks, and crosswalks.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Rafi Ibn Sultan , Chengyin Li , Hui Zhu , Prashant Khanduri , Marco Brocanelli , Dongxiao Zhu

Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or texts. However, the immersiveness and expressiveness of the…

Multimedia · Computer Science 2025-08-13 Wei Guo , Heng Wang , Jianbo Ma , Weidong Cai

For immersive applications, the generation of binaural sound that matches its visual counterpart is crucial to bring meaningful experiences to people in a virtual environment. Recent studies have shown the possibility of using neural…

Sound · Computer Science 2023-05-22 Francesc Lluís , Vasileios Chatziioannou , Alex Hofmann

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Rui Liu , Shuwei He , Yifan Hu , Haizhou Li

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

Multimedia · Computer Science 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook…

Sound · Computer Science 2026-03-02 Siyi Xie , Hanxin Zhu , Xinyi Chen , Tianyu He , Xin Li , Zhibo Chen