English
Related papers

Related papers: SoundPlot: An Open-Source Framework for Birdsong A…

200 papers

Understanding a 3D scene immediately with its exploration is essential for embodied tasks, where an agent must construct and comprehend the 3D scene in an online and nearly real-time manner. In this study, we propose EmbodiedSplat, an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Seungjun Lee , Zihan Wang , Yunsong Wang , Gim Hee Lee

The objective of this paper is to perform audio-visual sound source separation, i.e.~to separate component audios from a mixture based on the videos of sound sources. Moreover, we aim to pinpoint the source location in the input video…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Lingyu Zhu , Esa Rahtu

We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and song coherently synchronized with the input video. AudioGen-Omni introduces a novel…

Sound · Computer Science 2025-08-08 Le Wang , Jun Wang , Chunyu Qiang , Feng Deng , Chen Zhang , Di Zhang , Kun Gai

Music creation is typically composed of two parts: composing the musical score, and then performing the score with instruments to make sounds. While recent work has made much progress in automatic music generation in the symbolic domain,…

Sound · Computer Science 2018-11-13 Bryan Wang , Yi-Hsuan Yang

This paper proposes a real-time system integrating an acoustic material estimation from visual appearance and an on-the-fly mapping in the 3-dimension. The proposed method estimates the acoustic materials of surroundings in indoor scenes…

Robotics · Computer Science 2019-09-17 Taeyoung Kim , Youngsun Kwon , Sung-eui Yoon

Audio event has a hierarchical architecture in both time and frequency and can be grouped together to construct more abstract semantic audio classes. In this work, we develop a multiscale audio spectrogram Transformer (MAST) that employs…

Sound · Computer Science 2023-03-21 Wentao Zhu , Mohamed Omar

Sound source localization (SSL) is a critical technology for determining the position of sound sources in complex environments. However, existing methods face challenges such as high computational costs and precise calibration requirements,…

Sound · Computer Science 2025-05-28 Yiyuan Yang , Shitong Xu , Niki Trigoni , Andrew Markham

Vision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However,…

Robotics · Computer Science 2024-03-04 Amandine Brunetto , Sascha Hornauer , Stella X. Yu , Fabien Moutarde

Deep neural networks have been applied to audio spectrograms for respiratory sound classification. Existing models often treat the spectrogram as a synthetic image while overlooking its physical characteristics. In this paper, a Multi-View…

Sound · Computer Science 2024-05-31 Wentao He , Yuchen Yan , Jianfeng Ren , Ruibin Bai , Xudong Jiang

A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial…

Advances in computational chemistry have produced high-dimensional datasets on atmospherically relevant molecules. To aid exploration of such datasets, particularly for the study of atmospheric aerosol formation, we introduce PhiPlot: a…

Human-Computer Interaction · Computer Science 2026-03-13 Matias Loukojärvi , Ananth Mahadevan , Katsiaryna Haitsiukevich , Kai Puolamäki

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Jaemin Jung , Junseok Ahn , Chaeyoung Jung , Tan Dat Nguyen , Youngjoon Jang , Joon Son Chung

Acoustic metamaterials have become a novel and effective way to control sound waves and design acoustic devices. In this study, we design a 3D acoustic metamaterial lens (AML) to achieve point-to-point acoustic communication in air: any…

General Physics · Physics 2023-01-31 Fei Sun , Shuwei Guo , Borui Li , Yichao Liu , Sailing He

Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Risa Shinoda , Kaede Shiohara , Nakamasa Inoue , Kuniaki Saito , Hiroaki Santo , Fumio Okura

Few-shot audio-visual acoustics modeling seeks to synthesize the room impulse response in arbitrary locations with few-shot observations. To sufficiently exploit the provided few-shot data for accurate acoustic modeling, we present a…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Diwei Huang , Kunyang Lin , Peihao Chen , Qing Du , Mingkui Tan

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

Sound · Computer Science 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

We consider and propose a new problem of retrieving audio files relevant to multimodal design document inputs comprising both textual elements and visual imagery, e.g., birthday/greeting cards. In addition to enhancing user experience,…

Multimedia · Computer Science 2023-03-01 Prachi Singh , Srikrishna Karanam , Sumit Shekhar

Vision-Language Navigation (VLN) aims to guide agents by leveraging language instructions and visual cues, playing a pivotal role in embodied AI. Indoor VLN has been extensively studied, whereas outdoor aerial VLN remains underexplored. The…

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D Visual…

Machine Learning · Computer Science 2025-08-14 Duc Cao-Dinh , Khai Le-Duc , Anh Dao , Bach Phan Tat , Chris Ngo , Duy M. H. Nguyen , Nguyen X. Khanh , Thanh Nguyen-Tang

While 3D Vision Foundation Models (3DVFMs) have demonstrated remarkable zero-shot capabilities in visual geometry estimation, their direct application to generalizable novel view synthesis (NVS) remains challenging. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Minh-Quan Viet Bui , Jaeho Moon , Munchurl Kim