English
Related papers

Related papers: Both Ears Wide Open: Towards Language-Driven Spati…

200 papers

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio…

Sound · Computer Science 2019-05-15 Yu-Ding Lu , Hsin-Ying Lee , Hung-Yu Tseng , Ming-Hsuan Yang

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Hyundo Lee , Suhyung Choi , Inwoo Hwang , Byoung-Tak Zhang

Most existing text-to-audio (TTA) generation methods produce mono outputs, neglecting essential spatial information for immersive auditory experiences. To address this issue, we propose a cascaded method for text-to-multisource binaural…

Sound · Computer Science 2025-11-06 Yuxuan He , Xiaoran Yang , Ningning Pan , Gongping Huang

Interactive audio spatialization technology previously developed for video game authoring and rendering has evolved into an essential component of platforms enabling shared immersive virtual experiences for future co-presence, remote…

Sound · Computer Science 2021-09-28 Jean-Marc Jot , Rémi Audfray , Mark Hertensteiner , Brian Schmidt

This study aims to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video. To achieve this, we propose a novel method that guides single-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Akio Hayakawa , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Diffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diffusion models to other tasks, such as speech enhancement. A…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-10 Philippe Gonzalez , Zheng-Hua Tan , Jan Østergaard , Jesper Jensen , Tommy Sonne Alstrøm , Tobias May

While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Lijiang Li , Zuwei Long , Yunhang Shen , Heting Gao , Haoyu Cao , Xing Sun , Caifeng Shan , Ran He , Chaoyou Fu

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Rishabh Garg , Ruohan Gao , Kristen Grauman

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning.…

Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional…

Sound · Computer Science 2024-08-30 Tiantian Feng , Dimitrios Dimitriadis , Shrikanth Narayanan

In a range of recent works, object-centric architectures have been shown to be suitable for unsupervised scene decomposition in the vision domain. Inspired by these methods we present AudioSlots, a slot-centric generative model for blind…

Sound · Computer Science 2023-05-10 Pradyumna Reddy , Scott Wisdom , Klaus Greff , John R. Hershey , Thomas Kipf

Large audio-language models have made rapid progress in recognizing what is present in an audio clip, but spatial audio-language understanding still lacks a clear task interface. A model must also decide where sound events occur, which…

Sound · Computer Science 2026-05-12 Yuhuan You , Lai Wei , Xihong Wu , Tianshu Qu

Autoregressive language models dominate modern text generation, yet their sequential nature introduces fundamental limitations: decoding is slow, and maintaining global coherence remains challenging. Diffusion models offer a promising…

Computation and Language · Computer Science 2026-01-06 Viacheslav Meshchaninov , Egor Chimbulatov , Alexander Shabalin , Aleksandr Abramov , Dmitry Vetrov

This research introduces an innovative AI-driven multi-agent framework specifically designed for creating immersive audiobooks. Leveraging neural text-to-speech synthesis with FastSpeech 2 and VALL-E for expressive narration and…

Sound · Computer Science 2025-05-09 Shaja Arul Selvamani , Nia D'Souza Ganapathy

We address the challenge of making spatial audio datasets by proposing a shared mechanized recording space that can run custom acoustic experiments: a Mechatronic Acoustic Research System (MARS). To accommodate a wide variety of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-03 Austin Lu , Ethaniel Moore , Arya Nallanthighall , Kanad Sarkar , Manan Mittal , Ryan M. Corey , Paris Smaragdis , Andrew Singer

This survey paper provides a comprehensive overview of the recent advancements and challenges in applying large language models to the field of audio signal processing. Audio processing, with its diverse signal representations and a wide…

Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been proposed to model music as a mixture of multiple instrumental…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-18 Zhongweiyang Xu , Debottam Dutta , Yu-Lin Wei , Romit Roy Choudhury

Audiobook generation aims to create rich, immersive listening experiences from multimodal inputs, but current approaches face three critical challenges: (1) the lack of synergistic generation of diverse audio types (e.g., speech, sound…

Sound · Computer Science 2025-08-13 Yan Rong , Shan Yang , Chenxing Li , Dong Yu , Li Liu

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

Sound · Computer Science 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang
‹ Prev 1 4 5 6 7 8 10 Next ›