English
Related papers

Related papers: AudioScene: Integrating Object-Event Audio into 3D…

200 papers

Although 360\textdegree{} cameras ease the capture of panoramic footage, it remains challenging to add realistic 360\textdegree{} audio that blends into the captured scene and is synchronized with the camera motion. We present a method for…

Graphics · Computer Science 2018-05-15 Dingzeyu Li , Timothy R. Langlois , Changxi Zheng

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

Cinematic video production requires control over scene-subject composition and camera movement, but live-action shooting remains costly due to the need for constructing physical sets. To address this, we introduce the task of cinematic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Kaiyi Huang , Yukun Huang , Yu Li , Jianhong Bai , Xintao Wang , Zinan Lin , Xuefei Ning , Jiwen Yu , Pengfei Wan , Yu Wang , Xihui Liu

Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object grounding and contextual reasoning, limiting their ability…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Haifeng Huang , Yilun Chen , Zehan Wang , Jiangmiao Pang , Zhou Zhao

This paper proposes a benchmark of submissions to Detection and Classification Acoustic Scene and Events 2021 Challenge (DCASE) Task 4 representing a sampling of the state-of-the-art in Sound Event Detection task. The submissions are…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Francesca Ronchini , Romain Serizel

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given region for spatial audio recorded by a microphone array. The…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-12 Jinzheng Zhao , Yong Xu , Haohe Liu , Davide Berghi , Xinyuan Qian , Qiuqiang Kong , Junqi Zhao , Mark D. Plumbley , Wenwu Wang

Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plausible locations and scales without interpenetration, aligning…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Qingyang Liu , Bingjie Gao , Weiheng Huang , Jun Zhang , Zhongqian Sun , Yang Wei , Fengrui Liu , Zelin Peng , Qianli Ma , Shuai Yang , Zhaohe Liao , Haonan Zhao , Li Niu

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to…

Sound · Computer Science 2025-02-10 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

Panoptic scene understanding and tracking of dynamic agents are essential for robots and automated vehicles to navigate in urban environments. As LiDARs provide accurate illumination-independent geometric depictions of the scene, performing…

Computer Vision and Pattern Recognition · Computer Science 2021-12-28 Whye Kit Fong , Rohit Mohan , Juana Valeria Hurtado , Lubing Zhou , Holger Caesar , Oscar Beijbom , Abhinav Valada

The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts. However, existing datasets typically suffer from limitations in data scale or diversity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Weipeng Zhong , Peizhou Cao , Yichen Jin , Li Luo , Wenzhe Cai , Jingli Lin , Hanqing Wang , Zhaoyang Lyu , Tai Wang , Bo Dai , Xudong Xu , Jiangmiao Pang

In recent years the automotive industry has been strongly promoting the development of smart cars, equipped with multi-modal sensors to gather information about the surroundings, in order to aid human drivers or make autonomous decisions.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-31 Jun Yin , Stefano Damiano , Marian Verhelst , Toon van Waterschoot , Andre Guntoro

From whirling ceiling fans to ticking clocks, the sounds that we hear subtly vary as we move through a scene. We ask whether these ambient sounds convey information about 3D scene structure and, if so, whether they provide a useful learning…

Sound · Computer Science 2021-11-11 Ziyang Chen , Xixi Hu , Andrew Owens

Within a perception framework for autonomous mobile and robotic systems, semantic analysis of 3D point clouds typically generated by LiDARs is key to numerous applications, such as object detection and recognition, and scene reconstruction.…

Robotics · Computer Science 2024-10-14 Samir Abou Haidar , Alexandre Chariot , Mehdi Darouich , Cyril Joly , Jean-Emmanuel Deschaud

While 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Xudong Xu , Dejan Markovic , Jacob Sandakly , Todd Keebler , Steven Krenn , Alexander Richard

Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remove these background sounds using speech enhancement or train…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-04 Chaitanya Narisetty , Emiru Tsunoo , Xuankai Chang , Yosuke Kashiwagi , Michael Hentschel , Shinji Watanabe

Automatic detection and classification of animal sounds has many applications in biodiversity monitoring and animal behaviour. In the past twenty years, the volume of digitised wildlife sound available has massively increased, and automatic…

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xinhao Cai , Minghang Zheng , Xin Jin , Yang Liu

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-18 Shubo Lv , Yihui Fu , Yukai Jv , Lei Xie , Weixin Zhu , Wei Rao , Yannan Wang

3D Semantic Scene Completion (SSC) provides comprehensive scene geometry and semantics for autonomous driving perception, which is crucial for enabling accurate and reliable decision-making. However, existing SSC methods are limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Meng Wang , Fan Wu , Ruihui Li , Yunchuan Qin , Zhuo Tang , Kenli Li

The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and…

Computer Vision and Pattern Recognition · Computer Science 2022-07-08 Chuang Gan , Yi Gu , Siyuan Zhou , Jeremy Schwartz , Seth Alter , James Traer , Dan Gutfreund , Joshua B. Tenenbaum , Josh McDermott , Antonio Torralba