中文
相关论文

相关论文: Real Acoustic Fields: An Audio-Visual Room Acousti…

200 篇论文

Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation…

计算机视觉与模式识别 · 计算机科学 2021-08-20 Yudong Guo , Keyu Chen , Sen Liang , Yong-Jin Liu , Hujun Bao , Juyong Zhang

Neural radiance fields (NeRF) encode a scene into a neural representation that enables photo-realistic rendering of novel views. However, a successful reconstruction from RGB images requires a large number of input views taken under static…

计算机视觉与模式识别 · 计算机科学 2022-04-08 Barbara Roessle , Jonathan T. Barron , Ben Mildenhall , Pratul P. Srinivasan , Matthias Nießner

This paper proposes an adaptive near-field beam training method to enhance performance in multi-user and multipath environments. The approach identifies multiple strongest beams through beam sweeping and linearly combines their received…

信号处理 · 电气工程与系统科学 2025-05-14 Zijun Wang , Rama Kiran , Jinesh Nair , Chien-Hua Chen , Tzu-Han Chou , Shawn Tsai , Rui Zhang

We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Xinglong Luo , Kunming Luo , Ao Luo , Zhengning Wang , Ping Tan , Shuaicheng Liu

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Haytham M. Fayek , Anurag Kumar

Modern neural-network-based speech processing systems are typically required to be robust against reverberation, and the training of such systems thus needs a large amount of reverberant data. During the training of the systems, on-the-fly…

声音 · 计算机科学 2023-04-18 Yi Luo , Rongzhi Gu

With the advancement of IoT technology, recognizing user activities with machine learning methods is a promising way to provide various smart services to users. High-quality data with privacy protection is essential for deploying such…

人机交互 · 计算机科学 2024-01-18 Hyunju Kim , Geon Kim , Taehoon Lee , Kisoo Kim , Dongman Lee

This work proposes DOFS, a pilot dataset of 3D deformable objects (DOs) (e.g., elasto-plastic objects) with full spatial information (i.e., top, side, and bottom information) using a novel and low-cost data collection platform with a…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Zhen Zhang , Xiangyu Chu , Yunxi Tang , K. W. Samuel Au

Robust depth perception in visually-degraded environments is crucial for autonomous aerial systems. Thermal imaging cameras, which capture infrared radiation, are robust to visual degradation. However, due to lack of a large-scale dataset,…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Devansh Dhrafani , Yifei Liu , Andrew Jong , Ukcheol Shin , Yao He , Tyler Harp , Yaoyu Hu , Jean Oh , Sebastian Scherer

Dynamic objects in the environment, such as people and other agents, lead to challenges for existing simultaneous localization and mapping (SLAM) approaches. To deal with dynamic environments, computer vision researchers usually apply some…

机器人学 · 计算机科学 2021-08-04 Tianwei Zhang , Huayan Zhang , Xiaofei Li , Junfeng Chen , Tin Lun Lam , Sethu Vijayakumar

Accurate and efficient simulation of room impulse responses is crucial for spatial audio applications. However, existing acoustic ray-tracing tools often operate as black boxes and only output impulse responses (IRs), providing limited…

声音 · 计算机科学 2025-03-25 Yongyi Zang , Qiuqiang Kong

As audio-visual systems increasingly bring immersive and interactive capabilities into our work and leisure activities, so the need for naturalistic test material grows. New volumetric datasets have captured high-quality 3D video, but…

多媒体 · 计算机科学 2021-05-04 Hanne Stenzel , Davide Berghi , Marco Volino , Philip J. B. Jackson

Increasingly, phonetic research utilizes data collected from participants who record themselves on readily available devices. Though such recordings are convenient, their suitability for acoustic analysis remains an open question,…

声音 · 计算机科学 2024-10-08 Cong Zhang , Kathleen Jepson , Yu-Ying Chuang

The demand for realistic virtual immersive audio continues to grow, with Head-Related Transfer Functions (HRTFs) playing a key role. HRTFs capture how sound reaches our ears, reflecting unique anatomical features and enhancing spatial…

声音 · 计算机科学 2026-01-26 Xuyi Hu , Jian Li , Lorenzo Picinali , Aidan O. T. Hogg

Multi-focus image fusion, a technique to generate an all-in-focus image from two or more partially-focused source images, can benefit many computer vision tasks. However, currently there is no large and realistic dataset to perform…

计算机视觉与模式识别 · 计算机科学 2020-08-31 Juncheng Zhang , Qingmin Liao , Shaojun Liu , Haoyu Ma , Wenming Yang , Jing-Hao Xue

With the upcoming multitude of commercial and public applications envisioned in the mobile 6G radio landscape using unmanned aerial vehicles (UAVs), integrated sensing and communication (ISAC) plays a key role to enable the detection and…

We present an iVector based Acoustic Scene Classification (ASC) system suited for real life settings where active foreground speech can be present. In the proposed system, each recording is represented by a fixed-length iVector that models…

音频与语音处理 · 电气工程与系统科学 2021-08-03 Siyuan Song , Brecht Desplanques , Celest De Moor , Kris Demuynck , Nilesh Madhu

We explore active audio-visual separation for dynamic sound sources, where an embodied agent moves intelligently in a 3D environment to continuously isolate the time-varying audio stream being emitted by an object of interest. The agent…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Sagnik Majumder , Kristen Grauman

Hearing aids use dynamic range compression (DRC), a form of automatic gain control, to make quiet sounds louder and loud sounds quieter. Compression can improve listening comfort, but it can also cause distortion in noisy environments. It…

音频与语音处理 · 电气工程与系统科学 2021-07-28 Ryan M. Corey , Andrew C. Singer

For augmented (AR) and virtual reality (VR) applications, accurate estimates of the acoustic characteristics of a scene are critical for creating a sense of immersion. However, directly estimating Room-impulse Responses (RIRs) from scene…

音频与语音处理 · 电气工程与系统科学 2025-11-20 Ricardo Falcon-Perez , Ruohan Gao , Gregor Mueckl , Sebastia V. Amengual Gari , Ishwarya Ananthabhotla