English
Related papers

Related papers: Acoustic Volume Rendering for Neural Impulse Respo…

200 papers

The spatial impulse response (SIR) method is a well-known approach to calculate transient acoustic fields of arbitrary-shape transducers. It involves the evaluation of a time-dependent surface integral. Although analytic expressions of the…

Numerical Analysis · Mathematics 2021-11-01 Dimitris Perdios , Florian Martinez , Marcel Arditi , Jean-Philippe Thiran

Obtaining high-quality 3D reconstructions of room-scale scenes is of paramount importance for upcoming applications in AR or VR. These range from mixed reality applications for teleconferencing, virtual measuring, virtual room planing, to…

Computer Vision and Pattern Recognition · Computer Science 2022-03-16 Dejan Azinović , Ricardo Martin-Brualla , Dan B Goldman , Matthias Nießner , Justus Thies

Audio adversarial examples are audio files that have been manipulated to fool an automatic speech recognition (ASR) system, while still sounding benign to a human listener. Most methods to generate such samples are based on a two-step…

Sound · Computer Science 2023-10-06 Armin Ettenhofer , Jan-Philipp Schulze , Karla Pizzi

Spatial audio in Extended Reality (XR) provides users with better awareness of where virtual elements are placed, and efficiently guides them to events such as notifications, system alerts from different windows, or approaching avatars.…

Human-Computer Interaction · Computer Science 2024-08-20 Hyunsung Cho , Alexander Wang , Divya Kartik , Emily Liying Xie , Yukang Yan , David Lindlbauer

Sim2real transfer has received increasing attention lately due to the success of learning robotic tasks in simulation end-to-end. While there has been a lot of progress in transferring vision-based navigation policies, the existing sim2real…

Sound · Computer Science 2024-09-12 Changan Chen , Jordi Ramos , Anshul Tomar , Kristen Grauman

Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Philipp Götz , Gloria Dal Santo , Sebastian J. Schlecht , Vesa Välimäki , Emanuël A. P. Habets

We present Exact Volumetric Ellipsoid Rendering (EVER), a method for real-time differentiable emission-only volume rendering. Unlike recent rasterization based approach by 3D Gaussian Splatting (3DGS), our primitive based representation…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Alexander Mai , Peter Hedman , George Kopanas , Dor Verbin , David Futschik , Qiangeng Xu , Falko Kuester , Jonathan T. Barron , Yinda Zhang

Automatic speech recognition (ASR) allows a natural and intuitive interface for robotic educational applications for children. However there are a number of challenges to overcome to allow such an interface to operate robustly in realistic…

Computation and Language · Computer Science 2016-11-10 Samuel Fernando , Roger K. Moore , David Cameron , Emily C. Collins , Abigail Millings , Amanda J. Sharkey , Tony J. Prescott

We present AudioMiXR, an augmented reality (AR) interface intended to assess how users manipulate virtual audio objects situated in their physical space using six degrees of freedom (6DoF) deployed on a head-mounted display (Apple Vision…

Human-Computer Interaction · Computer Science 2025-08-07 Brandon Woodard , Margarita Geleta , Joseph J. LaViola , Andrea Fanelli , Rhonda Wilson

Acoustical behavior of a room for a given position of microphone and sound source is usually described using the room impulse response. If we rely on the standard uniform sampling, the estimation of room impulse response for arbitrary…

Audio and Speech Processing · Electrical Eng. & Systems 2018-02-19 Helena Peić Tukuljac , Thach Pham Vu , Hervé Lissek , Pierre Vandergheynst

The creation of virtual humans increasingly leverages automated synthesis of speech and gestures, enabling expressive, adaptable agents that effectively engage users. However, the independent development of voice and gesture generation…

Graphics · Computer Science 2025-07-02 Haoyang Du , Kiran Chhatre , Christopher Peters , Brian Keegan , Rachel McDonnell , Cathy Ennis

The paper presents results from a project aiming to create horizontally distributed surround sound sources and virtual sound images as auditory BCI (aBCI) stimuli. The purpose is to create evoked brain wave response patterns depending on…

Human-Computer Interaction · Computer Science 2012-10-11 Nozomu Nishikawa , Yoshihiro Matsumoto , Shoji Makino , Tomasz M. Rutkowski

The rise of deep learning algorithms has led many researchers to withdraw from using classic signal processing methods for sound generation. Deep learning models have achieved expressive voice synthesis, realistic sound textures, and…

Sound · Computer Science 2022-01-10 Anastasia Natsiou , Sean O'Leary

From the existing research it has been observed that many techniques and methodologies are available for performing every step of Automatic Speech Recognition (ASR) system, but the performance (Minimization of Word Error Recognition-WER and…

Computation and Language · Computer Science 2013-03-25 Urmila Shrawankar , Vilas Thakare

We present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms…

Graphics · Computer Science 2025-11-04 Jiye Lee , Chenghui Li , Linh Tran , Shih-En Wei , Jason Saragih , Alexander Richard , Hanbyul Joo , Shaojie Bai

Stereo matching is a fundamental building block for many vision and robotics applications. An informative and concise cost volume representation is vital for stereo matching of high accuracy and efficiency. In this paper, we present a novel…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Gangwei Xu , Yun Wang , Junda Cheng , Jinhui Tang , Xin Yang

The goal of voice conversion (VC) is to convert input voice to match the target speaker's voice while keeping text and prosody intact. VC is usually used in entertainment and speaking-aid systems, as well as applied for speech data…

Sound · Computer Science 2022-04-01 A. Kashkin , I. Karpukhin , S. Shishkin

Modeling room acoustics in a field setting involves some degree of blind parameter estimation from noisy and reverberant audio. Modern approaches leverage convolutional neural networks (CNNs) in tandem with time-frequency representation.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-15 Christopher Ick , Adib Mehrabi , Wenyu Jin

This research introduces an innovative AI-driven multi-agent framework specifically designed for creating immersive audiobooks. Leveraging neural text-to-speech synthesis with FastSpeech 2 and VALL-E for expressive narration and…

Sound · Computer Science 2025-05-09 Shaja Arul Selvamani , Nia D'Souza Ganapathy

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram
‹ Prev 1 8 9 10 Next ›