English
Related papers

Related papers: GWA: A Large High-Quality Acoustic Dataset for Aud…

200 papers

This Ph.D. thesis focuses on developing a system for high-quality speech synthesis and voice conversion. Vocoder-based speech analysis, manipulation, and synthesis plays a crucial role in various kinds of statistical parametric speech…

Sound · Computer Science 2021-01-26 Mohammed Salah Al-Radhi

Natural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our…

Full waveform inversion (FWI) is widely used in geophysics to reconstruct high-resolution velocity maps from seismic data. The recent success of data-driven FWI methods results in a rapidly increasing demand for open datasets to serve the…

Machine Learning · Computer Science 2023-06-27 Chengyuan Deng , Shihang Feng , Hanchen Wang , Xitong Zhang , Peng Jin , Yinan Feng , Qili Zeng , Yinpeng Chen , Youzuo Lin

Accurately modeling sound propagation with complex real-world environments is essential for Novel View Acoustic Synthesis (NVAS). While previous studies have leveraged visual perception to estimate spatial acoustics, the combined use of…

Multimedia · Computer Science 2025-03-18 Hadam Baek , Hannie Shin , Jiyoung Seo , Chanwoo Kim , Saerom Kim , Hyeongbok Kim , Sangpil Kim

Millimeter-wave (mmWave) radar-based gesture recognition is gaining attention as a key technology to enable intuitive human-machine interaction. Nevertheless, the significant challenge lies in obtaining large-scale, high-quality mmWave…

Human-Computer Interaction · Computer Science 2024-12-23 Huanqi Yang , Mingda Han , Xinyue Li , Di Duan , Tianxing Li , Weitao Xu

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

In their simplest form, bulk acoustic wave (BAW) devices consist of a piezoelectric crystal between two electrodes that transduce the material's vibrations into electrical signals. They are adopted in frequency control and metrology, with…

We present the Neural Waveshaping Unit (NEWT): a novel, lightweight, fully causal approach to neural audio synthesis which operates directly in the waveform domain, with an accompanying optimisation (FastNEWT) for efficient CPU inference.…

Sound · Computer Science 2021-07-28 Ben Hayes , Charalampos Saitis , György Fazekas

We present DeepSSM, an open-source code powered by neural networks (NNs) to emulate gravitational wave (GW) spectra produced by sound waves during cosmological first-order phase transitions in the radiation-dominated era. The training data…

Cosmology and Nongalactic Astrophysics · Physics 2025-08-26 Chi Tian , Xiao Wang , Csaba Balázs

High-frequency acoustic wave transducers, vibrating at gigahertz (GHz), favored for their compact size, are not only dominating the front-end of mobile handsets but are also expanding into various interdisciplinary fields, including quantum…

Signal Processing · Electrical Eng. & Systems 2026-04-28 Fangsheng Qian , Shuhan Chen , Wei Wei , Jiashuai Xu , Kai Yang , Junyan Zheng , Zijun Ren , Xingyu Liu , Yansong Yang

This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and diffusion probabilistic models. It starts from a Gaussian…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-12 Nanxin Chen , Yu Zhang , Heiga Zen , Ron J. Weiss , Mohammad Norouzi , William Chan

This paper presents UPV_RIR_DB, a structured database of measured room impulse responses (RIRs) designed to provide acoustic data with explicit spatial metadata and traceable acquisition parameters. The dataset currently contains 166…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-30 Jesús García-Gamborino , Laura Fuster , Daniel de la Prida , Luis A. Azpicueta-Ruiz , Gema Piñero

The MWA is a next-generation radio interferometer under construction in remote Western Australia. The data rate from the correlator makes storing the raw data infeasible, so the data must be processed in real-time. The processing task is of…

Instrumentation and Methods for Astrophysics · Physics 2009-02-06 S. Ord , L. Greenhill , R. Wayth , D. Mitchell , K. Dale , H. Pfister , R. G. Edgar

With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio…

Sound · Computer Science 2024-06-13 Zeyu Xie , Baihan Li , Xuenan Xu , Zheng Liang , Kai Yu , Mengyue Wu

The field of sound healing includes ancient practices coming from a broad range of cultures. Across such practices there is a variety of acoustic instrumentation utilised. Practitioners suggest that sound has the ability to target both…

Sound · Computer Science 2019-10-23 Alice Baird , Bjoern Schuller

Multimodal Deep Learning enhances decision-making by integrating diverse information sources, such as texts, images, audio, and videos. To develop trustworthy multimodal approaches, it is essential to understand how uncertainty impacts…

Machine Learning · Computer Science 2025-08-14 Grigor Bezirganyan , Sana Sellami , Laure Berti-Équille , Sébastien Fournier

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a…

Sound · Computer Science 2025-03-18 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Jiankang Deng , Xiatian Zhu

Spatial reasoning is fundamental to auditory perception, yet current audio large language models (ALLMs) largely rely on unstructured binaural cues and single step inference. This limits both perceptual accuracy in direction and distance…

Sound · Computer Science 2025-10-01 Subrata Biswas , Mohammad Nur Hossain Khan , Bashima Islam

Room acoustics measurements are used in many areas of audio research, from physical acoustics modelling and speech enhancement to virtual reality applications. This paper documents the technical specifications and choices made in the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-24 Thomas McKenzie , Leo McCormack , Christoph Hold

A method is presented for estimating and reconstructing the sound field within a room using physics-informed neural networks. By incorporating a limited set of experimental room impulse responses as training data, this approach combines…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-03 Xenofon Karakonstantis , Diego Caviedes-Nozal , Antoine Richard , Efren Fernandez-Grande
‹ Prev 1 4 5 6 7 8 10 Next ›