English
Related papers

Related papers: GWA: A Large High-Quality Acoustic Dataset for Aud…

200 papers

The phenomenon of Gravitational Wave (GW) analysis has grown in popularity as technology has advanced and the process of observing gravitational waves has become more precise. Although the sensitivity and the frequency of observation of GW…

Machine Learning · Computer Science 2023-11-07 Elena-Simona Apostol , Ciprian-Octavian Truică

This paper introduces BIRD, the Big Impulse Response Dataset. This open dataset consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open…

In this paper, we present HOMULA-RIR, a dataset of room impulse responses (RIRs) acquired using both higher-order microphones (HOMs) and a uniform linear array (ULA), in order to model a remote attendance teleconferencing scenario.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-22 Federico Miotello , Paolo Ostan , Mirco Pezzoli , Luca Comanducci , Alberto Bernardini , Fabio Antonacci , Augusto Sarti

The LISA telescope will provide the first opportunity to probe the scenario of a first-order phase transition happening close to the electroweak scale. By now, it is evident that the main contribution to the GW spectrum comes from the sound…

Cosmology and Nongalactic Astrophysics · Physics 2021-04-14 Ryusuke Jinno , Thomas Konstandin , Henrique Rubira

This paper proposes a novel way of doing audio synthesis at the waveform level using Transformer architectures. We propose a deep neural network for generating waveforms, similar to wavenet. This is fully probabilistic, auto-regressive, and…

Sound · Computer Science 2021-07-09 Prateek Verma , Chris Chafe

Acoustical behavior of a room for a given position of microphone and sound source is usually described using the room impulse response. If we rely on the standard uniform sampling, the estimation of room impulse response for arbitrary…

Audio and Speech Processing · Electrical Eng. & Systems 2018-02-19 Helena Peić Tukuljac , Thach Pham Vu , Hervé Lissek , Pierre Vandergheynst

Moving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Changan Chen , Unnat Jain , Carl Schissler , Sebastia Vicenc Amengual Gari , Ziad Al-Halah , Vamsi Krishna Ithapu , Philip Robinson , Kristen Grauman

The Laser Interferometer Space Antenna (LISA) will open the mHz frequency window of the gravitational wave (GW) landscape. Among all the new GW sources expected to emit in this frequency band, extreme mass-ratio inspirals (EMRIs) constitute…

Cosmology and Nongalactic Astrophysics · Physics 2021-11-03 Danny Laghi , Nicola Tamanini , Walter Del Pozzo , Alberto Sesana , Jonathan Gair , Stanislav Babak , David Izquierdo-Villalba

Deep generative modeling has the potential to cause significant harm to society. Recognizing this threat, a magnitude of research into detecting so-called "Deepfakes" has emerged. This research most often focuses on the image domain, while…

Machine Learning · Computer Science 2021-11-05 Joel Frank , Lea Schönherr

Audio adversarial examples are audio files that have been manipulated to fool an automatic speech recognition (ASR) system, while still sounding benign to a human listener. Most methods to generate such samples are based on a two-step…

Sound · Computer Science 2023-10-06 Armin Ettenhofer , Jan-Philipp Schulze , Karla Pizzi

Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. An ideal dataset is…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-01 Philipp Klumpp , Tomás Arias-Vergara , Paula Andrea Pérez-Toro , Elmar Nöth , Juan Rafael Orozco-Arroyave

Ultrasound imaging is a widely used, non-invasive diagnostic tool in modern medicine. A crucial assumption is a constant sound speed in the observed medium. For large scale sound speed variations, this assumption leads to blurred and…

Numerical Analysis · Mathematics 2025-10-08 Simon Hackl , Simon Hubmer , Ronny Ramlau

The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling,…

Sound · Computer Science 2024-02-05 Shijia Liao , Shiyi Lan , Arun George Zachariah

The standard model of particle physics is known to be intriguingly successful. However their rich phenomena represented by the phase transitions (PTs) have not been completely understood yet, including the possibility of the existence of…

Cosmology and Nongalactic Astrophysics · Physics 2021-07-07 Katsuya T. Abe , Yuichiro Tada , Ikumi Ueda

Large-scale 3D generative models require substantial computational resources yet often fall short in capturing fine details and complex geometries at high resolutions. We attribute this limitation to the inefficiency of current…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Aditya Sanghi , Aliasghar Khani , Pradyumna Reddy , Arianna Rampini , Derek Cheung , Kamal Rahimi Malekshan , Kanika Madan , Hooman Shayani

Audio Super-Resolution is a set of techniques aimed at high-quality estimation of the given signal as if it would be sampled with higher sample rate. Among suggested methods there are diffusion and flow models (which are considered slower),…

Sound · Computer Science 2026-03-05 Nikita Kuznetsov , Maksim Kaledin

In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time…

Acoustic word embeddings (AWEs) are vector representations such that different acoustic exemplars of the same word are projected nearby in the embedding space. In addition to their use in speech technology applications such as spoken term…

Computation and Language · Computer Science 2023-01-10 Badr M. Abdullah , Dietrich Klakow

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 Guangyao Li , Yake Wei , Yapeng Tian , Chenliang Xu , Ji-Rong Wen , Di Hu

We present an open-source differentiable acoustic simulator, j-Wave, which can solve time-varying and time-harmonic acoustic problems. It supports automatic differentiation, which is a program transformation technique that has many…

Computational Physics · Physics 2022-07-05 Antonio Stanziola , Simon R. Arridge , Ben T. Cox , Bradley E. Treeby
‹ Prev 1 3 4 5 6 7 10 Next ›