Related papers: An acoustic glottal source for vocal tract physica…
Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not…
Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional…
Acoustic solitons can be obtained by considering the propagation of large amplitude sound waves across a set of Helmholtz resonators. The model proposed by Sugimoto and his coauthors has been validated experimentally in previous works. Here…
This paper is concerned with the inverse acoustic scattering problems of reconstructing time-dependent multiple point sources and sources on a curve $L$ of the form $\lambda(t)\tau(x)\delta_L(x)$. A direct sampling method with a novel…
An onomatopoeic word, which is a character sequence that phonetically imitates a sound, is effective in expressing characteristics of sound such as duration, pitch, and timbre. We propose an environmental-sound-extraction method using…
The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt…
Articulatory-to-acoustic inversion strongly depends on the type of data used. While most previous studies rely on EMA, which is limited by the number of sensors and restricted to accessible articulators, we propose an approach aiming at a…
Solid materials may appear static, but at the atomic scale they are in constant vibrational motion. These vibrations, described by phonons, govern many key material properties, including structural stability, mechanical strength, optical…
With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences. Therefore, we propose a new task, deepfake general audio…
Earables (ear wearables) is rapidly emerging as a new platform encompassing a diverse range of personal applications. The traditional authentication methods hence become less applicable and inconvenient for earables due to their limited…
Sound absorbing materials are usually defined by five parameters: open porosity, static airflow resistivity, tortuosity, and two characteristic lengths. In recent decades, different methods have been developed in order to characterize these…
The simulation of two-dimensional (2D) wave propagation is an affordable computational task and its use can potentially improve time performance in vocal tracts' acoustic analysis. Several models have been designed that rely on 2D wave…
Wind-driven sound generation is a source of anger and pleasure, depending on the situation: airframe and car noise, or combustion noise are some of the most disturbing environmental pollutions, whereas musical instruments are sources of…
The way infants use auditory cues to learn to speak despite the acoustic mismatch of their vocal apparatus is a hot topic of scientific debate. The simulation of early vocal learning using articulatory speech synthesis offers a way towards…
Many hearables contain an in-ear microphone, which may be used to capture the own voice of its user in noisy environments. Since the in-ear microphone mostly records body-conducted speech due to ear canal occlusion, it suffers from…
The photoacoustic signal in a closed T-cell resonator is generated and measured using laser based photoacoustic spectroscopy. The signal is modelled using the amplitude mode expansion method, which is based on eigenmode expansion and…
Recent strides in neural speech synthesis technologies, while enjoying widespread applications, have nonetheless introduced a series of challenges, spurring interest in the defence against the threat of misuse and abuse. Notably, source…
The imitation of percussive sounds via the human voice is a natural and effective tool for communicating rhythmic ideas on the fly. Thus, the automatic retrieval of drum sounds using vocal percussion can help artists prototype drum patterns…
In this work, a new Physics laboratory experiment on Acoustics beats is presented. We have designed a simple experimental setup to study superposition of sound waves of slightly different frequencies (acoustic beat). The microphone of a…
This paper introduces GlOttal-flow LPC Filter (GOLF), a novel method for singing voice synthesis (SVS) that exploits the physical characteristics of the human voice using differentiable digital signal processing. GOLF employs a glottal…