English
Related papers

Related papers: Quantifying and Correlating Rhythm Formants in Spe…

200 papers

Large language models (LLMs) have exhibited impressive performance and surprising emergent properties. However, their effectiveness remains limited by the fixed context window of the transformer architecture, posing challenges for…

Computation and Language · Computer Science 2025-06-16 Tianqi Du , Haotian Huang , Yifei Wang , Yisen Wang

Realistic sound propagation is essential for immersion in a virtual scene, yet physically accurate wave-based simulations remain computationally prohibitive for real-time applications. Wave coding methods address this limitation by…

Sound · Computer Science 2026-02-09 Hugo Seuté , Pranai Vasudev , Etienne Richan , Louis-Xavier Buffoni

The short-time Fourier transform (STFT) represents a window of audio samples as a set of complex coefficients. These are advantageously viewed as magnitudes and phases and the overall distribution of phases is very often assumed to be…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-16 Stephen D. Voran

Modulation classification plays a crucial role in wireless communication systems, enabling applications such as cognitive radio, spectrum monitoring, and electronic warfare. Conventional techniques often involve deep learning or complex…

Signal Processing · Electrical Eng. & Systems 2025-06-26 Srinivas Rahul Sapireddy , Mostafizur Rahman

Acoustic scene classification (ASC) aims to identify the type of scene (environment) in which a given audio signal is recorded. The log-mel feature and convolutional neural network (CNN) have recently become the most popular time-frequency…

Sound · Computer Science 2021-08-12 Yuzhong Wu , Tan Lee

Frequency modulation (FM) is a form of radio broadcasting which is widely used nowadays and has been for almost a century. We suggest a software-defined-radio (SDR) receiver for FM demodulation that adopts an end-to-end learning based…

Machine Learning · Computer Science 2017-10-10 Dan Elbaz , Michael Zibulevsky

While transformers demonstrate outstanding performance across various audio tasks, their application to neural vocoders remains challenging. Neural vocoders require the generation of long audio signals at the sample level, which demands…

Sound · Computer Science 2025-12-30 Seongho Hong , Yong-Hoon Choi

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-19 Edresson Casanova , Ryan Langman , Paarth Neekhara , Shehzeen Hussain , Jason Li , Subhankar Ghosh , Ante Jukić , Sang-gil Lee

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

Sound · Computer Science 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approaches aim to isolate…

Computation and Language · Computer Science 2025-06-23 Daejin Jo , Jeeyoung Yun , Byungseok Roh , Sungwoong Kim

Frequency locking to an external forcing frequency is a {well} known phenomenon. In the auditory system, it results in a localized traveling wave, the shape of which is essential for efficient discrimination between incoming frequencies. An…

Pattern Formation and Solitons · Physics 2018-09-12 Yuval Edri , Dolores Bozovic , Ehud Meron , Arik Yochelis

Passive resonators-systems that exhibit loss but no gain-are foundational elements across nearly every domain of physics and many types of of systems such as subwavelength particles, dielectric slabs, electric circuits, biological…

Optics · Physics 2025-06-05 Asaf Farhi , Dror Hershkovitz , Haim Suchowski

The all-temperature magnon (ATM) theory [J. Phys. Condens. Matter 21, 336003/1-14, 2009] has been used to analyze the temperature dependence of magnetization as well as the internal energy components of a mono-domain ferromagnetic solid.…

Materials Science · Physics 2023-06-30 Sambhu N. Datta

Chord recognition serves as a critical task in music information retrieval due to the abstract and descriptive nature of chords in music analysis. While audio chord recognition systems have achieved significant accuracy for small…

The angular and frequency correlation functions of the transmission coefficient for light propagation through a strongly scattering amplifying medium are considered. It is found that just as in the case of an elastic scattering medium the…

Mesoscale and Nanoscale Physics · Physics 2009-10-31 A. A. Burkov , A. Yu. Zyuzin

Prosody is usually defined in terms of the three distinct but interacting domains of pitch, intensity and duration patterning, or, more generally, as phonological and phonetic properties of 'suprasegmentals', speech segments which are…

Computation and Language · Computer Science 2018-05-16 Dafydd Gibbon

The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional generative formulations often treat each dataset sample as a fixed…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-24 Doyeop Kwak , Youngjoon Jang , Joon Son Chung

Language models (LMs) for text data have been studied extensively for their usefulness in language generation and other downstream tasks. However, language modelling purely in the speech domain is still a relatively unexplored topic, with…

Computation and Language · Computer Science 2021-11-02 Anurag Katakkar , Alan W Black

The aim of this study is to determine the effect of language varieties on the spectral distribution of stressed and unstressed sonorants (nasals /m, n/, lateral approximants /l/, and rhotics /r/) and on their coarticulatory effects on…

Computation and Language · Computer Science 2021-10-11 Charalambos Themistocleous , Valantis Fyndanis , Kyrana Tsapkini

Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal representations align with human neural dynamics during…

Sound · Computer Science 2026-02-04 Haoyun Yang , Xin Xiao , Jiang Zhong , Yu Tian , Dong Xiaohua , Yu Mao , Hao Wu , Kaiwen Wei