English
Related papers

Related papers: Quantifying and Correlating Rhythm Formants in Spe…

200 papers

Large-scale latent diffusion models (LDMs) excel in content generation across various modalities, but their reliance on phonemes and durations in text-to-speech (TTS) limits scalability and access from other fields. While recent studies…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-18 Keon Lee , Dong Won Kim , Jaehyeon Kim , Seungjun Chung , Jaewoong Cho

Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates the actual audio. High-quality spectrum models usually…

Sound · Computer Science 2021-04-05 Qing He , Zhiping Xiu , Thilo Koehler , Jilong Wu

We consider a round turbulent jet grazing a rectangular plate angled at $45^\circ$. Through sound pressure measurements, the tonal dynamics associated with jet-edge interaction are explored in a parameter space comprising jet Mach number,…

Fluid Dynamics · Physics 2026-03-26 Michael N. Stavropoulos , André V. G. Cavalieri , Lutz Lesshafft , Peter Jordan

During voiced speech, the human vocal folds interact with the vocal tract acoustics. The resulting glottal source-resonator coupling has been observed using mathematical and physical models as well as in in vivo phonation. We propose a…

Fluid Dynamics · Physics 2017-03-16 Atte Aalto , Tiina Murtola , Jarmo Malinen , Daniel Aalto , Martti Vainio

Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-17 Zhong Meng , Shinji Watanabe , John R. Hershey , Hakan Erdogan

Inspired by the recent experimental signatures of relativistic hydrodynamics in graphene, we investigate theoretically the behavior of hydrodynamic sound modes in such quasi-relativistic fluids near charge neutrality, within linear…

Strongly Correlated Electrons · Physics 2016-07-06 Andrew Lucas

Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a…

Sound · Computer Science 2026-04-21 Hao Meng , Siyuan Zheng , Shuran Zhou , Qiangqiang Wang , Yang Song

Large language models (LLMs) and multimodal models have become powerful general-purpose reasoning systems. However, radio-frequency (RF) signals, which underpin wireless systems, are still not natively supported by these models. Existing…

Signal Processing · Electrical Eng. & Systems 2026-02-17 Hang Zou , Yu Tian , Bohao Wang , Lina Bariah , Samson Lasaulce , Chongwen Huang , Mérouane Debbah

This study experimentally validates a numerical model of electromagnetic propagation through the human head during the pronunciation of different vowels, with the goal of improving our understanding of the underlying physical phenomena. A…

Formant synthesis aims to generate speech with controllable formant structures, enabling precise control of vocal resonance and phonetic features. However, while existing formant synthesis approaches enable precise formant manipulation,…

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is key for multi-channel enhancement. Deep learning shows great potential on multi-channel…

Sound · Computer Science 2023-09-21 Jiahui Pan , Pengjie Shen , Hui Zhang , Xueliang Zhang

Mechanical resonators that possess coupled modes with harmonic frequency relations have recently sparked interest due to their suitability for controllable energy transfer and non-Hermitian dynamics. Here, we show coupling between high…

LLM-based automatic speech recognition models demonstrate strong performance by connecting audio encoders and LLMs. However, data scarcity of paired speech and transcription often hinders their adaptation to new domains, making text-only…

Sound · Computer Science 2026-05-15 Ryo Magoshi , Takashi Maekaku , Yusuke Shinohara

A new mechanism of low frequency (1/f-like) noise generation is described and analyzed. It is attributed to higher frequency asymmetric resonance modes, which are stimulated by a random factor, e.g. due to thermal excitation.…

Applied Physics · Physics 2018-06-08 Andrzej Stankiewicz

Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ability of such systems remains underexplored, primarily due…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-18 Haoyu Li , Mingyang Han , Yu Xi , Dongxiao Wang , Hankun Wang , Haoxiang Shi , Boyu Li , Jun Song , Bo Zheng , Shuai Wang , Kai Yu

Various resonant and near-resonant patterns form in a light-sensitive Belousov-Zhabotinsky (BZ) reaction in response to a spatially-homogeneous time-periodic perturbation with light. The regions (tongues) in the forcing frequency and…

Pattern Formation and Solitons · Physics 2009-11-10 Anna L. Lin , Aric Hagberg , Ehud Meron , Harry L. Swinney

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

Sound · Computer Science 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

This paper proposes a novel time-frequency warped waveform for short symbols, massive machine-type communication (mMTC), and internet of things (IoT) applications. The waveform is composed of asymmetric raised cosine (RC) pulses to increase…

Signal Processing · Electrical Eng. & Systems 2024-10-28 Mostafa Ibrahim , Huseyin Arslan , Hakan Ali Cirpan , Sabit Ekin

Low frequency oscillator (LFO) driven audio effects such as phaser, flanger, and chorus, modify an input signal using time-varying filters and delays, resulting in characteristic sweeping or widening effects. It has been shown that these…

Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speech of an unseen speaker as acoustic prompt. However, they…

Sound · Computer Science 2024-06-07 Jinlong Xue , Yayue Deng , Yicheng Han , Yingming Gao , Ya Li