English
Related papers

Related papers: Quantifying and Correlating Rhythm Formants in Spe…

200 papers

This article focuses on the research tool for investigating the fundamental frequencies of voiced sounds. We introduce an objective and informative measurement method of pitch extractors' response to frequency-modulated tones. The method…

The speech-to-singing (STS) voice conversion task aims to generate singing samples corresponding to speech recordings while facing a major challenge: the alignment between the target (singing) pitch contour and the source (speech) content…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-25 Ruiqi Li , Rongjie Huang , Lichao Zhang , Jinglin Liu , Zhou Zhao

In this project, we wanted to discover an analog topology that could effectively convert amplitude-modulated (AM) signals to frequency-modulated (FM) signals, while also ensuring that both sets of signals were within their respective radio…

Signal Processing · Electrical Eng. & Systems 2025-09-16 Rishab Parthasarathy , Michael Popik , Noah Haefner

The resonances of forced dynamical systems occur when either the amplitude of the frequency response undergoes a local maximum (amplitude resonance) or phase lag quadrature takes places (phase resonance). This study focuses on the phase…

Dynamical Systems · Mathematics 2021-08-25 Martin Volvert , Gaetan Kerschen

While large language models (LLMs) have been applied to automatic speech recognition (ASR), the task of making the model streamable remains a challenge. This paper proposes a novel model architecture, Transducer-Llama, that integrates LLMs…

Computation and Language · Computer Science 2024-12-24 Keqi Deng , Jinxi Guo , Yingyi Ma , Niko Moritz , Philip C. Woodland , Ozlem Kalinli , Mike Seltzer

This work builds together two popular blocks of neural architecture, namely convolutional layers and Transformers, for large language models (LLMs). Non-causal conformers are used ubiquitously in automatic speech recognition. This work aims…

Computation and Language · Computer Science 2023-07-04 Prateek Verma

In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry information in long-form audio via a recurrent form, such that the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-03 Quan Wang , Yang Yu , Jason Pelecanos , Yiling Huang , Ignacio Lopez Moreno

We present a study of morphological irregularity. Following recent work, we define an information-theoretic measure of irregularity based on the predictability of forms in a language. Using a neural transduction model, we estimate this…

Computation and Language · Computer Science 2019-06-28 Shijie Wu , Ryan Cotterell , Timothy J. O'Donnell

The AdS/CFT correspondence between string theory in AdS space and conformal field theories in physical space-time leads to an analytic, semi-classical model for strongly-coupled QCD which has scale invariance and dimensional counting at…

High Energy Physics - Phenomenology · Physics 2008-11-26 Stanley J. Brodsky , Guy F. de Teramond

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

Oscillation synchronization phenomenon is widely observed in natural systems through frequency modulated signals, especially in biological neural networks. Frequency modulation is also one of most widely used technologies in engineering.…

Optimization and Control · Mathematics 2021-08-17 Zhiyong Chen

For spoken dialog systems to conduct fluid conversational interactions with users, the systems must be sensitive to turn-taking cues produced by a user. Models should be designed so that effective decisions can be made as to when it is…

Computation and Language · Computer Science 2018-07-02 Matthew Roddy , Gabriel Skantze , Naomi Harte

Wireless spectrum regulation is a complex and demanding process due to the rapid pace of technological progress, increasing demand for spectrum, and a multitude of stakeholders with potentially conflicting interests, alongside significant…

Networking and Internet Architecture · Computer Science 2024-03-27 Amir Ghasemi , Paul Guinand

Thousands of individuals need surgical removal of their larynx due to critical diseases every year and therefore, require an alternative form of communication to articulate speech sounds after the loss of their voice box. This work…

Image and Video Processing · Electrical Eng. & Systems 2020-07-01 Pramit Saha , Yadong Liu , Bryan Gick , Sidney Fels

The recently introduced acoustic ray-tracing semiclassical (RTS) method is validated for a set of practically relevant boundary conditions. RTS is a frequency domain geometrical method which directly reproduces the acoustic Green's…

Computational Physics · Physics 2018-10-17 Rok Prislan , Daniel Svenšek

Textless spoken language models (SLMs) are generative models of speech that do not rely on text supervision. Most textless SLMs learn to predict the next semantic token, a discrete representation of linguistic content, and rely on a…

Computation and Language · Computer Science 2025-10-23 Ju-Chieh Chou , Jiawei Zhou , Karen Livescu

Looped Transformers have emerged as an efficient and powerful class of models for reasoning in the language domain. Recent studies show that these models achieve strong performance on algorithmic and reasoning tasks, suggesting that looped…

Computation and Language · Computer Science 2026-02-13 Ahmadreza Jeddi , Marco Ciccone , Babak Taati

All-neural end-to-end (E2E) automatic speech recognition (ASR) systems that use a single neural network to transduce audio to word sequences have been shown to achieve state-of-the-art results on several tasks. In this work, we examine the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-28 Arun Narayanan , Rohit Prabhavalkar , Chung-Cheng Chiu , David Rybach , Tara N. Sainath , Trevor Strohman

This paper discusses and contains offensive content. Language models (LMs) are used in decision-making systems and as interactive assistants. However, how well do these models making judgements align with the diversity of human values,…

Computation and Language · Computer Science 2025-04-17 Michael Galarnyk , Agam Shah , Dipanwita Guhathakurta , Poojitha Nandigam , Sudheer Chava

Directly learning to generate audio waveforms in an autoregressive manner is a challenging task, due to the length of the raw sequences and the existence of important structure on many different timescales. Traditional approaches based on…

Sound · Computer Science 2025-10-06 Konrad Szewczyk , Daniel Gallo Fernández , James Townsend