English
Related papers

Related papers: Analyzing long-term rhythm variations in Mising an…

200 papers

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA)…

Sound · Computer Science 2024-03-22 Samuel Pegg , Kai Li , Xiaolin Hu

We study the problem of estimating the frequencies of several complex sinusoids with constant amplitude (CA) (also called constant modulus) from multichannel signals of their superposition. To exploit the CA property for frequency…

Optimization and Control · Mathematics 2024-04-23 Xunmeng Wu , Zai Yang , Zongben Xu

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Language Model (ALM) with…

Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is treated as a fine-tuning task. To investigate the effects of…

This paper investigates part-of-speech tagging, an important task in Natural Language Processing (NLP) for the Nagamese language. The Nagamese language, a.k.a. Naga Pidgin, is an Assamese-lexified Creole language developed primarily as a…

Computation and Language · Computer Science 2025-10-14 Alovi N Shohe , Chonglio Khiamungam , Teisovi Angami

How important are different temporal speech modulations for speech recognition? We answer this question from two complementary perspectives. Firstly, we quantify the amount of phonetic \textit{information} in the modulation spectrum of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-24 Samik Sadhu , Hynek Hermansky

Tonogenesis-the historical process by which segmental contrasts evolve into lexical tone-has traditionally been studied through comparative reconstruction and acoustic phonetics. We introduce a computational approach that quantifies the…

Computation and Language · Computer Science 2025-10-28 Siyu Liang , Zhaxi Zerong

The present contribution is a tutorial on selected aspects of prosody, the rhythms and melodies of speech, based on a course of the same name at the Summer School on Contemporary Phonetics and Phonology at Tongji University, Shanghai, China…

Computation and Language · Computer Science 2017-04-28 Dafydd Gibbon

In this paper, we give estimators of the frequency, amplitude and phase of a noisy sinusoidal signal with time-varying amplitude by using the algebraic parametric techniques introduced by Fliess and Sira-Ramirez. We apply a similar strategy…

Numerical Analysis · Mathematics 2011-05-17 Da-Yan Liu , Olivier Gibaru , Wilfrid Perruquetti

Birdsong often contains large amounts of rapid frequency modulation (FM). It is believed that the use or otherwise of FM is adaptive to the acoustic environment, and also that there are specific social uses of FM such as trills in…

Sound · Computer Science 2015-09-22 Dan Stowell , Mark D. Plumbley

Modern automatic speech recognition (ASR) systems have been observed to function better for certain speaker groups (SGs) than others, despite recent gains in overall performance. One potential impediment to progress towards fairer ASR is a…

Computation and Language · Computer Science 2026-04-27 Felix Herron , Solange Rossato , Alexandre Allauzen , François Portet

With the rapid development of radar jamming systems, especially digital radio frequency memory (DRFM), the electromagnetic environment has become increasingly complex. In recent years, most existing studies have focused solely on either…

Signal Processing · Electrical Eng. & Systems 2025-06-10 Huake Wang , Xudong Han , Bairui Cai , Guisheng Liao , Yinghui Quan

Recent speech foundation models excel at multilingual automatic speech recognition (ASR) for high-resource languages, but adapting them to low-resource languages remains challenging due to data scarcity and efficiency constraints.…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-03 Yang Xiao , Eun-Jung Holden , Ting Dang

Large audio language models (ALMs) extend LLMs with auditory understanding. A common approach freezes the LLM and trains only an adapter on self-generated targets. However, this fails for reasoning LLMs (RLMs) whose built-in…

Computation and Language · Computer Science 2026-03-11 Petr Grinberg , Hassan Shahmohammadi

Accurate medical image segmentation remains challenging due to blurred lesion boundaries (LBA), loss of high-frequency details (LHD), and difficulty in modeling long-range anatomical structures (DC-LRSS). Vision Mamba employs…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Ze Rong , ZiYue Zhao , Zhaoxin Wang , Lei Ma

Phone level localization of mis-articulation is a key requirement for an automatic articulation error assessment system. A robust phone segmentation technique is essential to aid in real-time assessment of phone level mis-articulations of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Bhavik Vachhani , Chitralekha Bhat , Sunil Kopparapu

Improvement in time resolution sometimes introduces short-range random noises into temporal data sequences. These noises affect the results of power-spectrum analyses and the Detrended Fluctuation Analysis (DFA). The DFA is one of useful…

Data Analysis, Statistics and Probability · Physics 2009-02-05 Shin-ichi Tadaki

Conversational speech normally is embodied with loose syntactic structures at the utterance level but simultaneously exhibits topical coherence relations across consecutive utterances. Prior work has shown that capturing longer context…

Computation and Language · Computer Science 2022-06-02 Bi-Cheng Yan , Hsin-Wei Wang , Shih-Hsuan Chiu , Hsuan-Sheng Chiu , Berlin Chen

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic…

Computation and Language · Computer Science 2021-02-08 Yanpei Shi , Thomas Hain

Music is a string of some of the notes out of 12 notes (Sa, Komal_re, Re, Komal_ga, Ga, Ma, Kari_ma, Pa, Komal_dha, Dha, Komal_ni, Ni) and their harmonics. Each note corresponds to a particular frequency. When such strings are encoded to…

Sound · Computer Science 2011-09-29 Avishek Ghosh , Joydeep Banerjee , Sk. S. Hassan , P. Pal Choudhury
‹ Prev 1 4 5 6 7 8 10 Next ›