English
Related papers

Related papers: Analyzing long-term rhythm variations in Mising an…

200 papers

The data heterogeneity across devices and the limited communication resources, e.g., bandwidth and energy, are two of the main bottlenecks for wireless federated learning (FL). To tackle these challenges, we first devise a novel FL…

Machine Learning · Computer Science 2023-02-21 Zhixiong Chen , Wenqiang Yi , Arumugam Nallanathan , Geoffrey Ye Li

Affine frequency division multiplexing (AFDM) is a recently proposed communication waveform for time-varying channel scenarios. As a chirp-based multicarrier modulation technique it can not only satisfy the needs of multiple scenarios in…

Signal Processing · Electrical Eng. & Systems 2024-01-01 Jiajun Zhu , Yanqun Tang , Xizhang Wei , Haoran Yin , Jinming Du , Zhengpeng Wang , Yuqinng Liu

Recently, pre-trained models with phonetic supervision have demonstrated their advantages for crosslingual speech recognition in data efficiency and information sharing across languages. However, a limitation is that a pronunciation lexicon…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-17 Saierdaer Yusuyin , Te Ma , Hao Huang , Zhijian Ou

In molecular communications (MC), inter-symbol interference (ISI) and noise are key factors that degrade communication reliability. Although time-domain equalization can effectively mitigate these effects, it often entails high…

Subcellular Processes · Quantitative Biology 2025-11-25 Cheng Xiang , Yu Huang , Miaowen Wen , Weiqiang Tan , Chan-Byoung Chae

Word frequency is a strong predictor in most lexical processing tasks. Thus, any model of word recognition needs to account for how word frequency effects arise. The Discriminative Lexicon Model (DLM; Baayen et al., 2018a, 2019) models…

Computation and Language · Computer Science 2024-03-19 Maria Heitmeier , Yu-Ying Chuang , Seth D. Axen , R. Harald Baayen

Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and…

Sound · Computer Science 2026-02-27 Sanjid Hasan , Risalat Labib , A H M Fuad , Bayazid Hasan

The performance of text-to-speech (TTS) systems heavily depends on spectrogram to waveform generation, also known as the speech reconstruction phase. The time required for the same is known as synthesis delay. In this paper, an approach to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Ankit Sharma , Puneet Kumar , Vikas Maddukuri , Nagasai Madamshettib , Kishore KG , Sahit Sai Sriram Kavurub , Balasubramanian Raman , Partha Pratim Roy

A natural sound can be described by dynamic changes in envelope (amplitude) and carrier (frequency), corresponding to amplitude modulation (AM) and frequency modulation (FM) respectively. Although the neural responses to both AM and FM…

Neurons and Cognition · Quantitative Biology 2007-05-23 Huan Luo , Yadong Wang , David Poeppel , Jonathan Z. Simon

This article presents a whisper speech detector in the far-field domain. The proposed system consists of a long-short term memory (LSTM) neural network trained on log-filterbank energy (LFBE) acoustic features. This model is trained and…

Computation and Language · Computer Science 2020-04-07 Zeynab Raeesy , Kellen Gillespie , Zhenpei Yang , Chengyuan Ma , Thomas Drugman , Jiacheng Gu , Roland Maas , Ariya Rastrow , Björn Hoffmeister

Detrended fluctuation analysis (DFA) is a scaling analysis method used to quantify long-range power-law correlations in signals. Many physical and biological signals are ``noisy'', heterogeneous and exhibit different types of…

Data Analysis, Statistics and Probability · Physics 2009-11-07 Zhi Chen , Plamen Ch. Ivanov , Kun Hu , H. Eugene Stanley

Feature-mapping with deep neural networks is commonly used for single-channel speech enhancement, in which a feature-mapping network directly transforms the noisy features to the corresponding enhanced ones and is trained to minimize the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-01 Zhong Meng , Jinyu Li , Yifan Gong , Biing-Hwang , Juang

In music information retrieval (MIR) research, the use of pretrained foundational audio encoders (FAEs) has recently become a trend. FAEs pretrained on large amounts of music and audio data have been shown to improve performance on MIR…

Sound · Computer Science 2026-01-30 Keisuke Toyama , Zhi Zhong , Akira Takahashi , Shusuke Takahashi , Yuki Mitsufuji

The acoustic wave-propagation without mean flow and heat flux can be described in terms of velocity and pressure by the compressible nonlinear Navier-Stokes equations, where boundary layers appear at walls due to the viscosity and a…

Analysis of PDEs · Mathematics 2017-01-10 Anastasia Thoens-Zueva , Kersten Schmidt , Adrien Semin

Evolutionary models of languages are usually considered to take the form of trees. With the development of so-called tree constraints the plausibility of the tree model assumptions can be addressed by checking whether the moments of…

Applications · Statistics 2014-10-06 Nathaniel Shiers , John A. D. Aston , Jim Q. Smith , John S. Coleman

Online abusive content detection, particularly in low-resource settings and within the audio modality, remains underexplored. We investigate the potential of pre-trained audio representations for detecting abusive language in low-resource…

Computation and Language · Computer Science 2024-12-16 Aditya Narayan Sankaran , Reza Farahbakhsh , Noel Crespi

In clinical voice signal analysis, mishandling of subharmonic voicing may cause an acoustic parameter to signal false negatives. As such, the ability of a fundamental frequency estimator to identify speaking fundamental frequency is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-10 Takeshi Ikuma , Melda Kunduk , Andrew J. McWhorter

Developing a practical speech recognizer for a low resource language is challenging, not only because of the (potentially unknown) properties of the language, but also because test data may not be from the same domain as the available…

Computation and Language · Computer Science 2018-10-02 Siddharth Dalmia , Xinjian Li , Florian Metze , Alan W. Black

Time-series forecasting in real-world applications such as finance and energy often faces challenges due to limited training data and complex, noisy temporal dynamics. Existing deep forecasting models typically supervise predictions using…

Machine Learning · Computer Science 2026-01-14 Jiacheng You , Jingcheng Yang , Yuhang Xie , Zhongxuan Wu , Xiucheng Li , Feng Li , Pengjie Wang , Jian Xu , Bo Zheng , Xinyang Chen

This work presents a seemingly simple but effective technique to improve low-resource ASR systems for phonetic languages. By identifying sets of acoustically similar graphemes in these languages, we first reduce the output alphabet of the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-13 Anuj Diwan , Preethi Jyothi

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

Computation and Language · Computer Science 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich