English
Related papers

Related papers: Discovering phoneme-specific critical articulators…

200 papers

Though significant progress has been made for the voice conversion (VC) of typical speech, VC for atypical speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for atypical prosody…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-26 Disong Wang , Songxiang Liu , Lifa Sun , Xixin Wu , Xunying Liu , Helen Meng

One of the major challenges in acoustic modelling of child speech is the rapid changes that occur in the children's articulators as they grow up, their differing growth rates and the subsequent high variability in the same age group. These…

Sound · Computer Science 2022-11-08 Mostafa Shahin , Beena Ahmed , Julien Epps

Acoustic-to-articulatory speech inversion could enhance automated clinical mispronunciation detection to provide detailed articulatory feedback unattainable by formant-based mispronunciation detection algorithms; however, it is unclear the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-22 Nina R Benway , Yashish M Siriwardena , Jonathan L Preston , Elaine Hitchcock , Tara McAllister , Carol Espy-Wilson

Transformers have seen an unprecedented rise in Natural Language Processing and Computer Vision tasks. However, in audio tasks, they are either infeasible to train due to extremely large sequence length of audio waveforms or incur a…

Machine Learning · Computer Science 2022-02-02 Surya Kant Sahu , Sai Mitheran , Juhi Kamdar , Meet Gandhi

Attention-based encoder-decoder (AED) models learn an implicit internal language model (ILM) from the training transcriptions. The integration with an external LM trained on much more unpaired text usually leads to better performance. A…

Computation and Language · Computer Science 2021-06-18 Mohammad Zeineldeen , Aleksandr Glushko , Wilfried Michel , Albert Zeyer , Ralf Schlüter , Hermann Ney

Tokenization algorithms that merge the units of a base vocabulary into larger, variable-rate units have become standard in natural language processing tasks. This idea, however, has been mostly overlooked when the vocabulary consists of…

Sound · Computer Science 2024-06-11 Avihu Dekel , Raul Fernandez

Contextual automatic speech recognition, i.e., biasing recognition towards a given context (e.g. user's playlists, or contacts), is challenging in end-to-end (E2E) models. Such models maintain a limited number of candidates during…

Computation and Language · Computer Science 2019-07-23 Ke Hu , Antoine Bruguier , Tara N. Sainath , Rohit Prabhavalkar , Golan Pundak

Accent variability remains a major errors in automatic speech recognition, yet most adaptation methods rely on parameter fine-tuning without understanding where accent information is encoded. We treat accent variation as an interpretable…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-09 Jinuo Sun , Yang Xiao , Sung Kyun Chung , Qiuchi Hu , Gongping Huang , Eun-Jung Holden , Ting Dang

Cochlear implant (CI) users have considerable difficulty in understanding speech in reverberant listening environments. Time-frequency (T-F) masking is a common technique that aims to improve speech intelligibility by multiplying…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-01 Kevin M. Chu , Leslie M. Collins , Boyla O. Mainsah

Phones, the segmental units of the International Phonetic Alphabet (IPA), are used for lexical distinctions in most human languages; Tones, the suprasegmental units of the IPA, are used in perhaps 70%. Many previous studies have explored…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-29 Jialu Li , Mark Hasegawa-Johnson

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-12 Zehai Tu , Jack Deadman , Ning Ma , Jon Barker

Effective human-AI collaboration hinges on the ability to dynamically integrate the complementary strengths of human experts and AI models across diverse decision contexts. Context-aware weighted combination of human and AI outputs is a…

Human-Computer Interaction · Computer Science 2025-11-07 Renlong Jie

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

State of the art speech recognition systems use data-intensive context-dependent phonemes as acoustic units. However, these approaches do not translate well to low resourced languages where large amounts of training data is not available.…

Computation and Language · Computer Science 2016-06-21 Amir Hossein Harati Nejad Torbati , Joseph Picone

Accurately extracting clinical information from speech is critical to the diagnosis and treatment of many neurological conditions. As such, there is interest in leveraging AI for automatic, objective assessments of clinical speech to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Daniela A. Wiepert , Rene L. Utianski , Joseph R. Duffy , John L. Stricker , Leland R. Barnard , David T. Jones , Hugo Botha

In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked…

Sound · Computer Science 2024-10-18 Ashish Seth , Ramaneswaran Selvakumar , S Sakshi , Sonal Kumar , Sreyan Ghosh , Dinesh Manocha

Comparing spoken segments is a central operation to speech processing. Traditional approaches in this area have favored frame-level dynamic programming algorithms, such as dynamic time warping, because they require no supervision, but they…

Computation and Language · Computer Science 2023-08-30 Shane Settle

In the last decades several multi-microphone speech dereverberation algorithms have been proposed, among which the weighted prediction error (WPE) algorithm. In the WPE algorithm, a prediction delay is required to reduce the correlation…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Anselm Lohmann , Toon van Waterschoot , Joerg Bitzer , Simon Doclo

Acoustic word embeddings (AWEs) are vector representations such that different acoustic exemplars of the same word are projected nearby in the embedding space. In addition to their use in speech technology applications such as spoken term…

Computation and Language · Computer Science 2023-01-10 Badr M. Abdullah , Dietrich Klakow

Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained features like exact tempo or pitch, effective integration…

Sound · Computer Science 2026-02-17 Siqian Tong , Xuan Li , Yiwei Wang , Baolong Bi , Yujun Cai , Shenghua Liu , Yuchen He , Chengpeng Hao
‹ Prev 1 8 9 10 Next ›