English
Related papers

Related papers: Improving Factored Hybrid HMM Acoustic Modeling wi…

200 papers

This paper presents a new and flexible prognostics framework based on a higher order hidden semi-Markov model (HOHSMM) for systems or components with unobservable health states and complex transition dynamics. The HOHSMM extends the basic…

Applications · Statistics 2020-02-14 Ying Liao , Yisha Xiang , Min Wang

The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. While most prior work frames spoof detection as a binary…

Sound · Computer Science 2025-09-10 Bin Hu , Kunyang Huang , Daehan Kwak , Meng Xu , Kuan Huang

Hybrid State-Space models combine Attention with recurrent State-Space Model (SSM) layers, balancing eidetic memory from Attention with compressed fading memory from SSMs. This yields smaller Key-Value caches and faster decoding than…

Direct acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional sub-word based automatic speech recognition models using phones, characters, or context-dependent hidden Markov…

Computation and Language · Computer Science 2017-12-11 Kartik Audhkhasi , Brian Kingsbury , Bhuvana Ramabhadran , George Saon , Michael Picheny

The Hidden Markov Model (HMM) can predict the future value of a time series based on its current and previous values, making it a powerful algorithm for handling various types of time series. Numerous studies have explored the improvement…

Machine Learning · Computer Science 2024-02-28 YeXin Huang

It is well known that speaker identification yields very high performance in a neutral talking environment, on the other hand, the performance has been sharply declined in a shouted talking environment. This work aims at proposing,…

Sound · Computer Science 2017-07-07 Ismail Shahin

Token-based text-to-speech (TTS) models have emerged as a promising avenue for generating natural and realistic speech, yet they grapple with low pronunciation accuracy, speaking style and timbre inconsistency, and a substantial need for…

Sound · Computer Science 2024-03-12 Chunhui Wang , Chang Zeng , Bowen Zhang , Ziyang Ma , Yefan Zhu , Zifeng Cai , Jian Zhao , Zhonglin Jiang , Yong Chen

Automatically inducing the syntactic part-of-speech categories for words in text is a fundamental task in Computational Linguistics. While the performance of unsupervised tagging models has been slowly improving, current state-of-the-art…

Computation and Language · Computer Science 2014-02-27 Greg Dubbin , Phil Blunsom

We present HAFM, a system that generates instrumental music audio to accompany input vocals. Given isolated singing voice, HAFM produces a coherent instrumental accompaniment that can be directly mixed with the input to create complete…

Sound · Computer Science 2026-04-14 Jian Zhu , Jianwei Cui , Shihao Chen , Yubang Zhang , Cheng Luo

Most multi-domain machine translation models rely on domain-annotated data. Unfortunately, domain labels are usually unavailable in both training processes and real translation scenarios. In this work, we propose a label-free multi-domain…

Computation and Language · Computer Science 2023-05-09 Fan Zhang , Mei Tu , Sangha Kim , Song Liu , Jinyao Yan

This paper investigates a non-negative matrix factorization (NMF)-based approach to the semi-supervised single-channel speech enhancement problem where only non-stationary additive noise signals are given. The proposed method relies on…

Sound · Computer Science 2013-09-25 Nikolay Lyubimov , Mikhail Kotov

In this paper, we present several adaptation methods for non-native speech recognition. We have tested pronunciation modelling, MLLR and MAP non-native pronunciation adaptation and HMM models retraining on the HIWIRE foreign accented…

Computation and Language · Computer Science 2007-11-07 Ghazi Bouselmi , Dominique Fohr , Irina Illina

Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-17 Orchid Chetia Phukan , Swarup Ranjan Behera , Girish , Mohd Mujtaba Akhtar , Arun Balaji Buduru , Rajesh Sharma

Modeling continuous-time physiological processes that manifest a patient's evolving clinical states is a key step in approaching many problems in healthcare. In this paper, we develop the Hidden Absorbing Semi-Markov Model (HASMM): a…

Artificial Intelligence · Computer Science 2016-12-28 Ahmed M. Alaa , Mihaela van der Schaar

This work attempts to approximate a linear Gaussian system with a finite-state hidden Markov model (HMM), which is found useful in solving sophisticated event-based state estimation problems. An indirect modeling approach is developed,…

Systems and Control · Electrical Eng. & Systems 2020-07-10 Kaikai Zheng , Dawei Shi , Ling Shi

A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel internal acoustic model (IAM) training strategy to enhance…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-01 Takafumi Moriya , Takanori Ashihara , Masato Mimura , Hiroshi Sato , Kohei Matsuura , Ryo Masumura , Taichi Asami

Learning processes by exploiting restricted domain knowledge is an important task across a plethora of scientific areas, with more and more hybrid training methods additively combining data-driven and model-based approaches. Although the…

Machine Learning · Computer Science 2025-01-17 Yann Claes , Vân Anh Huynh-Thu , Pierre Geurts

With the increased popularity of smart phones, there is a greater need to have a robust authentication mechanism that handles various security threats and privacy leakages effectively. This paper studies continuous authentication for touch…

Cryptography and Security · Computer Science 2017-12-25 Aditi Roy , Tzipora Halevi , Nasir Memon

The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and…

Sound · Computer Science 2025-05-22 Kunyang Huang , Bin Hu

Given a nonparametric Hidden Markov Model (HMM) with two states, the question of constructing efficient multiple testing procedures is considered, treating one of the states as an unknown null hypothesis. A procedure is introduced, based on…

Statistics Theory · Mathematics 2021-01-12 Kweku Abraham , Ismael Castillo , Elisabeth Gassiat