English
Related papers

Related papers: Detection of transitions between broad phonetic cl…

200 papers

We study a phase transition in parameter learning of Hidden Markov Models (HMMs). We do this by generating sequences of observed symbols from given discrete HMMs with uniformly distributed transition probabilities and a noise level encoded…

Statistical Mechanics · Physics 2021-10-13 Nikita Rau , Jörg Lücke , Alexander K. Hartmann

In optical wireless scattering communication, received signal in each symbol interval is captured by a photomultiplier tube (PMT) and then sampled through very short but finite interval sampling. The resulting samples form a signal vector…

Information Theory · Computer Science 2016-12-14 Difan Zou , Chen Gong , Zhengyuan Xu

Many astronomical sources produce transient phenomena at radio frequencies, but the transient sky at low frequencies (<300 MHz) remains relatively unexplored. Blind surveys with new widefield radio instruments are setting increasingly…

Time delay estimation or Time-Difference-Of-Arrival estimates is a critical component for multiple localization applications such as multilateration, direction of arrival, and self-calibration. The task is to estimate the time difference…

Sound · Computer Science 2024-11-21 Erik Tegler , Magnus Oskarsson , Kalle Åström

Recent advances in song identification leverage deep neural networks to learn compact audio fingerprints directly from raw waveforms. While these methods perform well under controlled conditions, their accuracy drops significantly in…

Sound · Computer Science 2025-09-16 Christos Nikou , Theodoros Giannakopoulos

We introduce an unsupervised approach for correcting highly imperfect speech transcriptions based on a decision-level fusion of stemming and two-way phoneme pruning. Transcripts are acquired from videos by extracting audio using Ffmpeg…

Computation and Language · Computer Science 2021-07-28 Sunakshi Mehra , Seba Susan

Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the splitand-classify (i.e., frame-level) strategy or the more…

Sound · Computer Science 2023-08-21 Swapnil Bhosale , Sauradip Nag , Diptesh Kanojia , Jiankang Deng , Xiatian Zhu

Many voice disorders induce subharmonic phonation, but voice signal analysis is currently lacking a technique to detect the presence of subharmonics reliably. Distinguishing subharmonic phonation from normal phonation is a challenging task…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-17 Takeshi Ikuma , Melda Kunduk , Brad Story , Andrew J. McWhorter

Interactive voice assistants have been widely used as input interfaces in various scenarios, e.g. on smart homes devices, wearables and on AR devices. Detecting the end of a speech query, i.e. speech end-pointing, is an important task for…

Sound · Computer Science 2022-10-27 Dawei Liang , Hang Su , Tarun Singh , Jay Mahadeokar , Shanil Puri , Jiedan Zhu , Edison Thomaz , Mike Seltzer

Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by…

Sound · Computer Science 2023-10-18 Fernando López , Jordi Luque , Carlos Segura , Pablo Gómez

The detection of pathologies from speech features is usually defined as a binary classification task with one class representing a specific pathology and the other class representing healthy speech. In this work, we train neural networks,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-02 Dominik Wagner , Ilja Baumann , Franziska Braun , Sebastian P. Bayerl , Elmar Nöth , Korbinian Riedhammer , Tobias Bocklet

Detecting critical transitions in complex, noisy time-series data is a fundamental challenge across science and engineering. Such transitions may be anticipated by the emergence of a low-dimensional order parameter, whose signature is often…

Machine Learning · Computer Science 2025-12-16 Wenqi Fang , Ye Li

In label-noise learning, the transition matrix plays a key role in building statistically consistent classifiers. Existing consistent estimators for the transition matrix have been developed by exploiting anchor points. However, the…

Machine Learning · Computer Science 2021-10-22 Xuefeng Li , Tongliang Liu , Bo Han , Gang Niu , Masashi Sugiyama

Speech disfluency commonly occurs in conversational and spontaneous speech. However, standard Automatic Speech Recognition (ASR) models struggle to accurately recognize these disfluencies because they are typically trained on fluent…

Computation and Language · Computer Science 2024-09-18 Robin Amann , Zhaolin Li , Barbara Bruno , Jan Niehues

A human action can be seen as transitions between one's body poses over time, where the transition depicts a temporal relation between two poses. Recognizing actions thus involves learning a classifier sensitive to these pose transitions as…

Computer Vision and Pattern Recognition · Computer Science 2017-04-03 Guillermo Garcia-Hernando , Tae-Kyun Kim

In this paper, we analyse the error patterns of the raw waveform acoustic models in TIMIT's phone recognition task. Our analysis goes beyond the conventional phone error rate (PER) metric. We categorise the phones into three groups:…

Sound · Computer Science 2024-06-04 Erfan Loweimi , Andrea Carmantini , Peter Bell , Steve Renals , Zoran Cvetkovic

The study of speech disorders can benefit greatly from time-aligned data. However, audio-text mismatches in disfluent speech cause rapid performance degradation for modern speech aligners, hindering the use of automatic approaches. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Theodoros Kouzelis , Georgios Paraskevopoulos , Athanasios Katsamanis , Vassilis Katsouros

The associationist account for early word-learning is based on the co-occurrence between objects and words. Here we examine the performance of a simple associative learning algorithm for acquiring the referents of words in a…

Physics and Society · Physics 2012-10-16 P. F. C. Tilles , J. F. Fontanari

We introduce a novel and inexpensive approach for the temporal alignment of speech to highly imperfect transcripts from automatic speech recognition (ASR). Transcripts are generated for extended lecture and presentation videos, which in…

Sound · Computer Science 2007-05-23 Alexander Haubold , John R. Kender

Language identification is a crucial first step in multilingual systems such as chatbots and virtual assistants, enabling linguistically and culturally accurate user experiences. Errors at this stage can cascade into downstream failures,…