English
Related papers

Related papers: Detection of transitions between broad phonetic cl…

200 papers

Textless self-supervised speech models have grown in capabilities in recent years, but the nature of the linguistic information they encode has not yet been thoroughly examined. We evaluate the extent to which these models' learned…

Computation and Language · Computer Science 2023-06-13 Kinan Martin , Jon Gauthier , Canaan Breiss , Roger Levy

Formants are the spectral maxima that result from acoustic resonances of the human vocal tract, and their accurate estimation is among the most fundamental speech processing problems. Recent work has been shown that those frequencies can…

Sound · Computer Science 2022-06-24 Yosi Shrem , Felix Kreuk , Joseph Keshet

Acoustic sensing has proved effective as a foundation for numerous applications in health and human behavior analysis. In this work, we focus on the problem of detecting in-person social interactions in naturalistic settings from audio…

Sound · Computer Science 2022-03-23 Dawei Liang , Zifan Xu , Yinuo Chen , Rebecca Adaimi , David Harwath , Edison Thomaz

The problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear…

Computation and Language · Computer Science 2023-10-11 Piyush Singh Pasi , Karthikeya Battepati , Preethi Jyothi , Ganesh Ramakrishnan , Tanmay Mahapatra , Manoj Singh

Extraction of symbolic information from signals is an active field of research enabling numerous applications especially in the Musical Information Retrieval domain. This complex task, that is also related to other topics such as pitch…

Languages have long been described according to their perceived rhythmic attributes. The associated typologies are of interest in psycholinguistics as they partly predict newborns' abilities to discriminate between languages and provide…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-29 François Deloche , Laurent Bonnasse-Gahot , Judit Gervain

The focus of this paper is on classification of different vehicles using sound emanated from the vehicles. In this paper,quadratic discriminant analysis classifies audio signals of passing vehicles to bus, car, motor, and truck categories…

Sound · Computer Science 2018-04-05 Ali Dalir , Ali Asghar Beheshti , Morteza Hoseini Masoom

Transformer based end-to-end modelling approaches with multiple stream inputs have been achieved great success in various automatic speech recognition (ASR) tasks. An important issue associated with such approaches is that the intermediate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-11 Jin Li , Rongfeng Su , Xurong Xie , Nan Yan , Lan Wang

We present an approach for the detection of sharp change points (short-lived and persistent) in nonlinear and nonstationary dynamic systems under high levels of noise by tracking the local phase and amplitude synchronization among the…

Data Analysis, Statistics and Probability · Physics 2020-08-04 Ashif Sikandar Iquebal , Satish Bukkapatnam , Arun Srinivasa

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

We consider the problem of detecting, isolating and classifying elephant calls in continuously recorded audio. Such automatic call characterisation can assist conservation efforts and inform environmental management strategies. In contrast…

Sound · Computer Science 2025-04-03 Christiaan M. Geldenhuys , Thomas R. Niesler

Speech separation has been studied in time domain because of lower latency and higher performance compared to time-frequency domain. The masking-based method has been mostly used in time domain, and the other common method (mapping-based)…

Sound · Computer Science 2022-03-22 Chenyang Gao , Yue Gu , Ivan Marsic

Speech signals are complex composites of various information, including phonetic content, speaker traits, channel effect, etc. Decomposing this complicated mixture into independent factors, i.e., speech factorization, is fundamentally…

Sound · Computer Science 2019-10-30 Haoran Sun , Yunqi Cai , Lantian Li , Dong Wang

Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for…

Sound · Computer Science 2019-04-17 Gene-Ping Yang , Chao-I Tuan , Hung-Yi Lee , Lin-shan Lee

Voice activity and overlapped speech detection (respectively VAD and OSD) are key pre-processing tasks for speaker diarization. The final segmentation performance highly relies on the robustness of these sub-tasks. Recent studies have shown…

The performance of most emotion recognition systems degrades in real-life situations ('in the wild' scenarios) where the audio is contaminated by reverberation. Our study explores new methods to alleviate the performance degradation of SER…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Ohad Cohen , Gershon Hazan , Sharon Gannot

Previous work has shown that neural encoder-decoder speech recognition can be improved with hierarchical multitask learning, where auxiliary tasks are added at intermediate layers of a deep encoder. We explore the effect of hierarchical…

Computation and Language · Computer Science 2019-03-08 Kalpesh Krishna , Shubham Toshniwal , Karen Livescu

Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram…

Sound · Computer Science 2026-03-26 Qingkun Deng , Saturnino Luz , Sofia de la Fuente Garcia

Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current…

Noise-induced transitions between metastable fixed points in systems evolving on multiple time scales are analyzed in situations where the time scale separation gives rise to a slow manifold with bifurcation. This analysis is performed…

Statistical Mechanics · Physics 2017-10-05 Tobias Grafke , Eric Vanden-Eijnden
‹ Prev 1 4 5 6 7 8 10 Next ›