English
Related papers

Related papers: Software Corrections of Vocal Disorders

200 papers

Nonlinearities can be introduced into communication systems by the physical components such as the power amplifier, or during signal propagation through a nonlinear channel. These nonlinearities can be compensated by a nonlinear equalizer…

Signal Processing · Electrical Eng. & Systems 2019-05-15 Etsushi Yamazaki , Nariman Farsad , Andrea Goldsmith

Machine learning algorithms, when trained on audio recordings from a limited set of devices, may not generalize well to samples recorded using other devices with different frequency responses. In this work, a relatively straightforward…

Sound · Computer Science 2021-05-26 Michał Kośmider

Distortion of the underlying speech is a common problem for single-channel speech enhancement algorithms, and hinders such methods from being used more extensively. A dictionary based speech enhancement method that emphasizes preserving the…

Sound · Computer Science 2016-05-09 Eunjoon Cho , Bowon Lee , Ronald Schafer , Bernard Widrow

State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Adam Gabryś , Goeric Huybrechts , Manuel Sam Ribeiro , Chung-Ming Chien , Julian Roth , Giulia Comini , Roberto Barra-Chicote , Bartek Perz , Jaime Lorenzo-Trueba

This paper addresses the challenge of enhancing the realism of vocoder-generated singing voice audio by mitigating the distinguishable disparities between synthetic and real-life recordings, particularly in high-frequency spectrogram…

Sound · Computer Science 2025-08-05 Runxuan Yang , Kai Li , Guo Chen , Xiaolin Hu

Source-tract decomposition (or glottal flow estimation) is one of the basic problems of speech processing. For this, several techniques have been proposed in the literature. However studies comparing different approaches are almost…

Sound · Computer Science 2020-01-06 Thomas Drugman , Baris Bozkurt , Thierry Dutoit

Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prompt. Previous work has attempted to address this issue by steering the model using guidance…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Anne Harrington , A. Sophia Koepke , Shyamgopal Karthik , Trevor Darrell , Alexei A. Efros

This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic…

Sound · Computer Science 2025-07-16 Andrew Valdivia , Yueming Zhang , Hailu Xu , Amir Ghasemkhani , Xin Qin

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

Computation and Language · Computer Science 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

In a typical voice conversion system, vocoder is commonly used for speech-to-features analysis and features-to-speech synthesis. However, vocoder can be a source of speech quality degradation. This paper presents a vocoder-free voice…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-18 Xiaohai Tian , Eng Siong Chng , Haizhou Li

This article focuses on techniques for acoustic noise reduction, signal filters and source reconstruction. For noise reduction, bandpass filters and cross correlations are found to be efficient and fast ways to improve the signal to noise…

Instrumentation and Methods for Astrophysics · Physics 2009-06-10 C. Richardt , G. Anton , K. Graf , J. Hoessl , A. Kappes , U. Katz , R. Lahmann , Ch. Naumann , M. Neff , F. Schoeck

The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech…

Audio codecs are typically transform-domain based and efficiently code stationary audio signals, but they struggle with speech and signals containing dense transient events such as applause. Specifically, with these two classes of signals…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-28 Arijit Biswas , Dai Jia

The solution of the problem of assessing the quality of the pronunciation of syllables during speech rehabilitation after surgical treatment of oncological diseases of the organs of the speech-forming tract is considered in the work. The…

Machine Learning · Computer Science 2023-01-26 Evgeny Kostyuchenko

In this paper, we propose a classification based glottal closure instants (GCI) detection from pathological acoustic speech signal, which finds many applications in vocal disorder analysis. Till date, GCI for pathological disorder is…

Sound · Computer Science 2018-11-28 Gurunath Reddy M , Tanumay Mandal , Krothapalli Sreenivasa Rao

Audio or visual data analysis tasks usually have to deal with high-dimensional and nonnegative signals. However, most data analysis methods suffer from overfitting and numerical problems when data have more than a few dimensions needing a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-24 Sergio Muñoz-Romero , Jerónimo Arenas García , Vanessa Gómez-Verdejo

The physical understanding of a method of detecting mammalian cancer via vocalization during a normal echo-Doppler test is provided. The backscattered ultrasound frequency in the case of a vocal humming resonating in the chest wall is…

Quantum noise in a model of singly resonant frequency doubling including phase mismatch and driving in the harmonic mode is analyzed. The general formulae about the fixed points and their stability as well as the squeezing spectra…

Quantum Physics · Physics 2007-05-23 C. Cabrillo , J. L. Roldan , P. Garcia-Fernandez

The availability of digital devices operated by voice is expanding rapidly. However, the applications of voice interfaces are still restricted. For example, speaking in public places becomes an annoyance to the surrounding people, and…

Human-Computer Interaction · Computer Science 2023-03-06 Naoki Kimura , Michinari Kono , Jun Rekimoto

The total variation filtering technique emerges as a highly effective strategy for restoring signals with discontinuities in various parts of their structure. This study presents and implements a one-dimensional signal filtering algorithm…

Optimization and Control · Mathematics 2024-10-14 Joyce Oliveira dos Santos , Francisco Márcio Barboza