English
Related papers

Related papers: Vocal Breath Sound Based Gender Classification

200 papers

This paper presents a fully automated approach for identifying speech anomalies from voice recordings to aid in the assessment of speech impairments. By combining Connectionist Temporal Classification (CTC) and encoder-decoder-based…

Sound · Computer Science 2023-08-04 Laurin Wagner , Mario Zusag , Theresa Bloder

The range of potential applications of acoustic analysis is wide. Classification of sounds, in particular, is a typical machine learning task that received a lot of attention in recent years. The most common approaches to sound…

Cough is a primary symptom of most respiratory diseases, and changes in cough characteristics provide valuable information for diagnosing respiratory diseases. The characterization of cough sounds still lacks concrete evidence, which makes…

Sound · Computer Science 2023-08-08 Naveenkumar Vodnala , Pratap Reddy Lankireddy , Padmasai Yarlagadda

This paper investigates the temporal excitation patterns of creaky voice. Creaky voice is a voice quality frequently used as a phrase-boundary marker, but also as a means of portraying attitude, affective states and even social status.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-02 Thomas Drugman , John Kane , Christer Gobl

This paper examines gender and age salience and (stereo)typicality in British English talk with the aim to predict gender and age categories based on lexical, phrasal and turn-taking features. We examine the SpokenBNC, a corpus of around…

Computation and Language · Computer Science 2021-03-01 Andreas Liesenfeld , Gábor Parti , Yu-Yin Hsu , Chu-Ren Huang

Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, imprecise…

Theoretical background: early verbal development is not yet fully understood, especially in its formative phase. Research question: can a reliable, easy-to-use coding scheme for the classification of early infant vocalizations be defined…

A text-independent speaker recognition system relies on successfully encoding speech factors such as vocal pitch, intensity, and timbre to achieve good performance. A majority of such systems are trained and evaluated using spoken voice or…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Anurag Chowdhury , Austin Cozzo , Arun Ross

The speech signal is a consummate example of time-series data. The acoustics of the signal change over time, sometimes dramatically. Yet, the most common type of comparison we perform in phonetics is between instantaneous acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-18 Matthew C. Kelley

It is widely known that males and females typically possess different sound characteristics when singing, such as timbre and pitch, but it has never been explored whether these gender-based characteristics lead to a performance disparity in…

Sound · Computer Science 2023-08-08 Xiangming Gu , Wei Zeng , Ye Wang

This paper addresses the issue of cough detection using only audio recordings, with the ultimate goal of quantifying and qualifying the degree of pathology for patients suffering from respiratory diseases, notably mucoviscidosis. A large…

Sound · Computer Science 2020-01-06 Thomas Drugman , Jerome Urbain , Thierry Dutoit

Human speech contains paralinguistic cues that reflect a speaker's physiological and neurological state, potentially enabling non-invasive detection of various medical phenotypes. We introduce the Human Phenotype Project Voice corpus…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 David Krongauz , Hido Pinto , Sarah Kohn , Yanir Marmor , Eran Segal

In our previous work, we derived the acoustic features, that contribute to the perception of warmth and competence in synthetic speech. As an extension, in our current work, we investigate the impact of the derived vocal features in the…

Sound · Computer Science 2022-04-05 Sai Sirisha Rallabandi , Sebastian Möller

VoxCeleb datasets are widely used in speaker recognition studies. Our work serves two purposes. First, we provide speaker age labels and (an alternative) annotation of speaker gender. Second, we demonstrate the use of this metadata by…

Machine Learning · Computer Science 2021-12-21 Khaled Hechmi , Trung Ngo Trong , Ville Hautamaki , Tomi Kinnunen

In this paper, we present a detailed analysis on extracting soft biometric traits, age and gender, from ear images. Although there have been a few previous work on gender classification using ear images, to the best of our knowledge, this…

Computer Vision and Pattern Recognition · Computer Science 2018-06-18 Dogucan Yaman , Fevziye Irem Eyiokur , Nurdan Sezgin , Hazım Kemal Ekenel

The Mean Opinion Score (MOS) serves as the standard metric for speech quality assessment, yet biases in human annotations remain underexplored. We conduct the first systematic analysis of gender bias in MOS, revealing that male listeners…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Wenze Ren , Yi-Cheng Lin , Wen-Chin Huang , Erica Cooper , Ryandhimas E. Zezario , Hsin-Min Wang , Hung-yi Lee , Yu Tsao

Singing voices contain much richer information than common voices, including varied vocal and acoustic properties. However, current open-source audio-text datasets for singing voices capture only a narrow range of attributes and lack…

Computation and Language · Computer Science 2025-08-19 Hyunjong Ok , Jaeho Lee

While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work has indicated that…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 Nick Mehlman , Thomas Thebaud , Dani Byrd , Shri Narayanan

Most previous studies on automatic recognition model for bipolar disorder (BD) were based on both social media and linguistic features. The present study investigates the possibility of adopting only language-based features, namely the…

Information Retrieval · Computer Science 2019-07-18 Yen-Hao Huang , Yi-Hsin Chen , Fernando Henrique Calderon Alvarado , Ssu-Rui Lee , Shu-I Wu , Yuwen Lai , Yi-Shin Chen

During voiced speech, the human vocal folds interact with the vocal tract acoustics. The resulting glottal source-resonator coupling has been observed using mathematical and physical models as well as in in vivo phonation. We propose a…

Fluid Dynamics · Physics 2017-03-16 Atte Aalto , Tiina Murtola , Jarmo Malinen , Daniel Aalto , Martti Vainio