English
Related papers

Related papers: Speech-based Age and Gender Prediction with Transf…

200 papers

Recent research has shown that state-of-the-art (SotA) Automatic Speech Recognition (ASR) systems, such as Whisper, often exhibit predictive biases that disproportionately affect various demographic groups. This study focuses on identifying…

Computation and Language · Computer Science 2024-11-15 Rik Raes , Saskia Lensink , Mykola Pechenizkiy

Existing methods for speaker age estimation usually treat it as a multi-class classification or a regression problem. However, precise age identification remains a challenge due to label ambiguity, \emph{i.e.}, utterances from adjacent age…

Sound · Computer Science 2022-02-24 Shijing Si , Jianzong Wang , Junqing Peng , Jing Xiao

While demographic factors like age and gender change the way people talk, and in particular, the way people talk to machines, there is little investigation into how large pre-trained language models (LMs) can adapt to these changes. To…

Computation and Language · Computer Science 2024-02-06 Anthony Sicilia , Jennifer C. Gates , Malihe Alikhani

Facial age estimation has achieved considerable success under controlled conditions. However, in unconstrained real-world scenarios, which are often referred to as 'in the wild', age estimation remains challenging, especially when faces are…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Waqar Tanveer , Laura Fernández-Robles , Eduardo Fidalgo , Víctor González-Castro , Enrique Alegre

This work studies the information freshness of the vehicle-to-infrastructure status updating in Internet of vehicles, which is modeled as a multi-source Ber/Geo/1/1 preemptive queueing system with heterogeneous service time. We pay…

Information Theory · Computer Science 2023-08-22 Tianci Zhang , Zhengchuan Chen

Facial analysis permits many investigations some of the most important of which are craniofacial identification, facial recognition, and age and sex estimation. In forensics, photo-anthropometry describes the study of facial growth and…

Computer Vision and Pattern Recognition · Computer Science 2020-09-10 Lucas F. Porto , Laise N. Correia Lima , Ademir Franco , Donald M. Pianto , Carlos Eduardo Machado Palhares , Donald M. Pianto , Flavio de Barros Vidal

Self-supervised learning (SSL) is a powerful tool that allows learning of underlying representations from unlabeled data. Transformer based models such as wav2vec 2.0 and HuBERT are leading the field in the speech domain. Generally these…

Computation and Language · Computer Science 2022-02-08 Bethan Thomas , Samuel Kessler , Salah Karout

Predictive modeling using structural magnetic resonance imaging (MRI) data is a prominent approach to study brain-aging. Machine learning algorithms and feature extraction methods have been employed to improve predictions and explore…

Machine Learning · Computer Science 2025-01-20 Georgios Antonopoulos , Shammi More , Simon B. Eickhoff , Federico Raimondo , Kaustubh R. Patil

Non-reference speech quality models are important for a growing number of applications. The VoiceMOS 2022 challenge provided a dataset of synthetic voice conversion and text-to-speech samples with subjective labels. This study looks at the…

Sound · Computer Science 2022-09-15 Michael Chinen , Jan Skoglund , Chandan K A Reddy , Alessandro Ragano , Andrew Hines

Gender bias in machine translation (MT) systems has been extensively documented, but bias in automatic quality estimation (QE) metrics remains comparatively underexplored. Existing studies suggest that QE metrics can also exhibit gender…

Computation and Language · Computer Science 2025-10-09 Giorgos Filandrianos , Orfeas Menis Mastromichalakis , Wafaa Mohammed , Giuseppe Attanasio , Chrysoula Zerva

Age estimation has attracted attention for its various medical applications. There are many studies on human age estimation from biomedical images. However, there is no research done on mammograms for age estimation, as far as we know. The…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Charitha Dissanayake Lekamlage , Fabia Afzal , Erik Westerberg , Abbas Cheddad

Gender recognition is an essential component of automatic speech recognition and interactive voice response systems. Determining gender of the speaker reduces the computational burden of such systems for any further processing. Typical…

Sound · Computer Science 2016-01-08 Jamil Ahmad , Mustansar Fiaz , Soon-il Kwon , Maleerat Sodanil , Bay Vo , Sung Wook Baik

Recently, audio-visual speech enhancement has been tackled in the unsupervised settings based on variational auto-encoders (VAEs), where during training only clean data is used to train a generative model for speech, which at test time is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Mostafa Sadeghi , Xavier Alameda-Pineda

We propose a novel transfer learning method for speech emotion recognition allowing us to obtain promising results when only few training data is available. With as low as 125 examples per emotion class, we were able to reach a higher…

Machine Learning · Computer Science 2020-11-12 Jonathan Boigne , Biman Liyanage , Ted Östrem

ASR systems designed for native English (L1) usually underperform on non-native English (L2). To address this performance gap, \textbf{(i)} we extend our previous work to investigate fine-tuning of a pre-trained wav2vec 2.0 model…

Computation and Language · Computer Science 2022-02-11 Peter Sullivan , Toshiko Shibano , Muhammad Abdul-Mageed

There has been a recent surge of interest in time series modeling using the Transformer architecture. However, forecasting multivariate time series with Transformer presents a unique challenge as it requires modeling both temporal…

Machine Learning · Computer Science 2025-07-04 Yu-Hsiang Lan , Eric K. Oermann

Pre-trained speech Transformers have facilitated great success across various speech processing tasks. However, fine-tuning these encoders for downstream tasks require sufficiently large training data to converge or to achieve…

Computation and Language · Computer Science 2022-10-25 Hao Yang , Jinming Zhao , Gholamreza Haffari , Ehsan Shareghi

Research on non-verbal behavior generation for social interactive agents focuses mainly on the believability and synchronization of non-verbal cues with speech. However, existing models, predominantly based on deep learning architectures,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Alice Delbosc , Magalie Ochs , Nicolas Sabouret , Brian Ravenet , Stephane Ayache

Automatic speaker verification has achieved remarkable progress in recent years. However, there is little research on cross-age speaker verification (CASV) due to insufficient relevant data. In this paper, we mine cross-age test sets based…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-14 Xiaoyi Qin , Na Li , Chao Weng , Dan Su , Ming Li

In this paper, we analyze status update systems modeled through the Stochastic Hybrid Systems (SHSs) tool. Contrary to previous works, we allow the system's transition dynamics to be polynomial functions of the Age of Information (AoI).…

Information Theory · Computer Science 2022-04-29 Ali Maatouk , Mohamad Assaad , Anthony Ephremides