English
Related papers

Related papers: HPP-Voice: A Large-Scale Evaluation of Speech Embe…

200 papers

This paper deals with a complex acoustic analysis of phonation in patients with Parkinson's disease (PD) with a special focus on estimation of disease progress that is described by 7 different clinical scales ,e. g. Unified Parkinson's…

Cross-modal associations between voice and face from a person can be learnt algorithmically, which can benefit a lot of applications. The problem can be defined as voice-face matching and retrieval tasks. Much research attention has been…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Chuyuan Xiong , Deyuan Zhang , Tao Liu , Xiaoyong Du

We present a comprehensive evaluation of pretrained speech embedding systems for the detection of dysarthric speech using existing accessible data. Dysarthric speech datasets are often small and can suffer from recording biases as well as…

Disordered speech recognition profound implications for improving the quality of life for individuals afflicted with, for example, dysarthria. Dysarthric speech recognition encounters challenges including limited data, substantial…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Yicong Jiang , Tianzi Wang , Xurong Xie , Juan Liu , Wei Sun , Nan Yan , Hui Chen , Lan Wang , Xunying Liu , Feng Tian

Depression is a common mental disorder which has been affecting millions of people around the world and becoming more severe with the arrival of COVID-19. Nevertheless proper diagnosis is not accessible in many regions due to a severe…

Multi-speaker spoken datasets enable the creation of text-to-speech synthesis (TTS) systems which can output several voice identities. The multi-speaker (MSPK) scenario also enables the use of fewer training samples per speaker. However, in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-04 Beata Lorincz , Adriana Stan , Mircea Giurgiu

Using speech samples as a biomarker is a promising avenue for detecting and monitoring the progression of Parkinson's disease (PD), but there is considerable disagreement in the literature about how best to collect and analyze such data.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-24 Peter Plantinga , Briac Cordelle , Dominique Louër , Mirco Ravanelli , Denise Klein

Traditional audiometry often provides an incomplete characterization of the functional impact of hearing loss on speech understanding, particularly for supra-threshold deficits common in presbycusis. This motivates the development of more…

Sound · Computer Science 2025-06-16 Stefan Bleeck

Depression detection from speech has attracted a lot of attention in recent years. However, the significance of speaker-specific information in depression detection has not yet been explored. In this work, we analyze the significance of…

Computers and Society · Computer Science 2021-07-30 Sri Harsha Dumpala , Sebastian Rodriguez , Sheri Rempel , Rudolf Uher , Sageev Oore

Speaker recognition performance in emotional talking environments is not as high as it is in neutral talking environments. This work focuses on proposing, implementing, and evaluating a new approach to enhance the performance in emotional…

Sound · Computer Science 2017-06-30 Ismail Shahin

Background: Depression is a major public health concern, affecting an estimated five percent of the global population. Early and accurate diagnosis is essential to initiate effective treatment, yet recognition remains challenging in many…

Signal Processing · Electrical Eng. & Systems 2025-11-21 Jana Weber , Marcel Weber , Juan Miguel Lopez Alcaraz

Parkinson's disease (PD) is a progressive neurodegenerative disorder that impacts motor functions and speech characteristics This study focuses on differentiating individuals with Parkinson's disease from healthy controls through the…

Machine Learning · Computer Science 2025-01-27 Burak Çelik , Ayhan Akbal

Personas are useful for dialogue response prediction. However, the personas used in current studies are pre-defined and hard to obtain before a conversation. To tackle this issue, we study a new task, named Speaker Persona Detection (SPD),…

Computation and Language · Computer Science 2021-09-06 Jia-Chen Gu , Zhen-Hua Ling , Yu Wu , Quan Liu , Zhigang Chen , Xiaodan Zhu

Speaker embedding has been a fundamental feature for speaker-related tasks such as verification, clustering, and diarization. Traditionally, speaker embeddings are represented as fixed vectors in high-dimensional space. This could lead to…

Sound · Computer Science 2022-06-28 Siqi Zheng , Hongbin Suo , Qian Chen

Speaker embeddings (x-vectors) extracted from very short segments of speech have recently been shown to give competitive performance in speaker diarization. We generalize this recipe by extracting from each speech segment, in parallel with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-09 Anna Silnova , Niko Brümmer , Johan Rohdin , Themos Stafylakis , Lukáš Burget

Intelligent systems are transforming the world, as well as our healthcare system. We propose a deep learning-based cough sound classification model that can distinguish between children with healthy versus pathological coughs such as…

Due to the substantial number of clinicians, patients, and data collection environments involved in clinical trials, gathering data of superior quality poses a significant challenge. In clinical trials, patients are assessed based on their…

Machine Learning · Computer Science 2024-04-09 Ali Akram , Marija Stanojevic , Malikeh Ehghaghi , Jekaterina Novikova

During psychiatric assessment, clinicians observe not only what patients report, but important nonverbal signs such as tone, speech rate, fluency, responsiveness, and body language. Weighing and integrating these different information…

Machine Learning · Computer Science 2025-12-19 Agnes Norbury , George Fairs , Alexandra L. Georgescu , Matthew M. Nour , Emilia Molimpakis , Stefano Goria

Voice trigger detection is an important task, which enables activating a voice assistant when a target user speaks a keyword phrase. A detector is typically trained on speech data independent of speaker information and used for the voice…