English
Related papers

Related papers: Pre-Trained Foundation Model representations to un…

200 papers

Compared with invasive examinations that require tissue sampling, respiratory sound testing is a non-invasive examination method that is safer and easier for patients to accept. In this study, we introduce Rene, a pioneering large-scale…

Sound · Computer Science 2024-06-10 Pengfei Zhang , Zhihang Zheng , Shichen Zhang , Minghao Yang , Shaojun Tang

Estimating human vital signs in a contactless non-invasive method using radar provides a convenient method in the medical field to conduct several health checkups easily and quickly. In addition to monitoring while sitting and sleeping, the…

Signal Processing · Electrical Eng. & Systems 2022-06-14 Tassneem Helal , Fady Aziz , Omar Metwally , Marco F. Huber , Dominik Alscher , Christoph Wasser , Urs Schneider

Vocal tract configurations play a vital role in generating distinguishable speech sounds, by modulating the airflow and creating different resonant cavities in speech production. They contain abundant information that can be utilized to…

Sound · Computer Science 2018-07-31 Pramit Saha , Praneeth Srungarapu , Sidney Fels

A crucial part of an accurate and reliable spoken language assessment system is the underlying ASR model. Recently, large-scale pre-trained ASR foundation models such as Whisper have been made available. As the output of these models is…

Computation and Language · Computer Science 2023-10-11 Rao Ma , Mengjie Qian , Mark J. F. Gales , Kate M. Knill

Recent advances in unsupervised speech representation learning discover new approaches and provide new state-of-the-art for diverse types of speech processing tasks. This paper presents an investigation of using wav2vec 2.0 deep speech…

Deepfake speech represents a real and growing threat to systems and society. Many detectors have been created to aid in defense against speech deepfakes. While these detectors implement myriad methodologies, many rely on low-level fragments…

In most automatic speech recognition (ASR) systems, the audio signal is processed to produce a time series of sensor measurements (e.g., filterbank outputs). This time series encodes semantic information in a speaker-dependent way. An…

Sound · Computer Science 2019-05-10 David N. Levin

Respiratory airflow signals provide critical insight into breathing mechanics, yet conventional analysis methods remain limited in their ability to characterize the internal structure of individual breaths. Traditional approaches treat…

Signal Processing · Electrical Eng. & Systems 2026-04-27 Victoria Ribeiro Rodrigues , Paul W. Davenport , Nicholas J. Napoli

Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emotional categories, or…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-13 Anna Derington , Hagen Wierstorf , Ali Özkil , Florian Eyben , Felix Burkhardt , Björn W. Schuller

This paper presents a robust deep learning framework developed to detect respiratory diseases from recordings of respiratory sounds. The complete detection process firstly involves front end feature extraction where recordings are…

Sound · Computer Science 2020-02-11 Lam Pham , Ian McLoughlin , Huy Phan , Minh Tran , Truc Nguyen , Ramaswamy Palaniappan

A precise spatial delivery of the radiation dose is crucial for the treatment success in radiotherapy. In the lung and upper abdominal region, respiratory motion introduces significant treatment uncertainties, requiring special motion…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Jan Boysen , Hristina Uzunova , Heinz Handels , Jan Ehrhardt

Measuring the amount of speech production in daily life is important for understanding communication in organizations and identifying mental disorders. However, measuring the amount of speech production can be problematic in terms of…

Human-Computer Interaction · Computer Science 2024-05-13 Rintaro Katagiri , Yutaka Arakawa , Yugo Nakamura

Continuous monitoring of respiratory activity is desirable in many clinical applications to detect respiratory events. Non-contact monitoring of respiration can be achieved with near- and far-infrared spectrum cameras. However, current…

Signal Processing · Electrical Eng. & Systems 2020-06-02 Gaetano Scebba , Giulia Da Poian , Walter Karlen

The purpose of this study is to provide means to physicians for automated and fast recognition of airways diseases. In this work, we mainly focus on measures that can be easily recorded using a spirometer. The signals used in this framework…

Machine Learning · Computer Science 2021-11-09 Riccardo Dio , André Galligo , Angelos Mantzaflaris , Benjamin Mauroy

Speech Emotion Recognition (SER) is the use of machines to detect the emotional state of humans based on the speech, which is gaining importance in natural human-computer interaction. Speech is a very valuable source of information, as…

This paper explores speculative speech recognition (SSR), where we empower conventional automatic speech recognition (ASR) with speculation capabilities, allowing the recognizer to run ahead of audio. We introduce a metric for measuring SSR…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-08 Bolaji Yusuf , Murali Karthick Baskar , Andrew Rosenberg , Bhuvana Ramabhadran

Recent research has highlighted the detection of human respiration rate using commodity WiFi devices. Nevertheless, these devices encounter challenges in accurately discerning human respiration amidst the prevailing human motion…

Signal Processing · Electrical Eng. & Systems 2024-01-08 Kehan Wu , Renqi Chen , Haiyu Wang , Chenqing Ji , Jiayuan Zhu , Guang Wu

Respiratory ailments afflict a wide range of people and manifests itself through conditions like asthma and sleep apnea. Continuous monitoring of chronic respiratory ailments is seldom used outside the intensive care ward due to the large…

Recent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-grained variational autoencoder (VAE) structure, extracting…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-11 Guangzhi Sun , Yu Zhang , Ron J. Weiss , Yuan Cao , Heiga Zen , Andrew Rosenberg , Bhuvana Ramabhadran , Yonghui Wu

This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in 128…