English
Related papers

Related papers: Comparing acoustic analyses of speech data collect…

200 papers

This work unveils the enigmatic link between phonemes and facial features. Traditional studies on voice-face correlations typically involve using a long period of voice input, including generating face images from voices and reconstructing…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Liao Qu , Xianwei Zou , Xiang Li , Yandong Wen , Rita Singh , Bhiksha Raj

A microphone array can provide a mobile robot with the capability of localizing, tracking and separating distant sound sources in 2D, i.e., estimating their relative elevation and azimuth. To combine acoustic data with visual information in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-23 Simon Michaud , Samuel Faucher , François Grondin , Jean-Samuel Lauzon , Mathieu Labbé , Dominic Létourneau , François Ferland , François Michaud

Outbound AI calling systems must distinguish voicemail greetings from live human answers in real time to avoid wasted agent interactions and dropped calls. We present a lightweight approach that extracts 15 temporal features from the speech…

Sound · Computer Science 2026-04-14 Kumar Saurav

The Grapheme-to-Phoneme (G2P) task aims to convert orthographic input into a discrete phonetic representation. G2P conversion is beneficial to various speech processing applications, such as text-to-speech and speech recognition. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-01 Manuel Sam Ribeiro , Giulia Comini , Jaime Lorenzo-Trueba

Contact tracing has been globally adopted in the fight to control the infection rate of COVID-19. Thanks to digital technologies, such as smartphones and wearable devices, contacts of COVID-19 patients can be easily traced and informed…

Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-source models against a very challenging real-world medical form that mixes dates;…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Nicholas Pather , Joshua Fouché , Sitwala Mundia , Karl-Günter Technau , Thokozile Malaba , Alex Welte , Ushma Mehta , Bruce A. Bassett

Over the past year, remote speech intelligibility testing has become a popular and necessary alternative to traditional in-person experiments due to the need for physical distancing during the COVID-19 pandemic. A remote framework was…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-01 Kevin M. Chu , Leslie M. Collins , Boyla O. Mainsah

We describe the speech activity detection (SAD), speaker diarization (SD), and automatic speech recognition (ASR) experiments conducted by the Behavox team for the Interspeech 2020 Fearless Steps Challenge (FSC-2). A relatively small amount…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-05 Arseniy Gorin , Daniil Kulko , Steven Grima , Alex Glasman

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Qiang Huang , Thomas Hain

In this work, we propose a speaker anonymization pipeline that leverages high quality automatic speech recognition and synthesis systems to generate speech conditioned on phonetic transcriptions and anonymized speaker embeddings. Using…

Sound · Computer Science 2022-07-12 Sarina Meyer , Florian Lux , Pavel Denisov , Julia Koch , Pascal Tilli , Ngoc Thang Vu

In mobile speech communication applications, wind noise can lead to a severe reduction of speech quality and intelligibility. Since the performance of speech enhancement algorithms using acoustic microphones tends to substantially degrade…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-15 Marvin Tammen , Xilin Li , Simon Doclo , Lalin Theverapperuma

The speech signal is a consummate example of time-series data. The acoustics of the signal change over time, sometimes dramatically. Yet, the most common type of comparison we perform in phonetics is between instantaneous acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-18 Matthew C. Kelley

Only a handful of the world's languages are abundant with the resources that enable practical applications of speech processing technologies. One of the methods to overcome this problem is to use the resources existing in other languages to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Piotr Żelasko , Laureano Moro-Velázquez , Mark Hasegawa-Johnson , Odette Scharenborg , Najim Dehak

We present LibriWASN, a data set whose design follows closely the LibriCSS meeting recognition data set, with the marked difference that the data is recorded with devices that are randomly positioned on a meeting table and whose sampling…

Sound · Computer Science 2023-08-22 Joerg Schmalenstroeer , Tobias Gburrek , Reinhold Haeb-Umbach

Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, imprecise…

We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition…

Speaker anonymization systems continue to improve their ability to obfuscate the original speaker characteristics in a speech signal, but often create processing artifacts and unnatural sounding voices as a tradeoff. Many of those systems…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-23 Ünal Ege Gaznepoglu , Nils Peters

We report a user study of over four months on the non-voice usage of mobile phones by teens from an underserved urban community in the USA where a community-wide, open-access Wi-Fi network exists. We instrumented the phones to record…

Human-Computer Interaction · Computer Science 2010-12-14 Ahmad Rahmati , Lin Zhong

Objective: Distorted loudness perception is one of the main complaints of hearing aid users. Being able to measure loudness perception correctly in the clinic is essential for fitting hearing aids. For this, experiments in the clinic should…

Neurons and Cognition · Quantitative Biology 2022-05-05 Gerard Llorach , Dirk Oetting , Matthias Vormann , Markus Meis , Volker Hohmann

The lockdowns and travel restrictions in current coronavirus pandemic situation has replaced face-to-face teaching and meeting with online teaching and meeting. Recently, the video conferencing tool Zoom has become extremely popular for its…

Cryptography and Security · Computer Science 2020-05-22 Manoranjan Mohanty , Waheeb Yaqub
‹ Prev 1 3 4 5 6 7 10 Next ›