English
Related papers

Related papers: Comparison of remote experiments using crowdsourci…

200 papers

The Lip Reading Sentences-3 (LRS3) benchmark has primarily been the focus of intense research in visual speech recognition (VSR) during the last few years. As a result, there is an increased risk of overfitting to its excessively used test…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Yasser Abdelaziz Dahou Djilali , Sanath Narayan , Eustache Le Bihan , Haithem Boussaid , Ebtessam Almazrouei , Merouane Debbah

Speech communication systems based on Voice-over-IP technology are frequently used by native as well as non-native speakers of a target language, e.g. in international phone calls or telemeetings. Frequently, such calls also occur in a…

Multimedia · Computer Science 2020-10-27 Babak Naderi , Gabriel Mittag , Rafael Zequeira Jim\a'enez , Sebastian Möller

Machine Translation (MT) has achieved remarkable performance, with growing interest in speech translation and multimodal approaches. However, despite these advancements, MT quality assessment remains largely text centric, typically relying…

Computation and Language · Computer Science 2025-09-18 Sami Ul Haq , Sheila Castilho , Yvette Graham

We investigate the opportunities and challenges of running virtual reality (VR) studies remotely. Today, many consumers own head-mounted displays (HMDs), allowing them to participate in scientific studies from their homes using their own…

The field of prosody transfer in speech synthesis systems is rapidly advancing. This research is focused on evaluating learning methods for adapting pre-trained monolingual text-to-speech (TTS) models to multilingual conditions, i.e.,…

Computation and Language · Computer Science 2024-06-19 Arnav Goel , Medha Hira , Anubha Gupta

Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech…

Sound · Computer Science 2025-06-23 Shoutrik Das , Nishant Singh , Arjun Gangwar , S Umesh

The wisdom of crowds has been shown to operate not only for factual judgments but also in matters of taste, where accuracy is defined relative to an individual's preferences. However, it remains unclear how different types of social signals…

Physics and Society · Physics 2026-02-12 Itsuki Fujisaki , Kunhao Yang

When evaluating the effectiveness of a drug, a Randomized Controlled Trial (RCT) is often considered the gold standard due to its perfect randomization. While RCT assures strong internal validity, its restricted external validity poses…

Applications · Statistics 2024-06-07 Kuan Jiang , Xin-xing Lai , Shu Yang , Ying Gao , Xiao-Hua Zhou

This paper introduces a unified framework for the detection of a source with a sensor array in the context where the noise variance and the channel between the source and the sensors are unknown at the receiver. The Generalized Maximum…

Probability · Mathematics 2010-06-16 Pascal Bianchi , Merouane Debbah , Mylène Maïda , Jamal Najim

Randomized Controlled Trials (RCTs) are pivotal in generating internally valid estimates with minimal assumptions, serving as a cornerstone for researchers dedicated to advancing causal inference methods. However, extending these findings…

Methodology · Statistics 2024-05-28 Melody Y Huang , Harsh Parikh

An accurate objective speech intelligibility prediction algorithms is of great interest for many applications such as speech enhancement for hearing aids. Most algorithms measures the signal-to-noise ratios or correlations between the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-07 Zehai Tu , Ning Ma , Jon Barker

Social robots are required not only to understand human intentions but also to effectively communicate their intentions or own internal states to users. This study explores the use of sonification to provide explicit auditory feedback,…

Robotics · Computer Science 2024-11-15 Simone Arreghini , Antonio Paolillo , Gabriele Abbate , Alessandro Giusti

Reverberation, especially in large rooms, severely degrades speech recognition performance and speech intelligibility. Since direct measurement of room characteristics is usually not possible, blind estimation of reverberation-related…

Sound · Computer Science 2015-10-19 M. Senoussaoui , J. F. Santos , T. H. Falk

The problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear…

Computation and Language · Computer Science 2023-10-11 Piyush Singh Pasi , Karthikeya Battepati , Preethi Jyothi , Ganesh Ramakrishnan , Tanmay Mahapatra , Manoj Singh

Security especially in the fields of IoT, industrial automation and critical infrastructure is paramount nowadays and a hot research topic. In order to ensure confidence in research results they need to be reproducible. In the past we…

Hardware Architecture · Computer Science 2024-07-10 Dmytro Petryk , Ievgen Kabin , Peter Langendörfer , Zoya Dyka

Software-defined radios (SDRs) are often used in the experimental evaluation of next-generation wireless technologies. While crowd-sourced spectrum monitoring is an important component of future spectrum-agile technologies, there is no…

Networking and Internet Architecture · Computer Science 2019-05-31 Phillip Smith , Anh Luong , Shamik Sarkar , Harsimran Singh , Neal Patwari , Sneha Kasera , Kurt Derr , Samuel Ramirez

A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Remote Sensing MLLMs fail to meet this requirement. Trained on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Rui Min , Liang Yao , Shiyu Miao , Shengxiang Xu , Yuxuan Liu , Chuanyi Zhang , Shimin Di , Fan Liu

Speech recognition (ASR) and speaker diarization (SD) models have traditionally been trained separately to produce rich conversation transcripts with speaker labels. Recent advances have shown that joint ASR and SD models can learn to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-06 Huanru Henry Mao , Shuyang Li , Julian McAuley , Garrison Cottrell

Traditional ASR metrics like WER and CER fail to capture intelligibility, especially for dysarthric and dysphonic speech, where semantic alignment matters more than exact word matches. ASR systems struggle with these speech types, often…

Machine Learning · Computer Science 2025-12-12 Bornali Phukon , Xiuwen Zheng , Mark Hasegawa-Johnson

Self-supervised speech representations (SSSRs) have been successfully applied to a number of speech-processing tasks, e.g. as feature extractor for speech quality (SQ) prediction, which is, in turn, relevant for assessment and training…

Sound · Computer Science 2023-12-08 George Close , Thomas Hain , Stefan Goetze
‹ Prev 1 8 9 10 Next ›