English
Related papers

Related papers: Moravec's Paradox: Towards an Auditory Turing Test

200 papers

World models have demonstrated impressive performance on robotic learning tasks. Many such tasks inherently demand multimodal reasoning; for example, filling a bottle with water will lead to visual information alone being ambiguous or…

Robotics · Computer Science 2025-12-10 Fan Zhang , Michael Gienger

Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error types. Challenge sets are used to probe specific dimensions of…

Computation and Language · Computer Science 2024-01-30 Nikita Moghe , Arnisa Fazla , Chantal Amrhein , Tom Kocmi , Mark Steedman , Alexandra Birch , Rico Sennrich , Liane Guillou

Audio deepfakes are increasingly in-differentiable from organic speech, often fooling both authentication systems and human listeners. While many techniques use low-level audio features or optimization black-box model training, focusing on…

Sound · Computer Science 2025-02-21 Kevin Warren , Daniel Olszewski , Seth Layton , Kevin Butler , Carrie Gates , Patrick Traynor

The Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC 2024) is part of the ISCSLP 2024 Competitions and Challenges track. While current text-to-speech (TTS) technology can generate high-quality audio, its ability to convey…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-01 Ruibo Fu , Rui Liu , Chunyu Qiang , Yingming Gao , Yi Lu , Shuchen Shi , Tao Wang , Ya Li , Zhengqi Wen , Chen Zhang , Hui Bu , Yukun Liu , Xin Qi , Guanjun Li

Intelligent tutoring systems increasingly provide automated feedback on student work, but robust feedback requires assessing reasoning, not only final answers. We study a failure mode we call the correct answer trap (CAT): models…

Computers and Society · Computer Science 2026-05-26 Moiz Imran , Sahan Bulathwela

Recent advancements in large audio-language models (LALMs) have shown impressive capabilities in understanding and reasoning about audio and speech information. However, these models still face challenges, including hallucinating…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Chun-Yi Kuan , Hung-yi Lee

As dialogue systems and chatbots increasingly integrate into everyday interactions, the need for efficient and accurate evaluation methods becomes paramount. This study explores the comparative performance of human and AI assessments across…

Computation and Language · Computer Science 2024-09-11 Ike Ebubechukwu , Johane Takeuchi , Antonello Ceravola , Frank Joublin

The Automated Audio Captioning (AAC) task aims to describe an audio signal using natural language. To evaluate machine-generated captions, the metrics should take into account audio events, acoustic scenes, paralinguistics, signal…

Sound · Computer Science 2024-11-06 Satvik Dixit , Soham Deshmukh , Bhiksha Raj

We present Voice Evaluation of Reasoning Ability (VERA), a benchmark for evaluating reasoning ability in voice-interactive systems under real-time conversational constraints. VERA comprises 2,931 voice-native episodes derived from…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-01 Yueqian Lin , Zhengmian Hu , Qinsi Wang , Yudong Liu , Hengfan Zhang , Jayakumar Subramanian , Nikos Vlassis , Hai Helen Li , Yiran Chen

Speech technologies are deployed in high-stakes settings, yet fairness concerns remain fragmented across tasks and disciplines. Existing surveys either adopt a general machine-learning perspective that overlooks speech-specific properties…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-05 Yi-Cheng Lin , Yun-Shao Tsai , Kuan-Yu Chen , Hsiao-Ying Huang , Huang-Cheng Chou , Hung-yi Lee

Our brain learns to update its mental model of the environment by abstracting sensory experiences for adaptation and survival. Learning to categorize sounds is one essential abstracting process for high-level human cognition, such as speech…

Neurons and Cognition · Quantitative Biology 2025-10-22 Nan Wang , Gangyi Feng

People solve different problems and know that some of them are simple, some are complex and some insoluble. The main goal of this work is to develop a mathematical theory of algorithmic complexity for problems. This theory is aimed at…

Computational Complexity · Computer Science 2008-07-08 Mark Burgin

Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in…

Sound · Computer Science 2025-04-30 Zhicheng Lian , Lizhi Wang , Hua Huang

In this work, we demonstrate the existence of universal adversarial audio perturbations that cause mis-transcription of audio signals by automatic speech recognition (ASR) systems. We propose an algorithm to find a single…

Machine Learning · Computer Science 2019-08-16 Paarth Neekhara , Shehzeen Hussain , Prakhar Pandey , Shlomo Dubnov , Julian McAuley , Farinaz Koushanfar

Our prior experiments show that humans and machines seem to employ different approaches to speaker discrimination, especially in the presence of speaking style variability. The experiments examined read versus conversational speech.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 Amber Afshan , Abeer Alwan

Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate acoustic information, rather than verifying the audio…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xiaofei Wen , Wenjie Jacky Mo , Xingyu Fu , Rui Cai , Tinghui Zhu , Wendi Li , Yanan Xie , Muhao Chen , Peng Qi

Modern voice cloning, also known as zero-shot text-to-speech (TTS), can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing.…

Sound · Computer Science 2026-05-26 Ruinan Jin , Xinting Liao , Hanlin Yu , Deval Pandya , Xiaoxiao Li

This paper presents, a first of its kind, audio-visual (AV) speech enhacement challenge in real-noisy settings. A detailed description of the AV challenge, a novel real noisy AV corpus (ASPIRE), benchmark speech enhancement task, and…

Sound · Computer Science 2019-10-02 Mandar Gogate , Ahsan Adeel , Kia Dashtipour , Peter Derleth , Amir Hussain

Automatic Music Transcription (AMT) -- the task of converting music audio into note representations -- has seen rapid progress, driven largely by deep learning systems. Due to the limited availability of richly annotated music datasets,…

Sound · Computer Science 2026-01-27 Lukáš Samuel Marták , Patricia Hu , Gerhard Widmer

Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and…

‹ Prev 1 4 5 6 7 8 10 Next ›