English
Related papers

Related papers: Voice Evaluation of Reasoning Ability: Diagnosing …

200 papers

While Audio Large Models (ALMs) have achieved remarkable proficiency, their robustness remains brittle in real-world deployment. Existing evaluations largely rely on synthetic Gaussian noise or simplistic single-source interference, failing…

While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts or focus exclusively on human speech. This leaves a significant gap in real-time, sound-centric…

Robotics · Computer Science 2026-03-18 Chang Nie , Tianchen Deng , Guangming Wang , Zhe Liu , Hesheng Wang

Large Language Models (LLMs) have achieved remarkable performance on single-turn tasks, yet their effectiveness deteriorates in multi-turn conversations. We define this phenomenon as cumulative contextual decay - a progressive degradation…

Computation and Language · Computer Science 2025-12-09 Wanyang Hong , Zhaoning Zhang , Yi Chen , Libo Zhang , Baihui Liu , Linbo Qiao , Zhiliang Tian , Dongsheng Li

Reasoning over dynamic visual content remains a central challenge for multimodal large language models. Recent thinking models generate explicit reasoning traces for interpretability; however, their reasoning often appears convincing while…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Muhammad Maaz , Hanoona Rasheed , Fahad Shahbaz Khan , Salman Khan

Retrieval-Augmented Generation (RAG) enhances recency and factuality in answers. However, existing evaluations rarely test how well these systems cope with real-world noise, conflicting between internal and external retrieved contexts, or…

Computation and Language · Computer Science 2025-10-29 Yixiao Zeng , Tianyu Cao , Danqing Wang , Xinran Zhao , Zimeng Qiu , Morteza Ziyadi , Tongshuang Wu , Lei Li

Hallucination--defined here as generating statements unsupported or contradicted by available evidence or conversational context--remains a major obstacle to deploying conversational AI systems in settings that demand factual reliability.…

Computation and Language · Computer Science 2026-04-22 Ashley Lewis , Andrew Perrault , Eric Fosler-Lussier , Michael White

Effectively steering hearable devices requires understanding the acoustic environment around the user. In the computational analysis of sound scenes, foundation models have emerged as the state of the art to produce high-performance,…

Audio Question Answering (AQA) is a key task for evaluating Audio-Language Models (ALMs), yet assessing open-ended responses remains challenging. Existing metrics used for AQA such as BLEU, METEOR and BERTScore, mostly adapted from NLP and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Satvik Dixit , Soham Deshmukh , Bhiksha Raj

Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only final outcomes, which can lead to a failure mode where task…

Computation and Language · Computer Science 2026-05-14 Kyuyoung Kim , Kevin Wang , Yunfei Xie , Peiyang Xu , Peiyao Sheng , Chen Wei , Zhangyang Wang , Jinwoo Shin , Pramod Viswanath , Sewoong Oh

Explainability in automated student answer scoring systems is critical for building trust and enhancing usability among educators. Yet, generating high-quality assessment rationales remains challenging due to the scarcity of annotated data…

Computation and Language · Computer Science 2025-09-30 Jiazheng Li , Artem Bobrov , Runcong Zhao , Cesare Aloisi , Yulan He

Retrieval-Augmented Generation (RAG) systems typically face constraints because of their inherent mechanism: a simple top-k semantic search [1]. The approach often leads to the incorporation of irrelevant or redundant information in the…

Computation and Language · Computer Science 2025-09-03 Andreas Ottem

Recent advancements in multimodal large language models have driven breakthroughs in visual question answering. Yet, a critical gap persists, `conceptualization'-the ability to recognize and reason about the same concept despite variations…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Zahra Babaiee , Peyman M. Kiasari , Daniela Rus , Radu Grosu

Reasoning capabilities are crucial for reliable medical visual question answering (VQA); however, existing datasets rarely include reasoning explanations. We address this by generating reasoning trajectories for six medical VQA benchmarks…

Machine Learning · Computer Science 2026-05-07 Halil Ibrahim Gulluk , Olivier Gevaert

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce…

Computation and Language · Computer Science 2025-09-29 Ke Wang , Houxing Ren , Zimu Lu , Mingjie Zhan , Hongsheng Li

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Sirnam Swetha , Rohit Gupta , Parth Parag Kulkarni , David G Shatwell , Jeffrey A Chan Santiago , Nyle Siddiqui , Joseph Fioresi , Mubarak Shah

Existing audio question answering benchmarks largely emphasize sound event classification or caption-grounded queries, often enabling models to succeed through shortcut strategies, short-duration cues, lexical priors, dataset-specific…

Computation and Language · Computer Science 2026-04-24 Tasnim Kabir , Dmytro Kurdydyk , Aadi Palnitkar , Liam Dorn , Ahmed Haj Ahmed , Jordan Lee Boyd-Graber

Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks…

Machine Learning · Computer Science 2025-10-14 Christopher Chiu , Silviu Pitis , Mihaela van der Schaar

Deep neural network approaches to speaker verification have proven successful, but typical computational requirements of State-Of-The-Art (SOTA) systems make them unsuited for embedded applications. In this work, we present a two-stage…

Sound · Computer Science 2021-04-22 Julien Balian , Raffaele Tavarone , Mathieu Poumeyrol , Alice Coucke

Despite decades of research on reverberant speech, comparing methods remains difficult because most corpora lack per-file acoustic annotations or provide limited documentation for reproduction. We present RIR-Mega-Speech, a corpus of…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Mandip Goswami

Visual-Language-Action (VLA) models report impressive success rates on robotic manipulation benchmarks, yet these results may mask fundamental weaknesses in robustness. We perform a systematic vulnerability analysis by introducing…