English
Related papers

Related papers: Failure-Centered Runtime Evaluation for Deployed T…

200 papers

Automatic speech recognition (ASR) outcomes serve as input for downstream tasks, substantially impacting the satisfaction level of end-users. Hence, the diagnosis and enhancement of the vulnerabilities present in the ASR model bear…

Computation and Language · Computer Science 2024-01-29 Seonmin Koo , Chanjun Park , Jinsung Kim , Jaehyung Seo , Sugyeong Eo , Hyeonseok Moon , Heuiseok Lim

Risk decision systems in fraud detection and credit scoring operate under structural label absence: ground truth arrives weeks to months after decisions are made. During this blind period, model performance may degrade silently, eroding the…

Computers and Society · Computer Science 2026-04-21 Oleg Solozobov

Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers have committed to in the market. We introduce StakeBench, an evaluation framework for…

Computation and Language · Computer Science 2026-05-26 Yunhua Pei , Jingyu Hu , Yiwei Shi , Hongnan Ma , Weiru Liu , John Cartlidge

Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience across task boundaries. This paper formalizes the…

Artificial Intelligence · Computer Science 2026-05-26 Sihang Jiang , Lipeng Ma , Zhonghua Hong , Keyi Wang , Zhiyu Lu , Tengfei Wang , Shisong Chen , Jinghao Zhang , Tianjun Pan , Weijia Li , Jiaqing Liang , Yanghua Xiao

Large language models are increasingly deployed in multi-agent systems to overcome context limitations by distributing information across agents. Yet whether agents can reliably compute with distributed information, rather than merely…

Multiagent Systems · Computer Science 2026-04-15 Yuzhe Zhang , Feiran Liu , Yi Shan , Xinyi Huang , Xin Yang , Yueqi Zhu , Xuxin Cheng , Cao Liu , Ke Zeng , Terry Jingchen Zhang , Wenyuan Jiang

Standard benchmarks fixate on how well large language model (LLM) agents perform in finance, yet say little about whether they are safe to deploy. We argue that accuracy metrics and return-based scores provide an illusion of reliability,…

General Finance · Quantitative Finance 2025-06-03 Zichen Chen , Jiaao Chen , Jianda Chen , Misha Sra

Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervision for execution drift, while failed rollouts are often discarded. We introduce RePO-VLA,…

Language Identification (LID) is a core task in multilingual NLP, yet current systems often overfit to clean, monolingual data. This work introduces DIVERS-BENCH, a comprehensive evaluation of state-of-the-art LID models across diverse…

Computation and Language · Computer Science 2025-09-23 Jessica Ojo , Zina Kamel , David Ifeoluwa Adelani

Static benchmarks for LLMs are increasingly compromised by contamination and overfitting especially on knowledge intensive reasoning tasks While recent dynamic benchmarks can alleviate staleness they often increase difficulty at the expense…

Computation and Language · Computer Science 2026-05-05 Yongrui Chen , Yangyang Ma , Xiaoying Huang , Shenyu Zhang , Huajun Chen , Haofen Wang , Guilin Qi

Deep learning methods have shown promising performance in fault diagnosis for multimode process. Most existing studies assume that the collected health state categories from different operating modes are identical. However, in real…

Machine Learning · Computer Science 2025-10-30 Guangqiang Li , M. Amine Atoui , Xiangshun Li

Automatic Speech Recognition (ASR) in professional settings faces challenges that existing benchmarks underplay: dense domain terminology, formal register variation, and near-zero tolerance for critical entity errors. We present…

Computation and Language · Computer Science 2025-12-30 Deepak Babu Piskala

Context: Today's safety critical systems are increasingly reliant on software. Software becomes responsible for most of the critical functions of systems. Many different safety analysis techniques have been developed to identify hazards of…

Software Engineering · Computer Science 2016-12-02 Asim Abdulkhaleq , Stefan Wagner

Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational…

Computation and Language · Computer Science 2024-09-19 Gaurav Maheshwari , Dmitry Ivanov , Théo Johannet , Kevin El Haddad

This paper presents the design and refinement of automated Moodle-based Problem-Solving Assessments (PSAs) deployed across large-scale computing units. Developed to replace traditional exams, PSAs assess applied problem-solving skills…

Computers and Society · Computer Science 2025-08-26 Charith Jayasekara , Carlo Kopp , Vincent Lee , Chetan Arora

Developers need to perform adequate testing to ensure the quality of Automatic Speech Recognition (ASR) systems. However, manually collecting required test cases is tedious and time-consuming. Our recent work proposes CrossASR, a…

Software Engineering · Computer Science 2022-01-06 Muhammad Hilmi Asyrofi , Zhou Yang , David Lo

In this paper we describe the top-scoring IDLab submission for the text-independent task of the Short-duration Speaker Verification (SdSV) Challenge 2020. The main difficulty of the challenge exists in the large degree of varying phonetic…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Jenthe Thienpondt , Brecht Desplanques , Kris Demuynck

When multiple people share a single voice assistant, the system conflates their histories: one resident's preferences can leak into another's responses, eroding utility and trust. We call this failure mode persona confusion, and we show it…

Human-Computer Interaction · Computer Science 2026-04-29 Mohammad Al-Ratrout , Pavan Uttej Ravva , Shayla Sharmin , Aditya Raikwar , Ju Young Shin , Roghayeh Leila Barmaki

Autonomous language-model agents increasingly rely on installable skills and tools to complete user tasks. Static skill auditing can expose capability surface before deployment, but it cannot determine whether a particular invocation is…

Artificial Intelligence · Computer Science 2026-04-14 Guijia Zhang , Shu Yang , Xilin Gong , Di Wang

Federated and Continual Learning have emerged as potential paradigms for the robust and privacy-aware use of Deep Learning in dynamic environments. However, Client Drift and Catastrophic Forgetting are fundamental obstacles to guaranteeing…

Machine Learning · Computer Science 2023-09-06 Niklas Babendererde , Moritz Fuchs , Camila Gonzalez , Yuri Tolkach , Anirban Mukhopadhyay

Sign language translation has historically been peripheral to mainstream machine translation research. In order to help converge the fields, we introduce FLEURS-ASL, an extension of the multiway parallel benchmarks FLORES (for text) and…

Computation and Language · Computer Science 2024-08-27 Garrett Tanzer