English
Related papers

Related papers: ABRA: Agent Benchmark for Radiology Applications

200 papers

Crucial in disease analysis and surgical planning, manual segmentation of volumetric medical scans (e.g. MRI, CT) is laborious, error-prone, and challenging to master, while fully automatic algorithms can benefit from user feedback.…

Human-Computer Interaction · Computer Science 2025-05-27 Pascal Spiegler , Arash Harirpoush , Yiming Xiao

Causal analysis on relational databases is challenging, as analysis datasets must be repeatedly queried from complex schemas. Recent LLM systems can automate individual steps, but they hardly manage dependencies across analysis stages,…

Databases · Computer Science 2026-03-19 Joanie Hayoun Chung , Sumin Lee , Sungbin Lim

VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings. This mismatch creates a substantial reality gap, where…

Chest X-rays (CXRs) play an integral role in driving critical decisions in disease management and patient care. While recent innovations have led to specialized models for various CXR interpretation tasks, these solutions often operate in…

Machine Learning · Computer Science 2025-05-30 Adibvafa Fallahpour , Jun Ma , Alif Munim , Hongwei Lyu , Bo Wang

Deep Research Agents (DRAs) can autonomously conduct complex investigations and generate comprehensive reports, demonstrating strong real-world potential. However, existing evaluations mostly rely on close-ended benchmarks, while open-ended…

As AI agents proliferate across industries and applications, evaluating their performance based solely on infrastructural metrics such as latency, time-to-first-token, or token throughput is proving insufficient. These metrics fail to…

Artificial Intelligence · Computer Science 2025-11-12 Waseem AlShikh , Muayad Sayed Ali , Brian Kennedy , Dmytro Mozolevskyi

Large Language Models (LLMs) are increasingly used as autonomous agents in complex, long-horizon applications, where effective memory is critical for sustained performance. Yet existing memory benchmarks are largely dialogue-centric, while…

Environments built for people are increasingly operated by a new class of economic actors: LLM-powered software agents making decisions on our behalf. These decisions range from our purchases to travel plans to medical treatment selection.…

Artificial Intelligence · Computer Science 2026-02-25 Manuel Cherep , Chengtian Ma , Abigail Xu , Maya Shaked , Pattie Maes , Nikhil Singh

In this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. These tasks require agents to end-to-end solving complex tasks by interacting with an execution…

Computation and Language · Computer Science 2024-03-12 Xueyu Hu , Ziyu Zhao , Shuang Wei , Ziwei Chai , Qianli Ma , Guoyin Wang , Xuwu Wang , Jing Su , Jingjing Xu , Ming Zhu , Yao Cheng , Jianbo Yuan , Jiwei Li , Kun Kuang , Yang Yang , Hongxia Yang , Fei Wu

Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition. Real clinical workflows impose a different burden: readers must search through complete…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Weixiang Shen , Chengzhi Shen , Yanzhu Hu , Che Liu , Junde Wu , Jiayuan Zhu , Xiao Han , Zongyue Li , Jingpei Wu , Min Xu , Daguang Xu , Yueming Jin , Benedikt Wiestler , Daniel Rueckert , Jiazhen Pan

Retinal anomaly detection plays a pivotal role in screening ocular and systemic diseases. Despite its significance, progress in the field has been hindered by the absence of a comprehensive and publicly available benchmark, which is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Chenyu Lian , Hong-Yu Zhou , Zhanli Hu , Jing Qin

Evaluating the clinical correctness and reasoning fidelity of automatically generated medical imaging reports remains a critical yet unresolved challenge. Existing evaluation methods often fail to capture the structured diagnostic logic…

Artificial Intelligence · Computer Science 2026-01-26 Suzhong Fu , Jingqi Dong , Xuan Ding , Rui Sun , Yiming Yang , Shuguang Cui , Zhen Li

The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Existing benchmarks measure agent capabilities through retrospectively curated tasks with…

While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively understand their limits and failure modes. To address this, we introduce a comprehensive benchmark…

Robotics · Computer Science 2025-12-30 Borong Zhang , Jiahao Li , Jiachen Shen , Yishuai Cai , Yuhao Zhang , Yuanpei Chen , Juntao Dai , Jiaming Ji , Yaodong Yang

We propose a new model-based computer-aided diagnosis (CAD) system for tumor detection and classification (cancerous v.s. benign) in breast images. Specifically, we show that (x-ray, ultrasound and MRI) images can be accurately modeled by…

Artificial Intelligence · Computer Science 2009-06-22 Nidhal Bouaynaya , Jerzy Zielinski , Dan Schonfeld

Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks…

Artificial Intelligence · Computer Science 2026-01-15 Ananya Mantravadi , Shivali Dalmia , Abhishek Mukherji

Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external writes, ignored errors, weakly grounded commitments, or…

Artificial Intelligence · Computer Science 2026-05-29 Yibing Liu , Yangze Liu , Xiaolong Yin , Bin Wang , Chong Zhang , Hao Yin , Zhongyi Han

Quantifying the degree of atrophy is done clinically by neuroradiologists following established visual rating scales. For these assessments to be reliable the rater requires substantial training and experience, and even then the rating…

Medical image classification is a core task in computer-aided diagnosis (CAD), playing a pivotal role in early disease detection, treatment planning, and patient prognosis assessment. In ophthalmic practice, fluorescein fundus angiography…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Zhuonan Wang , Wenjie Yan , Wenqiao Zhang , Xiaohui Song , Jian Ma , Ke Yao , Yibo Yu , Beng Chin Ooi

As reinforcement learning continues to scale the training of large language model-based agents, reliably verifying agent behaviors in complex environments has become increasingly challenging. Existing approaches rely on rule-based verifiers…

Artificial Intelligence · Computer Science 2026-04-21 Wentao Shi , Yu Wang , Yuyang Zhao , Yuxin Chen , Fuli Feng , Xueyuan Hao , Xi Su , Qi Gu , Hui Su , Xunliang Cai , Xiangnan He
‹ Prev 1 3 4 5 6 7 10 Next ›