English
Related papers

Related papers: MedOpenClaw and MedFlowBench: Auditing Medical Age…

200 papers

We introduce MultiMedEval, an open-source toolkit for fair and reproducible evaluation of large, medical vision-language models (VLM). MultiMedEval comprehensively assesses the models' performance on a broad array of six multi-modal tasks,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Corentin Royer , Bjoern Menze , Anjany Sekuboyina

Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evaluating them on live services is risky due to potentially irreversible changes. Existing…

Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external writes, ignored errors, weakly grounded commitments, or…

Artificial Intelligence · Computer Science 2026-05-29 Yibing Liu , Yangze Liu , Xiaolong Yin , Bin Wang , Chong Zhang , Hao Yin , Zhongyi Han

The integration of Artificial Intelligence (AI), especially Large Language Models (LLMs), into the clinical diagnosis process offers significant potential to improve the efficiency and accessibility of medical care. While LLMs have shown…

Computation and Language · Computer Science 2024-10-15 Mingyu Derek Ma , Chenchen Ye , Yu Yan , Xiaoxuan Wang , Peipei Ping , Timothy S Chang , Wei Wang

Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through specialist roles, peer review and consensus formation. In clinical decision support, however,…

Computation and Language · Computer Science 2026-05-28 Yinghao Zhu , Lei Gu , Zixiang Wang , Haoran Sang , Dehao Sui , Wen Tang , Lan Mi , Yasha Wang , Junyi Gao , Liang Yao , Tianfan Fu , Ewen Harrison , Lequan Yu , Liantao Ma

Visual compliance verification is a critical yet underexplored problem in computer vision, especially in domains such as media, entertainment, and advertising where content must adhere to complex and evolving policy rules. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Rahul Ghosh , Baishali Chaudhury , Hari Prasanna Das , Meghana Ashok , Ryan Razkenari , Long Chen , Sungmin Hong , Chun-Hao Liu

Autonomous coding agents built on large language models (LLMs) can now solve many general software and machine learning tasks, but they remain ineffective on complex, domain-specific scientific problems. Medical imaging is a particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Roshan Kenia , Xiaoman Zhang , Pranav Rajpurkar

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows.…

Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud-based benchmarking…

As LLM-based agents are increasingly deployed in real-life scenarios, existing benchmarks fail to capture their inherent complexity of handling extensive information, leveraging diverse resources, and managing dynamic user interactions. To…

Computation and Language · Computer Science 2025-10-20 Wei He , Yueqing Sun , Hongyan Hao , Xueyuan Hao , Zhikang Xia , Qi Gu , Chengcheng Han , Dengchang Zhao , Hui Su , Kefeng Zhang , Man Gao , Xi Su , Xiaodong Cai , Xunliang Cai , Yu Yang , Yunke Zhao

Recent advances in vision-language models (VLMs) have achieved remarkable performance on standard medical benchmarks, yet their true clinical reasoning ability remains unclear. Existing datasets predominantly emphasize classification…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Miao Jing , Mengting Jia , Junling Lin , Zhongxia Shen , Huan Gao , Mingkun Xu , Shangyang Li

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents…

LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, making it…

Software Engineering · Computer Science 2026-05-04 Chenxin Li , Zhengyang Tang , Mingxin Huang , Yunlong Lin , Shijue Huang , Shengyuan Liu , Bowen Ye , Rang Li , Lei Li , Benyou Wang , Yixuan Yuan

Brain imaging analysis is crucial for diagnosing and treating brain disorders, and multimodal large language models (MLLMs) are increasingly supporting it. However, current brain imaging visual question-answering (VQA) benchmarks either…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Zhihao Peng , Cheng Wang , Shengyuan Liu , Zhiying Liang , Zanting Ye , Minjie Ju , PeterYM Woo , Yixuan Yuan

LLMs have achieved significant performance progress in various NLP applications. However, LLMs still struggle to meet the strict requirements for accuracy and reliability in the medical field and face many challenges in clinical…

Computation and Language · Computer Science 2024-10-11 Weixiang Yan , Haitian Liu , Tengxiao Wu , Qian Chen , Wen Wang , Haoyuan Chai , Jiayi Wang , Weishan Zhao , Yixin Zhang , Renjun Zhang , Li Zhu , Xuandong Zhao

Scientific research follows multi-turn, multi-step workflows that require proactively searching the literature, consulting figures and tables, and integrating evidence across papers to align experimental settings and support reproducible…

Computation and Language · Computer Science 2026-04-08 Xuan Dong , Huanyang Zheng , Tianhao Niu , Zhe Han , Pengzhan Li , Bofei Liu , Zhengyang Liu , Guancheng Li , Qingfu Zhu , Wanxiang Che

As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Wendong Bu , Yang Wu , Qifan Yu , Minghe Gao , Bingchen Miao , Zhenkui Zhang , Kaihang Pan , Yunfei Li , Mengze Li , Wei Ji , Juncheng Li , Siliang Tang , Yueting Zhuang

Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models (VLMs) as we find in preliminary evaluation. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Fenfen Lin , Yesheng Liu , Haiyu Xu , Chen Yue , Zheqi He , Mingxuan Zhao , Miguel Hu Chen , Jiakang Liu , JG Yao , Xi Yang

MedCalc-Bench is a widely used benchmark for evaluating LLM performance on clinical calculator tasks, with state-of-the-art direct prompting scores plateauing around 35% on the Verified split (HELM MedHELM leaderboard) and the best…

Machine Learning · Computer Science 2026-03-04 Artus Krohn-Grimberghe

As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answering involving domain knowledge and descriptive reasoning.…