中文
相关论文

相关论文: AcuityBench: Evaluating Clinical Acuity Identifica…

200 篇论文

Large language models (LLMs) are increasingly used for medical consultation and health information support. In this high-stakes setting, safety depends not only on medical knowledge, but also on how models respond when patient inputs are…

计算与语言 · 计算机科学 2026-04-01 Yahan Li , Xinyi Jie , Wanjia Ruan , Xubei Zhang , Huaijie Zhu , Yicheng Gao , Chaohao Du , Ruishan Liu

Recent benchmarks for medical Large Vision-Language Models (LVLMs) emphasize leaderboard accuracy, overlooking reliability and safety. We study sycophancy -- models' tendency to uncritically echo user-provided information -- in high-stakes…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Botai Yuan , Yutian Zhou , Yingjie Wang , Fushuo Huo , Yongcheng Jing , Li Shen , Ying Wei , Zhiqi Shen , Ziwei Liu , Tianwei Zhang , Jie Yang , Dacheng Tao

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench,…

计算与语言 · 计算机科学 2024-10-07 Zetian Ouyang , Yishuai Qiu , Linlin Wang , Gerard de Melo , Ya Zhang , Yanfeng Wang , Liang He

Large language models (LLMs) are increasingly integrated into legal drafting and research workflows, where incorrect citations or fabricated precedents can cause serious professional harm. Existing legal benchmarks largely emphasize…

计算与语言 · 计算机科学 2026-05-12 Sijia Chen , Hang Yin , Shunfan Zhou

Evaluating alignment in language models requires testing how they behave under realistic pressure, not just what they claim they would do. While alignment failures increasingly cause real-world harm, comprehensive evaluation frameworks with…

人工智能 · 计算机科学 2026-02-25 Nora Petrova , John Burden

Large Language Models (LLMs) are increasingly deployed as scientific AI as- sistants, and a growing body of benchmarks evaluates their capabilities across knowledge retrieval, reasoning, code generation, and tool use. These evaluations,…

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs'…

Demand for mental health support through AI chatbots is surging, though current systems present several limitations, like sycophancy or overvalidation, and reinforcement of maladaptive beliefs. A core obstacle to the creation of better…

计算与语言 · 计算机科学 2025-12-08 José Pombal , Maya D'Eon , Nuno M. Guerreiro , Pedro Henrique Martins , António Farinhas , Ricardo Rei

Large language models perform well on static medical examinations, yet clinical diagnosis often requires iterative evidence gathering under uncertainty. Building on prior interactive evaluation efforts, we introduce an OSCE-inspired…

人工智能 · 计算机科学 2026-05-22 Chen Zhan , Xihe Qiu , Xiaoyu Tan , Xibing Zhuang , Gengchen Ma , Yue Zhang , Shuo Li , Peifeng Liu , Xiaoxiao Ge , Liang Liu , Lu Gan

Intraoperative monitoring and prediction of vital signs are critical for ensuring patient safety and improving surgical outcomes. Despite recent advances in deep learning models for medical time-series forecasting, several challenges…

机器学习 · 计算机科学 2025-11-19 Xiuding Cai , Xueyao Wang , Sen Wang , Yaoyao Zhu , Jiao Chen , Yu Yao

Multimodal large language models (MLLMs) demonstrate considerable potential in clinical diagnostics, a domain that inherently requires synthesizing complex visual and textual data alongside consulting authoritative medical literature.…

计算与语言 · 计算机科学 2026-03-23 Yannian Gu , Zhongzhen Huang , Linjie Mu , Xizhuo Zhang , Shaoting Zhang , Xiaofan Zhang

Text-guided image editing has seen significant progress in natural image domains, but its application in medical imaging remains limited and lacks standardized evaluation frameworks. Such editing could revolutionize clinical practices by…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Minghao Liu , Zhitao He , Zhiyuan Fan , Qingyun Wang , Yi R. Fung

The rapid evolution of Large Language Models (LLMs) presents a promising solution to the global shortage of mental health professionals. However, their alignment with essential counseling competencies remains underexplored. We introduce…

Language models have demonstrated remarkable capabilities on standard benchmarks, yet they struggle increasingly from mode collapse, the inability to generate diverse and novel outputs. Our work introduces NoveltyBench, a benchmark…

计算与语言 · 计算机科学 2025-08-12 Yiming Zhang , Harshita Diddee , Susan Holm , Hanchen Liu , Xinyue Liu , Vinay Samuel , Barry Wang , Daphne Ippolito

Clinical language processing has received a lot of attention in recent years, resulting in new models or methods for disease phenotyping, mortality prediction, and other tasks. Unfortunately, many of these approaches are tested under…

计算与语言 · 计算机科学 2022-09-30 Travis R. Goodwin , Dina Demner-Fushman

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows.…

Large Language Models (LLMs) are increasingly being explored for clinical question answering and decision support, yet safe deployment critically requires reliable handling of patient measurements in heterogeneous clinical notes. Existing…

计算与语言 · 计算机科学 2026-04-16 Minh-Vuong Nguyen , Fatemeh Shiri , Zhuang Li , Karin Verspoor

As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly urgent. Existing benchmarks and evaluations often focus on…

As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric,…

计算与语言 · 计算机科学 2026-05-28 Junyu Liu , Zirui Li , Qian Niu , Zequn Zhang , Yue Xun , Wenlong Hou , Shujun Wang , Yusuke Iwasawa , Yutaka Matsuo , Kan Hatakeyama-Sato

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly…

计算与语言 · 计算机科学 2026-02-16 Yangzhuo Li , Shengpeng Ji , Yifu Chen , Tianle Liang , Haorong Ying , Yule Wang , Junbo Li , Jun Fang , Zhou Zhao