English
Related papers

Related papers: MedDialBench: Benchmarking LLM Diagnostic Robustne…

200 papers

Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-world clinical settings. However, clinical decision-making is inherently safety-critical,…

Computation and Language · Computer Science 2026-04-13 Xiaohan Ren , Chenxiao Fan , Wenyin Ma , Hongliang He , Chongming Gao , Xiaoyan Zhao , Fuli Feng

Extended interaction with large language models (LLMs) has been linked to the reinforcement of delusional beliefs, a phenomenon attracting growing clinical and public concern. Yet most empirical work evaluates model safety in brief…

Human-Computer Interaction · Computer Science 2026-04-24 Luke Nicholls , Robert Hutto , Zephrah Soto , Hamilton Morrin , Thomas Pollak , Raj Korpan , Cheryl Carmichael

LLMs are increasingly used as long-running conversational agents, yet every major benchmark evaluating their memory treats user information as static facts to be stored and retrieved. That's the wrong model. People change their minds, and…

Computation and Language · Computer Science 2026-03-26 Praveen Kumar Myakala , Manan Agrawal , Rahul Manche

Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For more realistic interactive experience with customized styles,…

Computation and Language · Computer Science 2026-03-10 Haishu Zhao , Aokai Hao , Yuan Ge , Zhenqiang Hong , Tong Xiao , Jingbo Zhu

Despite the impressive capabilities of Large Language Models (LLMs), existing Conversational Health Agents (CHAs) remain static and brittle, incapable of adaptive multi-turn reasoning, symptom clarification, or transparent decision-making.…

Computation and Language · Computer Science 2025-07-11 Xinyi Liu , Dachun Sun , Yi R. Fung , Dilek Hakkani-Tür , Tarek Abdelzaher

Reasoning about Actions and Change (RAC) has historically played a pivotal role in solving foundational AI problems, such as the frame problem. It has driven advancements in AI fields, such as non-monotonic and commonsense reasoning. RAC…

Computational Complexity · Computer Science 2025-03-04 Divij Handa , Pavel Dolin , Shrinidhi Kumbhar , Tran Cao Son , Chitta Baral

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique,…

Computation and Language · Computer Science 2024-01-23 Chen Zhang , Luis Fernando D'Haro , Yiming Chen , Malu Zhang , Haizhou Li

Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud-based benchmarking…

As the performance of large language models (LLMs) continues to advance, their adoption in the medical domain is increasing. However, most existing risk evaluations largely focused on general safety benchmarks. In the medical applications,…

Computation and Language · Computer Science 2026-01-12 Jean-Philippe Corbeil , Minseon Kim , Maxime Griot , Sheela Agarwal , Alessandro Sordoni , Francois Beaulieu , Paul Vozila

Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved from using probabilistic correctness estimates to, more recently, eliciting verbalized…

Computation and Language · Computer Science 2026-05-11 Sree Bhattacharyya , Samarth Khanna , Leona Chen , Lucas Craig , Tharun Dilliraj , James Z. Wang

Large language models (LLMs) have shown promise in clinical diagnosis but remain limited by unreliable report generation, weak evidence grounding, and opaque reasoning. We propose MedCollab, an IBIS-guided multi-agent framework for…

Multiagent Systems · Computer Science 2026-05-27 Yuqi Zhan , Xinyue Wu , Tianyu Lin , Yutong Bao , Xiaoyu Wang , Weihao Cheng , Huangwei Chen , Feiwei Qin , Zhu Zhu

People often encounter role conflicts -- social dilemmas where the expectations of multiple roles clash and cannot be simultaneously fulfilled. As large language models (LLMs) increasingly navigate these social dynamics, a critical research…

Computation and Language · Computer Science 2026-04-20 Jisu Shin , Hoyun Song , Juhyun Oh , Changgeon Ko , Eunsu Kim , Chani Jung , Alice Oh

We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs). Our benchmark comprises 660 prompts across six…

Computation and Language · Computer Science 2025-03-17 Esben Kran , Hieu Minh "Jord" Nguyen , Akash Kundu , Sami Jawhar , Jinsuk Park , Mateusz Maria Jurewicz

Recent advances in large language models (LLMs) and medical LLMs (Med-LLMs) have demonstrated strong performance on general medical benchmarks. However, their capabilities in specialized medical fields, such as dentistry which require…

Computation and Language · Computer Science 2025-08-29 Hengchuan Zhu , Yihuan Xu , Yichen Li , Zijie Meng , Zuozhu Liu

Current large language models (LLMs) excel in verifiable domains where outputs can be checked before action but prove less reliable for high-stakes strategic decisions with uncertain outcomes. This gap, driven by mutually reinforcing…

Artificial Intelligence · Computer Science 2025-11-12 Alejandro R. Jadad

Despite significant advancements in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), current models still face substantial challenges in handling complex, multi-turn, and visually-grounded tasks that demand deep…

Computation and Language · Computer Science 2025-08-22 Seungmin Han , Haeun Kwon , Ji-jun Park , Taeyang Yoon

Passively collected behavioral health data from ubiquitous sensors holds significant promise to provide mental health professionals insights from patient's daily lives; however, developing analysis tools to use this data in clinical…

Large language models (LLMs) have attracted growing interest as supportive tools for psychiatric assessment and clinical decision support. However, existing mental health benchmarks largely rely on social media data or supportive dialogue…

Computation and Language · Computer Science 2026-05-19 Hoyun Song , Migyeong Kang , Jisu Shin , Jihyun Kim , Chanbi Park , Hangyeol Yoo , Jihyun An , Alice Oh , Jinyoung Han , KyungTae Lim

Although large language models (LLMs) are increasingly trained using human feedback for safety and alignment with human values, alignment decisions often overlook human social diversity. This study examines how incorporating pluralistic…

Artificial Intelligence · Computer Science 2025-11-27 Dalia Ali , Dora Zhao , Allison Koenecke , Orestis Papakyriakopoulos

Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be objectively defined by the final answer, medical diagnosis…

Computation and Language · Computer Science 2025-05-21 Kevin Wu , Eric Wu , Rahul Thapa , Kevin Wei , Angela Zhang , Arvind Suresh , Jacqueline J. Tao , Min Woo Sun , Alejandro Lozano , James Zou