English
Related papers

Related papers: PsychEval: A Multi-Session and Multi-Therapy Bench…

200 papers

Learning therapeutic counseling involves significant role-play experience with mock patients, with current manual training methods providing only intermittent granular feedback. We seek to accelerate and optimize counselor training by…

Human-Computer Interaction · Computer Science 2025-02-27 Ian Steenstra , Farnaz Nouraei , Timothy W. Bickmore

The development of AI for mental health is hindered by a lack of authentic therapy dialogues, due to strict privacy regulations and the fact that clinical sessions were historically rarely recorded. We present an LLM-driven pipeline that…

Evaluating the clinical correctness and reasoning fidelity of automatically generated medical imaging reports remains a critical yet unresolved challenge. Existing evaluation methods often fail to capture the structured diagnostic logic…

Artificial Intelligence · Computer Science 2026-01-26 Suzhong Fu , Jingqi Dong , Xuan Ding , Rui Sun , Yiming Yang , Shuguang Cui , Zhen Li

AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the development of useful agents, we need benchmarks that are challenging, but more crucially, directly…

Computation and Language · Computer Science 2024-09-18 Zachary S. Siegel , Sayash Kapoor , Nitya Nagdir , Benedikt Stroebl , Arvind Narayanan

Traditional in-person psychological counseling remains primarily niche, often chosen by individuals with psychological issues, while online automated counseling offers a potential solution for those hesitant to seek help due to feelings of…

The rapid advancement in AI-generated video synthesis has led to a growth demand for standardized and effective evaluation metrics. Existing metrics lack a unified framework for systematically categorizing methodologies, limiting a holistic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Xinhao Xiang , Xiao Liu , Zizhong Li , Zhuosheng Liu , Jiawei Zhang

Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory benchmarks primarily target declarative memory, specifically…

The in-context learning capabilities of large language models (LLMs) show great potential in mental health support. However, the lack of counseling datasets, particularly in Chinese corpora, restricts their application in this field. To…

Computation and Language · Computer Science 2025-03-06 Keqi Chen , Zekai Sun , Yuhua Wen , Huijun Lian , Yingming Gao , Ya Li

AI chatbots already function as de facto mental health support tools for millions of people, including people in crisis. Yet, they lack the clinical validation, shared standards, and coordinated oversight that their societal role demands.…

Computers and Society · Computer Science 2026-05-11 Emily Saltz , Claire R. Leibowicz

Curated datasets for healthcare are often limited due to the need of human annotations from experts. In this paper, we present MedEval, a multi-level, multi-task, and multi-domain medical benchmark to facilitate the development of language…

Computation and Language · Computer Science 2023-11-16 Zexue He , Yu Wang , An Yan , Yao Liu , Eric Y. Chang , Amilcare Gentili , Julian McAuley , Chun-Nan Hsu

Inspite the emerging importance of Speech Emotion Recognition (SER), the state-of-the-art accuracy is quite low and needs improvement to make commercial applications of SER viable. A key underlying reason for the low accuracy is the…

Sound · Computer Science 2020-03-24 Siddique Latif , Rajib Rana , Sara Khalifa , Raja Jurdak , Julien Epps , Björn W. Schuller

AI Clones aim to simulate an individual's thoughts and behaviors to enable long-term, personalized interaction, placing stringent demands on memory systems to model experiences, emotions, and opinions over time. Existing memory benchmarks…

Artificial Intelligence · Computer Science 2026-01-13 Sen Hu , Zhiyu Zhang , Yuxiang Wei , Xueran Han , Zhenheng Tang , Huacan Wang , Ronghao Chen

Next-generation AI must manage vast personal data, diverse tools, and multi-step reasoning, yet most benchmarks remain context-free and single-turn. We present ASTRA-bench (Assistant Skills in Tool-use, Reasoning \& Action-planning), a…

Artificial Intelligence · Computer Science 2026-03-03 Zidi Xiu , David Q. Sun , Kevin Cheng , Maitrik Patel , Josh Date , Yizhe Zhang , Jiarui Lu , Omar Attia , Raviteja Vemulapalli , Oncel Tuzel , Meng Cao , Samy Bengio

Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential for real-world practice. While recent interactive benchmarks…

Artificial Intelligence · Computer Science 2026-01-21 Chuhan Qiao , Jianghua Huang , Daxing Zhao , Ziding Liu , Yanjun Shen , Bing Cheng , Wei Lin , Kai Wu

As AI becomes more deeply embedded in knowledge work, building assistants that support human creativity and expertise becomes more important. Yet achieving synergy in human-AI collaboration is not easy. Providing AI with detailed…

Human-Computer Interaction · Computer Science 2026-04-22 Sean Kelley , David De Cremer , Christoph Riedl

Voice AI agents are rapidly transitioning to production deployments, yet systematic methods for ensuring testing reliability remain underdeveloped. Organizations cannot objectively assess whether their testing approaches (internal tools or…

Artificial Intelligence · Computer Science 2026-01-15 Miguel E. Andres , Vadim Fedorov , Rida Sadek , Enric Spagnolo-Arrizabalaga , Nadescha Trudel

The advent of Large Language Models (LLMs) offers potential solutions to address problems such as shortage of medical resources and low diagnostic consistency in psychiatric clinical practice. Despite this potential, a robust and…

Computation and Language · Computer Science 2025-06-19 Shuyu Liu , Ruoxi Wang , Ling Zhang , Xuequan Zhu , Rui Yang , Xinzhu Zhou , Fei Wu , Zhi Yang , Cheng Jin , Gang Wang

As conversational AI-based dialogue management has increasingly become a trending topic, the need for a standardized and reliable evaluation procedure grows even more pressing. The current state of affairs suggests various evaluation…

Computation and Language · Computer Science 2020-06-12 Sarah E. Finch , Jinho D. Choi

Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, where each turn must be stepped (i.e., prompted) by the user. Open-ended, asynchronous…

End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules;…