English
Related papers

Related papers: DALPHIN: Benchmarking Digital Pathology AI Copilot…

200 papers

Automatic speech intelligibility assessment is crucial for monitoring speech disorders and therapy efficacy. However, existing methods are difficult to compare: research is fragmented across private datasets with inconsistent protocols. We…

Sound · Computer Science 2026-03-10 Bence Mark Halpern , Thomas Tienkamp , Defne Abur , Tomoki Toda

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, feature expert…

The complexity and variability inherent in high-resolution pathological images present significant challenges in computational pathology. While pathology foundation models leveraging AI have catalyzed transformative advancements, their…

AI based mental health diagnosis is often judged by benchmark accuracy, yet in practice its value depends on how psychologists respond whether they accept, adjust, or reject AI suggestions. Mental health makes this especially challenging:…

Human-Computer Interaction · Computer Science 2025-12-11 Filippo Cenacchi , Longbing Cao , Deborah Richards

The rapid emergence of large language models (LLMs) has raised urgent questions across the modern workforce about this new technology's strengths, weaknesses, and capabilities. For privacy professionals, the question is whether these AI…

Computers and Society · Computer Science 2025-08-13 Zane Witherspoon , Thet Mon Aye , YingYing Hao

Artificial intelligence systems are now deployed at scale across sectors, accompanied by a growing number of real-world incidents ranging from misinformation and cybercrime to autonomous-system failures. Databases of AI incidents index…

Computers and Society · Computer Science 2026-04-23 Sophia Abraham , Taiye Chen , Cyril Chhun , Giovanna Jaramillo-Gutierrez , Simon Mylius , Sayash Raaj , Peter Slattery , Sean McGregor

Artificial Intelligence (AI) chatbots leveraging Large Language Models (LLMs) are gaining traction in healthcare for their potential to automate patient interactions and aid clinical decision-making. This study examines the reliability of…

Artificial Intelligence · Computer Science 2024-05-24 Ayesha Siddika Nipu , K M Sajjadul Islam , Praveen Madiraju

We present DM-Bench, the first benchmark designed to evaluate large language model (LLM) performance across real-world decision-making tasks faced by individuals managing diabetes in their daily lives. Unlike prior health benchmarks that…

Machine Learning · Computer Science 2025-10-06 Maria Ana Cardei , Josephine Lamp , Mark Derdzinski , Karan Bhatia

This paper addresses complex challenges in histopathological image analysis through three key contributions. Firstly, it introduces a fast patch selection method, FPS, for whole-slide image (WSI) analysis, significantly reducing…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Saghir Alfasly , Abubakr Shafique , Peyman Nejat , Jibran Khan , Areej Alsaafin , Ghazal Alabtah , H. R. Tizhoosh

This study systematically evaluates 27 frontier Large Language Models on eight biology benchmarks spanning molecular biology, genetics, cloning, virology, and biosecurity. Models from major AI developers released between November 2022 and…

Machine Learning · Computer Science 2025-05-23 Lennart Justen

We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop dataset creation: annotators seek to create examples that a target…

The growing demand for accurate and equitable AI models in digital dermatology faces a significant challenge: the lack of diverse, high-quality labeled data. In this work, we investigate the potential of domain-specific foundation models…

Curating high-quality, domain-specific datasets is a major bottleneck for deploying robust vision systems, requiring complex trade-offs between data quality, diversity, and cost when researching vast, unlabeled data lakes. We introduce…

Rationale and Objectives: To develop and validate PARROT (Polyglottal Annotated Radiology Reports for Open Testing), a large, multicentric, open-access dataset of fictional radiology reports spanning multiple languages for testing natural…

Computation and Language · Computer Science 2025-08-26 Bastien Le Guellec , Kokou Adambounou , Lisa C Adams , Thibault Agripnidis , Sung Soo Ahn , Radhia Ait Chalal , Tugba Akinci D Antonoli , Philippe Amouyel , Henrik Andersson , Raphael Bentegeac , Claudio Benzoni , Antonino Andrea Blandino , Felix Busch , Elif Can , Riccardo Cau , Armando Ugo Cavallo , Christelle Chavihot , Erwin Chiquete , Renato Cuocolo , Eugen Divjak , Gordana Ivanac , Barbara Dziadkowiec Macek , Armel Elogne , Salvatore Claudio Fanni , Carlos Ferrarotti , Claudia Fossataro , Federica Fossataro , Katarzyna Fulek , Michal Fulek , Pawel Gac , Martyna Gachowska , Ignacio Garcia Juarez , Marco Gatti , Natalia Gorelik , Alexia Maria Goulianou , Aghiles Hamroun , Nicolas Herinirina , Krzysztof Kraik , Dominik Krupka , Quentin Holay , Felipe Kitamura , Michail E Klontzas , Anna Kompanowska , Rafal Kompanowski , Alexandre Lefevre , Tristan Lemke , Maximilian Lindholz , Lukas Muller , Piotr Macek , Marcus Makowski , Luigi Mannacio , Aymen Meddeb , Antonio Natale , Beatrice Nguema Edzang , Adriana Ojeda , Yae Won Park , Federica Piccione , Andrea Ponsiglione , Malgorzata Poreba , Rafal Poreba , Philipp Prucker , Jean Pierre Pruvo , Rosa Alba Pugliesi , Feno Hasina Rabemanorintsoa , Vasileios Rafailidis , Katarzyna Resler , Jan Rotkegel , Luca Saba , Ezann Siebert , Arnaldo Stanzione , Ali Fuat Tekin , Liz Toapanta Yanchapaxi , Matthaios Triantafyllou , Ekaterini Tsaoulia , Evangelia Vassalou , Federica Vernuccio , Johan Wasselius , Weilang Wang , Szymon Urban , Adrian Wlodarczak , Szymon Wlodarczak , Andrzej Wysocki , Lina Xu , Tomasz Zatonski , Shuhang Zhang , Sebastian Ziegelmayer , Gregory Kuchcinski , Keno K Bressem

We introduce DABstep, a novel benchmark for evaluating AI agents on realistic multi-step data analysis tasks. DABstep comprises over 450 real-world challenges derived from a financial analytics platform, requiring models to combine…

Machine Learning · Computer Science 2025-07-01 Alex Egg , Martin Iglesias Goyanes , Friso Kingma , Andreu Mora , Leandro von Werra , Thomas Wolf

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural…

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on…

Computation and Language · Computer Science 2025-10-13 Jie Liu , Wenxuan Wang , Zizhan Ma , Guolin Huang , Yihang SU , Kao-Jung Chang , Wenting Chen , Haoliang Li , Linlin Shen , Michael Lyu

Large Language Models (LLMs) like GPT-4, MedPaLM-2, and Med-Gemini achieve performance competitively with human experts across various medical benchmarks. However, they still face challenges in making professional diagnoses akin to…

Computation and Language · Computer Science 2024-08-23 Xiaohan Wang , Xiaoyan Yang , Yuqi Zhu , Yue Shen , Jian Wang , Peng Wei , Lei Liang , Jinjie Gu , Huajun Chen , Ningyu Zhang

Difficulty replicating baselines, high computational costs, and required domain expertise create persistent barriers to clinical AI research. To address these challenges, we introduce PyHealth 2.0, an enhanced clinical deep learning toolkit…

The microscopic examination of surgical tissue remains a cornerstone of disease classification but relies on subjective interpretations and access to highly specialized experts, which can compromise accuracy and clinical care. While…

‹ Prev 1 4 5 6 7 8 10 Next ›