English
Related papers

Related papers: HARE: an entity and relation centric evaluation fr…

200 papers

Finetuning specialized generative evaluators has emerged as a popular paradigm to meet the increasing demand for scalable evaluation during both training and test-time. However, recent work has largely focused on applying new methodology,…

Computation and Language · Computer Science 2025-11-20 Austin Xu , Xuan-Phi Nguyen , Yilun Zhou , Chien-Sheng Wu , Caiming Xiong , Shafiq Joty

This technical report introduces a Named Clinical Entity Recognition Benchmark for evaluating language models in healthcare, addressing the crucial natural language processing (NLP) task of extracting structured information from clinical…

Graph-based methods have been extensively applied to whole-slide histopathology image (WSI) analysis due to the advantage of modeling the spatial relationships among different entities. However, most of the existing methods focus on…

Computer Vision and Pattern Recognition · Computer Science 2023-07-11 Tsai Hor Chan , Fernando Julio Cendra , Lan Ma , Guosheng Yin , Lequan Yu

Evaluating automatically generated radiology reports remains a fundamental challenge due to the lack of clinically grounded, interpretable, and fine-grained metrics. Existing methods either produce coarse overall scores or rely on opaque…

Computation and Language · Computer Science 2025-08-22 Yingshu Li , Yunyi Liu , Lingqiao Liu , Lei Wang , Luping Zhou

Human evaluation of machine translation normally uses sentence-level measures such as relative ranking or adequacy scales. However, these provide no insight into possible errors, and do not scale well with sentence length. We argue for a…

Computation and Language · Computer Science 2016-09-28 Alexandra Birch , Omri Abend , Ondrej Bojar , Barry Haddow

Traditional automatic evaluation metrics for machine translation have been widely criticized by linguists due to their low accuracy, lack of transparency, focus on language mechanics rather than semantics, and low agreement with human…

Computation and Language · Computer Science 2021-12-28 Serge Gladkoff , Lifeng Han

As medical imaging is central to diagnostic processes, automating the generation of radiology reports has become increasingly relevant to assist radiologists with their heavy workloads. Most current methods rely solely on global image…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Hamza Kalisch , Fabian Hörst , Jens Kleesiek , Ken Herrmann , Constantin Seibold

Purpose: To develop and evaluate an automated system for extracting structured clinical information from unstructured radiology and pathology reports using open-weights large language models (LMs) and retrieval augmented generation (RAG),…

Phenotyping electronic health records (EHR) focuses on defining meaningful patient groups (e.g., heart failure group and diabetes group) and identifying the temporal evolution of patients in those groups. Tensor factorization has been an…

Machine Learning · Computer Science 2019-11-15 Ardavan Afshar , Ioakeim Perros , Haesun Park , Christopher deFilippi , Xiaowei Yan , Walter Stewart , Joyce Ho , Jimeng Sun

Retrieval-Augmented Generation (RAG) shows promise for enterprise knowledge work, yet it often underperforms in high-stakes decision settings that require deep synthesis, strict traceability, and recovery from underspecified prompts.…

Information Retrieval · Computer Science 2026-01-27 Xincheng You , Qi Sun , Neha Bora , Huayi Li , Shubham Goel , Kang Li , Sean Culatana

Automatic Speech Recognition (ASR) in medical contexts has the potential to save time, cut costs, increase report accuracy, and reduce physician burnout. However, the healthcare industry has been slower to adopt this technology, in part due…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-14 Joel Shor , Ruyue Agnes Bi , Subhashini Venugopalan , Steven Ibara , Roman Goldenberg , Ehud Rivlin

Identifying patient cohorts is fundamental to numerous healthcare tasks, including clinical trial recruitment and retrospective studies. Current cohort retrieval methods in healthcare organizations rely on automated queries of structured…

Recent advances in large language models have enabled deep research systems that generate expert-level reports through multi-step reasoning and evidence-based synthesis. However, evaluating such reports remains challenging: report quality…

Computation and Language · Computer Science 2026-03-11 Janghoon Han , Heegyu Kim , Changho Lee , Dahm Lee , Min Hyung Park , Hosung Song , Stanley Jungkyu Choi , Moontae Lee , Honglak Lee

Even though many machine algorithms have been proposed for entity resolution, it remains very challenging to find a solution with quality guarantees. In this paper, we propose a novel HUman and Machine cOoperation (HUMO) framework for…

Databases · Computer Science 2018-04-03 Zhaoqiang Chen , Qun Chen , Fengfeng Fan , Yanyan Wang , Zhuo Wang , Youcef Nafa , Zhanhuai Li , Hailong Liu , Wei Pan

Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on large-scale real-world data such as electronic health records…

Extracting causal relationships from a medical case report is essential for comprehending the case, particularly its diagnostic process. Since the diagnostic process is regarded as a bottom-up inference, causal relationships in cases…

Computation and Language · Computer Science 2025-03-04 Sakiko Yahata , Zhen Wan , Fei Cheng , Sadao Kurohashi , Hisahiko Sato , Ryozo Nagai

Many studies have examined the shortcomings of word error rate (WER) as an evaluation metric for automatic speech recognition (ASR) systems. Since WER considers only literal word-level correctness, new evaluation metrics based on semantic…

Computation and Language · Computer Science 2023-12-04 Zitha Sasindran , Harsha Yelchuri , T. V. Prabhakar , Supreeth Rao

Biomedical named entity recognition (NER) and entity linking (EL) strongly depend on annotated corpora, but the utility of these resources for benchmarking is often assumed rather than characterized. We present a corpus-centric framework…

Computation and Language · Computer Science 2026-05-21 Robert Leaman , Rezarta Islamaj , Zhiyong Lu

Objective: Currently, a major limitation for natural language processing (NLP) analyses in clinical applications is that a concept can be referenced in various forms across different texts. This paper introduces Multi-Ontology Refined…

Computation and Language · Computer Science 2020-04-15 Steven Jiang , Weiyi Wu , Naofumi Tomita , Craig Ganoe , Saeed Hassanpour

Explainable AI (XAI) in medical histopathology is essential for enhancing the interpretability and clinical trustworthiness of deep learning models in cancer diagnosis. However, the black-box nature of these models often limits their…

Image and Video Processing · Electrical Eng. & Systems 2025-05-06 Raktim Kumar Mondol , Ewan K. A. Millar , Peter H. Graham , Lois Browne , Arcot Sowmya , Erik Meijering