中文
相关论文

相关论文: Uncertainty-Aware Structured Data Extraction from …

200 篇论文

Electronic Health Records (EHR)-based disease prediction models have demonstrated significant clinical value in promoting precision medicine and enabling early intervention. However, existing large language models face two major challenges:…

计算与语言 · 计算机科学 2025-06-19 Junke Wang , Hongshun Ling , Li Zhang , Longqian Zhang , Fang Wang , Yuan Gao , Zhi Li

Europe's healthcare systems require enhanced interoperability and digitalization, driving a demand for innovative solutions to process legacy clinical data. This paper presents the results of our project, which aims to leverage Large…

计算与语言 · 计算机科学 2025-07-09 Aynur Guluzade , Naguib Heiba , Zeyd Boukhers , Florim Hamiti , Jahid Hasan Polash , Yehya Mohamad , Carlos A Velasco

Current LLMs for creating fully-structured reports face the challenges of formatting errors, content hallucinations, and privacy leakage issues when uploading data to external servers.We aim to develop an open-source, accurate LLM for…

Structured Outputs from current LLMs exhibit sporadic errors, hindering enterprise AI deployment. We present CONSTRUCT, a real-time uncertainty estimator that scores the trustworthiness of LLM Structured Outputs. Lower-scoring outputs are…

计算与语言 · 计算机科学 2026-04-01 Hui Wen Goh , Jonas Mueller

Large language models (LLMs), including zero-shot and few-shot paradigms, have shown promising capabilities in clinical text generation. However, real-world applications face two key challenges: (1) patient data is highly unstructured,…

计算与语言 · 计算机科学 2025-07-10 Garapati Keerthana , Manik Gupta

Combinatorial medication recommendation(CMR) is a fundamental task of healthcare, which offers opportunities for clinical physicians to provide more precise prescriptions for patients with intricate health conditions, particularly in the…

人工智能 · 计算机科学 2025-01-14 Jie Tan , Yu Rong , Kangfei Zhao , Tian Bian , Tingyang Xu , Junzhou Huang , Hong Cheng , Helen Meng

Deploying multimodal large language models (MLLMs) for clinical summarization demands not only fluent generation but also transparency about where each statement originates-and a mechanism to flag when statements lack evidential support. We…

计算与语言 · 计算机科学 2026-04-21 Qianqi Yan , Huy Nguyen , Sumana Srivatsa , Hari Bandi , Xin Eric Wang , Krishnaram Kenthapadi

The recently introduced Controlled Text Reduction (CTR) task isolates the text generation step within typical summarization-style tasks. It does so by challenging models to generate coherent text conforming to pre-selected content within…

计算与语言 · 计算机科学 2024-02-27 Aviv Slobodkin , Avi Caciularu , Eran Hirsch , Ido Dagan

Electronic Healthcare records contain large volumes of unstructured data in different forms. Free text constitutes a large portion of such data, yet this source of richly detailed information often remains under-used in practice because of…

计算与语言 · 计算机科学 2019-10-17 M. Tarik Altuncu , Erik Mayer , Sophia N. Yaliraki , Mauricio Barahona

Radiology reports capture crucial longitudinal information on tumor burden, treatment response, and disease progression, yet their unstructured narrative format complicates automated analysis. While large language models (LLMs) have…

计算与语言 · 计算机科学 2026-03-12 Luc Builtjes , Alessa Hering

This study applies Large Language Models (LLMs) to two foundational Electronic Health Record (EHR) data science tasks: structured data querying (using programmatic languages, Python/Pandas) and information extraction from unstructured…

计算与语言 · 计算机科学 2026-01-29 Juan Jose Rubio Jan , Jack Wu , Julia Ive

Knowing the effect of an intervention is critical for human decision-making, but current approaches for causal effect estimation rely on manual data collection and structuring, regardless of the causal assumptions. This increases both the…

机器学习 · 计算机科学 2024-10-29 Nikita Dhawan , Leonardo Cotta , Karen Ullrich , Rahul G. Krishnan , Chris J. Maddison

Pre-consultation is a critical component of effective healthcare delivery. However, generating comprehensive pre-consultation questionnaires from complex, voluminous Electronic Medical Records (EMRs) is a challenging task. Direct Large…

人工智能 · 计算机科学 2025-08-04 Ruiqing Ding , Qianfang Sun , Yongkang Leng , Hui Yin , Xiaojian Li

Electronic health records (EHRs) are invaluable for clinical research, yet privacy concerns severely restrict data sharing. Synthetic data generation offers a promising solution, but EHRs present unique challenges: they contain both…

机器学习 · 计算机科学 2026-03-26 Shaonan Liu , Yuichiro Iwashita , Soichiro Nakako , Masakazu Iwamura , Koichi Kise

Deploying accurate Text-to-SQL systems at the enterprise level faces a difficult trilemma involving cost, security and performance. Current solutions force enterprises to choose between expensive, proprietary Large Language Models (LLMs)…

计算与语言 · 计算机科学 2026-03-13 Khushboo Thaker , Yony Bresler

Spectral clustering is a popular unsupervised learning technique which is able to partition unlabelled data into disjoint clusters of distinct shapes. However, the data under consideration are often experimental data, implying that the data…

机器学习 · 统计学 2025-05-26 Jürgen Dölz , Jolanda Weygandt

Accurate and robust segmentation of lung cancers from CT, even those located close to mediastinum, is needed to more accurately plan and deliver radiotherapy and to measure treatment response. Therefore, we developed a new cross-modality…

图像与视频处理 · 电气工程与系统科学 2021-12-08 Jue Jiang , Andreas Rimner , Joseph O. Deasy , Harini Veeraraghavan

Numerical models are increasingly used for non-invasive diagnosis and treatment planning in coronary artery disease, where service-based technologies have proven successful in identifying hemodynamically significant and hence potentially…

医学物理 · 物理学 2020-05-01 Jongmin Seo , Casey Fleeter , Andrew M. Kahn , Alison L. Marsden , Daniele E. Schiavazzi

Deployed language models must decide not only what to answer but also when not to answer. We present UniCR, a unified framework that turns heterogeneous uncertainty evidence including sequence likelihoods, self-consistency dispersion,…

This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLMs) to predict data entries based on relevant sections of…