中文
相关论文

相关论文: ELMV: an Ensemble-Learning Approach for Analyzing …

200 篇论文

Linking textual values in tabular data to their corresponding entities in a Knowledge Base is a core task across a variety of data integration and enrichment applications. Although Large Language Models (LLMs) have shown State-of-The-Art…

计算与语言 · 计算机科学 2025-10-03 Carlo Bono , Federico Belotti , Matteo Palmonari

Due to the increasing adoption of electronic health records (EHR), large scale EHRs have become another rich data source for translational clinical research. Despite its potential, deriving generalizable knowledge from EHR data remains…

机器学习 · 统计学 2023-06-01 Junwei Lu , Jin Yin , Tianxi Cai

With the widespread of machine learning models for healthcare applications, there is increased interest in building applications for personalized medicine. Despite the plethora of proposed research for personalized medicine, very few focus…

机器学习 · 计算机科学 2024-02-27 Ghadeer O. Ghosheh , Jin Li , Tingting Zhu

Envelope method was recently proposed as a method to reduce the dimension of responses in multivariate regressions. However, when there exists missing data, the envelope method using the complete case observations may lead to biased and…

统计方法学 · 统计学 2021-03-25 Linquan Ma , Lan Liu , Wei Yang

Fraud in healthcare is widespread, as doctors could prescribe unnecessary treatments to increase bills. Insurance companies want to detect these anomalous fraudulent bills and reduce their losses. Traditional fraud detection methods use…

机器学习 · 计算机科学 2020-10-13 Victoria Snorovikhina , Alexey Zaytsev

We consider the problem of handling missing data with deep latent variable models (DLVMs). First, we present a simple technique to train DLVMs when the training set contains missing-at-random data. Our approach, called MIWAE, is based on…

机器学习 · 统计学 2019-02-05 Pierre-Alexandre Mattei , Jes Frellsen

The use of Large Language Models (LLMs) to support patients in addressing medical questions is becoming increasingly prevalent. However, most of the measures currently used to evaluate the performance of these models in this context only…

人机交互 · 计算机科学 2026-04-22 Abu Noman Md Sakib , Md. Main Oddin Chisty , Zijie Zhang

Large Language Models (LLMs) have shown strong promise for mining Electronic Health Records (EHRs) by reasoning over longitudinal clinical information to capture context-rich patient trajectories. However, leveraging LLMs for structured…

计算与语言 · 计算机科学 2026-04-21 Arya Hadizadeh Moghaddam , Drew Ross , Mohsen Nayebi Kerdabadi , Dongjie Wang , Zijun Yao

Distributed representations of medical concepts have been used to support downstream clinical tasks recently. Electronic Health Records (EHR) capture different aspects of patients' hospital encounters and serve as a rich source for…

计算与语言 · 计算机科学 2020-01-07 Shaika Chowdhury , Chenwei Zhang , Philip S. Yu , Yuan Luo

This paper aims to address the challenge of sparse and missing data in recommendation systems, a significant hurdle in the age of big data. Traditional imputation methods struggle to capture complex relationships within the data. We propose…

信息检索 · 计算机科学 2024-08-09 Zhicheng Ding , Jiahao Tian , Zhenkai Wang , Jinman Zhao , Siyang Li

Missing data are inevitable in longitudinal studies. Traditional methods, such as the full information maximum likelihood (FIML), are commonly used to handle ignorable missing data. However, they may lead to biased model estimation due to…

应用统计 · 统计学 2024-01-01 Dandan Tang , Xin Tong

Missing values exist in nearly all clinical studies because data for a variable or question are not collected or not available. Inadequate handling of missing values can lead to biased results and loss of statistical power in analysis.…

机器学习 · 计算机科学 2021-03-04 Narges Pourshahrokhi , Samaneh Kouchaki , Kord M. Kober , Christine Miaskowski , Payam Barnaghi

Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional data curation strategies that rely on coarse-grained partitioning by modality or…

计算与语言 · 计算机科学 2026-04-29 Jianghang Lin , Haihua Yang , Deli Yu , Kai Wu , Kai Ye , Jinghao Lin , Zihan Wang , Yuhang Wu , Liujuan Cao

Often in real-world datasets, especially in high dimensional data, some feature values are missing. Since most data analysis and statistical methods do not handle gracefully missing values, the first step in the analysis requires the…

机器学习 · 统计学 2016-12-08 Yehezkel S. Resheff , Daphna Weinshall

Product attribute value extraction is a pivotal component in Natural Language Processing (NLP) and the contemporary e-commerce industry. The provision of precise product attribute values is fundamental in ensuring high-quality…

信息检索 · 计算机科学 2024-06-21 Chenhao Fang , Xiaohan Li , Zezhong Fan , Jianpeng Xu , Kaushiki Nag , Evren Korpeoglu , Sushant Kumar , Kannan Achan

In an era defined by rapid data evolution, traditional Machine Learning (ML) models often struggle to adapt to dynamic environments. Evolving Machine Learning (EML) has emerged as a pivotal paradigm, enabling continuous learning and…

Commonly, AI or machine learning (ML) models are evaluated on benchmark datasets. This practice supports innovative methodological research, but benchmark performance can be poorly correlated with performance in real-world applications -- a…

机器学习 · 计算机科学 2024-06-18 Olivier Binette , Jerome P. Reiter

Electronic health records (EHR) often contain varying levels of missing data. This study compared different imputation strategies to identify the most suitable approach for predicting central line-associated bloodstream infection (CLABSI)…

Electronic Health Records (EHR) data analysis plays a crucial role in healthcare system quality. Because of its highly complex underlying causality and limited observable nature, causal inference on EHR is quite challenging. Deep Learning…

机器学习 · 计算机科学 2022-10-28 Jia Li , Haoyu Yang , Xiaowei Jia , Vipin Kumar , Michael Steinbach , Gyorgy Simon

Gathering relevant information to predict student academic progress is a tedious task. Due to the large amount of irrelevant data present in databases which provides inaccurate results. Currently, it is not possible to accurately measure…

机器学习 · 计算机科学 2021-09-02 Ali Jaber Almalki