中文
相关论文

相关论文: Model Evaluation in Medical Datasets Over Time

200 篇论文

Machine learning is traditionally studied at the model level: researchers measure and improve the accuracy, robustness, bias, efficiency, and other dimensions of specific models. In practice, the societal impact of machine learning is…

机器学习 · 计算机科学 2024-04-04 Connor Toups , Rishi Bommasani , Kathleen A. Creel , Sarah H. Bana , Dan Jurafsky , Percy Liang

The development of electronic health records (EHR) systems has enabled the collection of a vast amount of digitized patient data. However, utilizing EHR data for predictive modeling presents several challenges due to its unique…

机器学习 · 计算机科学 2024-08-14 Jiaqi Wang , Junyu Luo , Muchao Ye , Xiaochen Wang , Yuan Zhong , Aofei Chang , Guanjie Huang , Ziyi Yin , Cao Xiao , Jimeng Sun , Fenglong Ma

Modern healthcare is ripe for disruption by AI. A game changer would be automatic understanding the latent processes from electronic medical records, which are being collected for billions of people worldwide. However, these healthcare…

神经与进化计算 · 计算机科学 2018-02-06 Phuoc Nguyen , Truyen Tran , Svetha Venkatesh

Medical imaging analysis has witnessed remarkable advancements even surpassing human-level performance in recent years, driven by the rapid development of advanced deep-learning algorithms. However, when the inference dataset slightly…

图像与视频处理 · 电气工程与系统科学 2024-10-11 Pratibha Kumari , Joohi Chauhan , Afshin Bozorgpour , Boqiang Huang , Reza Azad , Dorit Merhof

While the volume of electronic health records (EHR) data continues to grow, it remains rare for hospital systems to capture dense physiological data streams, even in the data-rich intensive care unit setting. Instead, typical EHR records…

机器学习 · 计算机科学 2018-12-04 Satya Narayan Shukla , Benjamin M. Marlin

We introduce LeetCodeDataset, a high-quality benchmark for evaluating and training code-generation models, addressing two key challenges in LLM research: the lack of reasoning-focused coding benchmarks and self-contained training testbeds.…

机器学习 · 计算机科学 2025-04-22 Yunhui Xia , Wei Shen , Yan Wang , Jason Klein Liu , Huifeng Sun , Siyue Wu , Jian Hu , Xiaolong Xu

While many machine learning methods have been used for medical prediction and risk factor analysis on healthcare data, most prior research has involved single-task learning (STL) methods. However, healthcare research often involves multiple…

机器学习 · 计算机科学 2021-03-08 Lu Wang , Haoyan Jiang , Mark Chignell

Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional data curation strategies that rely on coarse-grained partitioning by modality or…

计算与语言 · 计算机科学 2026-04-29 Jianghang Lin , Haihua Yang , Deli Yu , Kai Wu , Kai Ye , Jinghao Lin , Zihan Wang , Yuhang Wu , Liujuan Cao

Multiple sclerosis is a disease that affects the brain and spinal cord, it can lead to severe disability and has no known cure. The majority of prior work in machine learning for multiple sclerosis has been centered around using Magnetic…

Large language models (LLMs) have emerged as promising tools for assisting in medical tasks, yet processing Electronic Health Records (EHRs) presents unique challenges due to their longitudinal nature. While LLMs' capabilities to perform…

人工智能 · 计算机科学 2025-03-07 Hejie Cui , Alyssa Unell , Bowen Chen , Jason Alan Fries , Emily Alsentzer , Sanmi Koyejo , Nigam Shah

Continuous-time series is essential for different modern application areas, e.g. healthcare, automobile, energy, finance, Internet of things (IoT) and other related areas. Different application needs to process as well as analyse a massive…

机器学习 · 计算机科学 2024-09-17 Mansura Habiba , Barak A. Pearlmutter , Mehrdad Maleki

Irregularly sampled time series (ISTS) data has irregular temporal intervals between observations and different sampling rates between sequences. ISTS commonly appears in healthcare, economics, and geoscience. Especially in the medical…

机器学习 · 计算机科学 2020-10-27 Chenxi Sun , Shenda Hong , Moxian Song , Hongyan Li

Many large language models (LLMs) for medicine have largely been evaluated on short texts, and their ability to handle longer sequences such as a complete electronic health record (EHR) has not been systematically explored. Assessing these…

计算与语言 · 计算机科学 2023-11-17 Mihir Parmar , Aakanksha Naik , Himanshu Gupta , Disha Agrawal , Chitta Baral

Objective: Temporal electronic health records (EHRs) can be a wealth of information for secondary uses, such as clinical events prediction or chronic disease management. However, challenges exist for temporal data representation. We…

机器学习 · 计算机科学 2024-06-11 Feng Xie , Han Yuan , Yilin Ning , Marcus Eng Hock Ong , Mengling Feng , Wynne Hsu , Bibhas Chakraborty , Nan Liu

Performance of text classification models tends to drop over time due to changes in data, which limits the lifetime of a pretrained model. Therefore an ability to predict a model's ability to persist over time can help design models that…

计算与语言 · 计算机科学 2022-11-22 Rabab Alkhalifa , Elena Kochkina , Arkaitz Zubiaga

This study explores the potential of using training dynamics as an automated alternative to human annotation for evaluating the quality of training data. The framework used is Data Maps, which classifies data points into categories such as…

机器学习 · 计算机科学 2024-11-05 Laura Wenderoth

Individual-level human mobility prediction has emerged as a significant topic of research with applications in infectious disease monitoring, child, and elderly care. Existing studies predominantly focus on the microscopic aspects of human…

机器学习 · 计算机科学 2025-08-20 Yueyang Liu , Lance Kennedy , Ruochen Kong , Joon-Seok Kim , Andreas Züfle

Predictive models for clinical outcomes that are accurate on average in a patient population may underperform drastically for some subpopulations, potentially introducing or reinforcing inequities in care access and quality. Model training…

机器学习 · 统计学 2022-02-03 Stephen R. Pfohl , Haoran Zhang , Yizhe Xu , Agata Foryciarz , Marzyeh Ghassemi , Nigam H. Shah

While developments in machine learning led to impressive performance gains on big data, many human subjects data are, in actuality, small and sparsely labeled. Existing methods applied to such data often do not easily generalize to…

机器学习 · 计算机科学 2023-04-04 Julie Jiang , Kristina Lerman , Emilio Ferrara

Language models (LMs) represent an emerging paradigm within artificial intelligence, with applications throughout the medical enterprise. A comprehensive understanding of the clinical task and awareness of the variability in performance…

机器学习 · 计算机科学 2026-03-09 Victor Garcia , Mariia Sidulova , Aldo Badano