English
Related papers

Related papers: Multi-lingual Multi-institutional Electronic Healt…

200 papers

Hierarchical attention networks have recently achieved remarkable performance for document classification in a given language. However, when multilingual document collections are considered, training such models separately for each language…

Computation and Language · Computer Science 2017-09-18 Nikolaos Pappas , Andrei Popescu-Belis

Electronic Health Records (EHRs) enable deep learning for clinical predictions, but the optimal method for representing patient data remains unclear due to inconsistent evaluation practices. We present the first systematic benchmark to…

Machine Learning · Computer Science 2025-10-13 Tianyi Chen , Mingcheng Zhu , Zhiyao Luo , Tingting Zhu

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to…

Computation and Language · Computer Science 2025-10-22 Yingli Shen , Wen Lai , Shuo Wang , Ge Gao , Kangyang Luo , Alexander Fraser , Maosong Sun

The integration of multimodal Electronic Health Records (EHR) data has significantly improved clinical predictive capabilities. Leveraging clinical notes and multivariate time-series EHR, existing models often lack the medical context…

Artificial Intelligence · Computer Science 2024-02-13 Yinghao Zhu , Changyu Ren , Shiyun Xie , Shukai Liu , Hangyuan Ji , Zixiang Wang , Tao Sun , Long He , Zhoujun Li , Xi Zhu , Chengwei Pan

The inherent complexity of structured longitudinal Electronic Health Records (EHR) data poses a significant challenge when integrated with Large Language Models (LLMs), which are traditionally tailored for natural language processing.…

Computation and Language · Computer Science 2024-02-13 Yinghao Zhu , Zixiang Wang , Junyi Gao , Yuning Tong , Jingkun An , Weibin Liao , Ewen M. Harrison , Liantao Ma , Chengwei Pan

Multilingual proficiency presents a significant challenge for large language models (LLMs). English-centric models are usually suboptimal in other languages, particularly those that are linguistically distant from English. This performance…

Computation and Language · Computer Science 2025-01-07 Geyu Lin , Bin Wang , Zhengyuan Liu , Nancy F. Chen

Foundation models trained on patient electronic health records (EHRs) require tokenizing medical data into sequences of discrete vocabulary items. Existing tokenizers treat medical codes from EHRs as isolated textual tokens. However, each…

Computation and Language · Computer Science 2025-07-01 Xiaorui Su , Shvat Messica , Yepeng Huang , Ruth Johnson , Lukas Fesser , Shanghua Gao , Faryad Sahneh , Marinka Zitnik

Electronic health record (EHR) systems contain a wealth of multimodal clinical data including structured data like clinical codes and unstructured data such as clinical notes. However, many existing EHR-focused studies has traditionally…

Machine Learning · Statistics 2025-08-20 Tianxi Cai , Feiqing Huang , Ryumei Nakada , Linjun Zhang , Doudou Zhou

Multimodal language modeling has enabled breakthroughs for representation learning, yet remains unexplored in the realm of functional brain data for clinical phenotyping. This paper pioneers EEG-language models (ELMs) trained on clinical…

Signal Processing · Electrical Eng. & Systems 2025-08-12 Sam Gijsen , Kerstin Ritter

Open-source, multilingual medical large language models (LLMs) have the potential to serve linguistically diverse populations across different regions. Adapting generic LLMs for healthcare often requires continual pretraining, but this…

Computation and Language · Computer Science 2024-09-10 Meng Zhou , Surajsinh Parmar , Anubhav Bhatti

Ultrasound (US) report generation is a challenging task due to the variability of US images, operator dependence, and the need for standardized text. Unlike X-ray and CT, US imaging lacks consistent datasets, making automation difficult. In…

Image and Video Processing · Electrical Eng. & Systems 2025-05-20 Peixuan Ge , Tongkun Su , Faqin Lv , Baoliang Zhao , Peng Zhang , Chi Hong Wong , Liang Yao , Yu Sun , Zenan Wang , Pak Kin Wong , Ying Hu

Deep-learning-based clinical decision support using structured electronic health records (EHR) has been an active research area for predicting risks of mortality and diseases. Meanwhile, large amounts of narrative clinical notes provide…

Computation and Language · Computer Science 2023-05-10 Weimin Lyu , Xinyu Dong , Rachel Wong , Songzhu Zheng , Kayley Abell-Hart , Fusheng Wang , Chao Chen

Healthcare data are inherently multimodal, including electronic health records (EHR), medical images, and multi-omics data. Combining these multimodal data sources contributes to a better understanding of human health and provides optimal…

Machine Learning · Computer Science 2022-10-28 Farida Mohsen , Hazrat Ali , Nady El Hajj , Zubair Shah

Large language models (LLMs) hold promise for transforming healthcare, from streamlining administrative and clinical workflows to enriching patient engagement and advancing clinical decision-making. However, their successful integration…

Computers and Society · Computer Science 2025-04-04 Mohammed Al-Garadi , Tushar Mungle , Abdulaziz Ahmed , Abeed Sarker , Zhuqi Miao , Michael E. Matheny

Pre-trained Large Language Models (LLMs) often struggle on out-of-domain datasets like healthcare focused text. We explore specialized pre-training to adapt smaller LLMs to different healthcare datasets. Three methods are assessed:…

Computation and Language · Computer Science 2024-04-01 Niall Taylor , Dan Schofield , Andrey Kormilitzin , Dan W Joyce , Alejo Nevado-Holgado

Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL) poses a new challenge in finding high-quality training data…

Computation and Language · Computer Science 2022-10-25 Chen Qiu , Dan Oneata , Emanuele Bugliarello , Stella Frank , Desmond Elliott

While the ICD code assignment problem has been widely studied, most works have focused on post-discharge document classification. Models for early forecasting of this information could be used for identifying health risks, suggesting…

Machine Learning · Computer Science 2025-08-18 Cindy Shih-Ting Huang , Clarence Boon Liang Ng , Marek Rei

The lack of standardized evaluation benchmarks in the medical domain for text inputs can be a barrier to widely adopting and leveraging the potential of natural language models for health-related downstream tasks. This paper revisited an…

Computation and Language · Computer Science 2025-04-30 Jesus Lovon , Thouria Ben-Haddi , Jules Di Scala , Jose G. Moreno , Lynda Tamine

The development of Electronic Health Records summarization systems has revolutionized patient data management. Previous research advanced this field by adapting Large Language Models for clinical tasks, using diverse datasets to generate…

Computation and Language · Computer Science 2024-10-15 Ruvarashe Madzime , Clement Nyirenda

The scarcity of data presents a critical obstacle to the efficacy of medical visionlanguage pre-training (VLP). A potential solution lies in the combination of datasets from various language communities. Nevertheless, the main challenge…

Computation and Language · Computer Science 2024-02-20 Zhongwei Wan , Che Liu , Mi Zhang , Jie Fu , Benyou Wang , Sibo Cheng , Lei Ma , César Quilodrán-Casas , Rossella Arcucci