English
Related papers

Related papers: PMOA-TTS: Introducing the PubMed Open Access Textu…

200 papers

Timing of clinical events is central to characterization of patient trajectories, enabling analyses such as process tracing, forecasting, and causal reasoning. However, structured electronic health records capture few data elements critical…

Computation and Language · Computer Science 2025-04-18 Jing Wang , Jeremy C Weiss

Clinical case reports and discharge summaries may be the most complete and accurate summarization of patient encounters, yet they are finalized, i.e., timestamped after the encounter. Complementary structured data streams become available…

Computation and Language · Computer Science 2026-05-13 Shahriar Noroozizadeh , Jeremy C. Weiss

Type 2 diabetes case reports describe complex clinical courses, but their timelines are often expressed in language that is difficult to reuse in longitudinal modeling. To address this gap, we developed a textual time-series corpus of 136…

Computation and Language · Computer Science 2026-04-09 Sayantan Kumar , Jeremy C. Weiss

Large language models (LLMs) have shown remarkable performance in vision-language tasks, but their application in the medical field remains underexplored, particularly for integrating structured time series data with unstructured clinical…

Computation and Language · Computer Science 2025-06-17 Shuai Niu , Jing Ma , Hongzhan Lin , Liang Bai , Zhihua Wang , Wei Bi , Yida Xu , Guo Li , Xian Yang

A crucial component for clinical risk prediction is developing a reliable prediction model is collecting high-quality time series clinical events. In this work, we release such a dataset that consists of 22,588,586 Clinical Time Series…

Artificial Intelligence · Computer Science 2025-11-19 Jing Wang , Xing Niu , Tong Zhang , Jie Shen , Juyong Kim , Jeremy C. Weiss

While many advances in time series models focus exclusively on numerical data, research on multimodal time series, particularly those involving contextual textual information, remains in its infancy. With recent progress in large language…

Machine Learning · Computer Science 2026-03-10 Zihao Li , Xiao Lin , Zhining Liu , Jiaru Zou , Ziwei Wu , Lecheng Zheng , Dongqi Fu , Yada Zhu , Hendrik Hamann , Hanghang Tong , Jingrui He

PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page image is annotated with Google Cloud Vision and released in a compact JSON schema with word-, line-, and paragraph-level…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Hunter Heidenreich , Yosheb Getachew , Olivia Dinica , Ben Elliott

We introduce Biomed-Enriched, a biomedical text dataset constructed from PubMed via a two-stage annotation process. In the first stage, a large language model annotates 400K paragraphs from PubMed scientific articles, assigning scores for…

Computation and Language · Computer Science 2025-06-26 Rian Touchent , Nathan Godey , Eric de la Clergerie

Timely detection of critical health conditions remains a major challenge in public health analytics, especially in Big Data environments characterized by high volume, rapid velocity, and diverse variety of clinical data. This study presents…

Databases · Computer Science 2025-10-14 Ritesh Chandra , Sonali Agarwal , Navjot Singh

Reconstructing precise clinical timelines is essential for modeling patient trajectories and forecasting risk in complex, heterogeneous conditions like sepsis. While unstructured clinical narratives offer semantically rich and contextually…

Computation and Language · Computer Science 2026-05-15 Sayantan Kumar , Shahriar Noroozizadeh , Juyong Kim , Jeremy C. Weiss

Vision-language models hold considerable promise for ophthalmology, but their development depends on large-scale, high-quality image-text datasets that remain scarce. We present PubMed-Ophtha, a hierarchical dataset of 102,023…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Verena Jasmin Hallitschke , Carsten Eickhoff , Philipp Berens

The time at which a message is communicated is a vital piece of metadata in many real-world natural language processing tasks such as Topic Detection and Tracking (TDT). TDT systems aim to cluster a corpus of news articles by event, and in…

Computation and Language · Computer Science 2024-03-27 Hang Jiang , Doug Beeferman , Weiquan Mao , Deb Roy

The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology…

Clinical case reports encode temporal patient trajectories that are often underexploited by traditional machine learning methods relying on structured data. In this work, we introduce the forecasting problem from textual time series, where…

Computation and Language · Computer Science 2025-12-30 Shahriar Noroozizadeh , Sayantan Kumar , Jeremy C. Weiss

PubTator 3.0 (https://www.ncbi.nlm.nih.gov/research/pubtator3/) is a biomedical literature resource using state-of-the-art AI techniques to offer semantic and relation searches for key concepts like proteins, genetic variants, diseases, and…

Computation and Language · Computer Science 2024-01-23 Chih-Hsuan Wei , Alexis Allot , Po-Ting Lai , Robert Leaman , Shubo Tian , Ling Luo , Qiao Jin , Zhizheng Wang , Qingyu Chen , Zhiyong Lu

Current forecasting approaches are largely unimodal and ignore the rich textual data that often accompany the time series due to lack of well-curated multimodal benchmark dataset. In this work, we develop TimeText Corpus (TTC), a carefully…

Artificial Intelligence · Computer Science 2024-11-22 Kai Kim , Howard Tsai , Rajat Sen , Abhimanyu Das , Zihao Zhou , Abhishek Tanpure , Mathew Luo , Rose Yu

Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and…

We introduce \textsc{ComplexTempQA},\footnote{Dataset and code available at: https://github.com/DataScienceUIBK/ComplexTempQA} a large-scale dataset consisting of over 100 million question-answer pairs designed to tackle the challenges in…

Computation and Language · Computer Science 2025-08-26 Raphael Gruber , Abdelrahman Abdallah , Michael Färber , Adam Jatowt

Embedding models group text by semantic content, what text is about. We show that temporal co-occurrence within texts discovers a different kind of structure: recurrent transition-structure concepts or what text does. We train a…

Artificial Intelligence · Computer Science 2026-03-20 Jason Dury

Foundation models trained on large-scale dataset gain a recent surge in CV and NLP. In contrast, development in biomedical domain lags far behind due to data scarcity. To address this issue, we build and release PMC-OA, a biomedical dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Weixiong Lin , Ziheng Zhao , Xiaoman Zhang , Chaoyi Wu , Ya Zhang , Yanfeng Wang , Weidi Xie
‹ Prev 1 2 3 10 Next ›