中文
相关论文

相关论文: MetaHQ: Harmonized, high-quality metadata annotati…

200 篇论文

Accessing sensitive patient data for machine learning is challenging due to privacy concerns. Datasets with annotations of personally identifiable information are crucial for developing and testing anonymization systems to enable safe data…

Large language models (LLMs) are increasingly used by researchers in the social sciences and humanities (SSH) for text analysis, particularly to automate text annotation. However, many researchers still face challenges in adopting LLMs,…

计算机与社会 · 计算机科学 2026-05-28 Qixiang Fang , Javier Garcia Bernardo , Erik-Jan van Kesteren

Foundation models have emerged as a powerful approach for processing electronic health records (EHRs), offering flexibility to handle diverse medical data modalities. In this study, we present a comprehensive benchmark that evaluates the…

机器学习 · 计算机科学 2025-07-22 Kunyu Yu , Rui Yang , Jingchi Liao , Siqi Li , Huitao Li , Irene Li , Yifan Peng , Rishikesan Kamaleswaran , Nan Liu

We present an analytical study of the quality of metadata about samples used in biomedical experiments. The metadata under analysis are stored in two well-known databases: BioSample---a repository managed by the National Center for…

数据库 · 计算机科学 2019-02-27 Rafael S. Gonçalves , Mark A. Musen

In the fast-moving world of AI, as organizations and researchers develop more advanced models, they face challenges due to their sheer size and computational demands. Deploying such models on edge devices or in resource-constrained…

机器学习 · 计算机科学 2025-09-23 Oussama Bouaggad , Natalia Grabar

Scientists increasingly recognize the importance of providing rich, standards-adherent metadata to describe their experimental results. Despite the availability of sophisticated tools to assist in the process of data annotation,…

Electronic health records (EHR) contain large volumes of unstructured text, requiring the application of Information Extraction (IE) technologies to enable clinical analysis. We present the open-source Medical Concept Annotation Toolkit…

Manually annotated data is key to developing text-mining and information-extraction algorithms. However, human annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature, it is paramount to…

人机交互 · 计算机科学 2020-04-27 Rezarta Islamaj , Dongseop Kwon , Sun Kim , Zhiyong Lu

Scientific metadata often suffer from incompleteness, inconsistency, and formatting errors, which hinder effective discovery and reuse of the associated datasets. We present a method that combines GPT-4 with structured metadata templates…

信息检索 · 计算机科学 2025-06-10 Sowmya S Sundaram , Rafael S. Gonçalves , Mark A Musen

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lacks the essential metadata required for researchers to find and search them effectively. The lack of metadata poses a significant…

Legacy scientific workflows, and the services within them, often present scarce and unstructured (i.e. textual) descriptions. This makes it difficult to find, share and reuse them, thus dramatically reducing their value to the community.…

信息检索 · 计算机科学 2014-07-02 Beatriz García-Jiménez , Mark D. Wilkinson

An increasing amount of research is being devoted to applying machine learning methods to electronic health record (EHR) data for various clinical purposes. This growing area of research has exposed the challenges of the accessibility of…

Framing the investigation of diverse cancers as a machine learning problem has recently shown significant potential in multi-omics analysis and cancer research. Empowering these successful machine learning models are the high-quality…

基因组学 · 定量生物学 2025-06-17 Ziwei Yang , Rikuto Kotoge , Xihao Piao , Zheng Chen , Lingwei Zhu , Peng Gao , Yasuko Matsubara , Yasushi Sakurai , Jimeng Sun

Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce…

Citation indexes are by now part of the research infrastructure in use by most scientists: a necessary tool in order to cope with the increasing amounts of scientific literature being published. Commercial citation indexes are designed for…

数字图书馆 · 计算机科学 2022-05-17 Giovanni Colavizza , Silvio Peroni , Matteo Romanello

The constant introduction of standardized benchmarks in the literature has helped accelerating the recent advances in meta-learning research. They offer a way to get a fair comparison between different algorithms, and the wide range of…

机器学习 · 计算机科学 2019-09-17 Tristan Deleu , Tobias Würfl , Mandana Samiei , Joseph Paul Cohen , Yoshua Bengio

Biomedical documents such as Electronic Health Records (EHRs) contain a large amount of information in an unstructured format. The data in EHRs is a hugely valuable resource documenting clinical narratives and decisions, but whilst the text…

Robust machine learning relies on access to data that can be used with standardized frameworks in important tasks and the ability to develop models whose performance can be reasonably reproduced. In machine learning for healthcare, the…

Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce CLINSQL, a benchmark of 633 expert-annotated tasks on MIMIC-IV…

计算与语言 · 计算机科学 2026-01-16 Yifei Shen , Yilun Zhao , Justice Ou , Tinglin Huang , Arman Cohan

Metadata quality is crucial for digital objects to be discovered through digital library interfaces. However, due to various reasons, the metadata of digital objects often exhibits incomplete, inconsistent, and incorrect values. We…