中文
相关论文

相关论文: MetaHQ: Harmonized, high-quality metadata annotati…

200 篇论文

The ability of large language models (LLMs) to follow natural language instructions with human-level fluency suggests many opportunities in healthcare to reduce administrative burden and improve quality of care. However, evaluating LLMs on…

This paper introduces a human-in-the-loop (HITL) data annotation pipeline to generate high-quality, large-scale speech datasets. The pipeline combines human and machine advantages to more quickly, accurately, and cost-effectively annotate…

音频与语音处理 · 电气工程与系统科学 2021-09-06 Mingkuan Liu , Chi Zhang , Hua Xing , Chao Feng , Monchu Chen , Judith Bishop , Grace Ngapo

Fine-tuning large language models (LLMs) to align with user preferences is challenging due to the high cost of quality human annotations in Reinforcement Learning from Human Feedback (RLHF) and the generalizability limitations of AI…

Electronic Health Record (EHR) retrieval plays a pivotal role in various clinical tasks, but its development has been severely impeded by the lack of publicly available benchmarks. In this paper, we introduce a novel public EHR retrieval…

信息检索 · 计算机科学 2025-04-09 Zhengyun Zhao , Hongyi Yuan , Jingjing Liu , Haichao Chen , Huaiyuan Ying , Songchi Zhou , Yue Zhong , Sheng Yu

Combining clinical and omics data can improve both daily clinical routines and research to gain more insights into complex medical procedures. We present the results of our first phase in a multi-year collaboration with analysts and…

人机交互 · 计算机科学 2023-09-25 Markus Höhn , Hendrik Lücke-Tieke , Jan Burmeister , Jörn Kohlhammer

Given the increasing complexity of omics datasets, a key challenge is not only improving classification performance but also enhancing the transparency and reliability of model decisions. Effective model performance and feature selection…

Medical Subject Heading (MeSH) indexing refers to the problem of assigning a given biomedical document with the most relevant labels from an extremely large set of MeSH terms. Currently, the vast number of biomedical articles in the PubMed…

计算与语言 · 计算机科学 2022-04-29 Xindi Wang , Robert E. Mercer , Frank Rudzicz

Fine-grained, span-level human evaluation has emerged as a reliable and robust method for evaluating text generation tasks such as summarization, simplification, machine translation and news generation, and the derived annotations have been…

计算与语言 · 计算机科学 2023-10-17 David Heineman , Yao Dou , Wei Xu

The extraction of phenotype information which is naturally contained in electronic health records (EHRs) has been found to be useful in various clinical informatics applications such as disease diagnosis. However, due to imprecise…

计算与语言 · 计算机科学 2019-11-12 Jingqing Zhang , Xiaoyu Zhang , Kai Sun , Xian Yang , Chengliang Dai , Yike Guo

Purpose. The increasing emphasis on data quantity in research infrastructures has highlighted the need for equally robust mechanisms ensuring data quality, particularly in bibliographic and citation datasets. This paper addresses the…

数字图书馆 · 计算机科学 2025-04-17 Ivan Heibi , Silvio Peroni , Elia Rizzetto

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often not generalizable, based on synthetic data, or not publicly…

数字图书馆 · 计算机科学 2026-03-27 Parth Sarin , Juan Pablo Alperin , Adam Buttrick , Dione Mentis

High computation costs and latency of large language models such as GPT-4 have limited their deployment in clinical settings. Small language models (SLMs) offer a cost-effective alternative, but their limited capacity requires biomedical…

The fast and affordable sequencing of large clinical and environmental metagenomic datasets opens up new horizons in medical and biotechnological applications. It is believed that today we have described only about 1\% of the microorganisms…

基因组学 · 定量生物学 2017-10-31 Balazs Szalkai , Vince Grolmusz

In the evolving landscape of clinical informatics, the integration and utilization of software tools developed through governmental funding represent a pivotal advancement in research and application. However, the dispersion of these tools…

数字图书馆 · 计算机科学 2024-03-28 Jeremy R. Harper

Motivation: Ontologies are widely used in biology for data annotation, integration, and analysis. In addition to formally structured axioms, ontologies contain meta-data in the form of annotation axioms which provide valuable pieces of…

计算与语言 · 计算机科学 2018-05-01 Fatima Zohra Smaili , Xin Gao , Robert Hoehndorf

Helix is an open-source, extensible, Python-based software framework to facilitate reproducible and interpretable machine learning workflows for tabular data. It addresses the growing need for transparent experimental data analytics…

The HuggingFace Datasets Hub hosts thousands of datasets, offering exciting opportunities for language model training and evaluation. However, datasets for a specific task type often have different schemas, making harmonization challenging.…

计算与语言 · 计算机科学 2023-05-17 Damien Sileo

A variety of schemas and ontologies are currently used for the machine-readable description of bibliographic entities and citations. This diversity, and the reuse of the same ontology terms with different nuances, generates inconsistencies…

This paper presents an approach for metadata reconciliation, curation and linking for Open Governamental Data Portals (ODPs). ODPs have been lately the standard solution for governments willing to put their public data available for the…

信息检索 · 计算机科学 2015-10-16 Alan Tygel , Sören Auer , Jeremy Debattista , Fabrizio Orlandi , Maria Luiza Machado Campos

Medical imaging papers often focus on methodology, but the quality of the algorithms and the validity of the conclusions are highly dependent on the datasets used. As creating datasets requires a lot of effort, researchers often use…