中文
相关论文

相关论文: RadAnnotate: Large Language Models for Efficient a…

200 篇论文

Many machine learning systems today are trained on large amounts of human-annotated data. Data annotation tasks that require a high level of competency make data acquisition expensive, while the resulting labels are often subjective,…

机器学习 · 计算机科学 2020-04-08 Emmanouil Antonios Platanios , Maruan Al-Shedivat , Eric Xing , Tom Mitchell

Radiology reports summarize key findings and differential diagnoses derived from medical imaging examinations. The extraction of differential diagnoses is crucial for downstream tasks, including patient management and treatment planning.…

计算与语言 · 计算机科学 2024-10-15 Luoyao Chen , Revant Teotia , Antonio Verdone , Aidan Cardall , Lakshay Tyagi , Yiqiu Shen , Sumit Chopra

As Large Language Model (LLM) capabilities advance, the demand for high-quality annotation of exponentially increasing text corpora has outpaced human capacity, leading to the widespread adoption of LLMs in automatic evaluation and…

计算与语言 · 计算机科学 2026-04-02 Jiayu Wang , Junyoung Lee

When developing new large language models (LLMs), a key step is evaluating their final performance, often by computing the win-rate against a reference model based on external feedback. Human feedback is the gold standard, particularly for…

机器学习 · 计算机科学 2025-02-26 Zhaoyi Zhou , Yuda Song , Andrea Zanette

Financial analysis relies heavily on the interpretation of earnings reports to assess company performance and guide decision-making. Traditional methods for generating such analyzes require significant financial expertise and are often…

统计金融 · 定量金融 2025-12-10 Hai-Thien To , Tien-Cuong Bui , Van-Duc Le

Large language models (LLMs) offer strategy researchers powerful tools for annotating text at scale, but treating LLM-generated labels as deterministic overlooks substantial instability. Grounded in content analysis and generalizability…

计算机与社会 · 计算机科学 2026-01-21 Arnaldo Camuffo , Alfonso Gambardella , Saeid Kazemi , Jakub Malachowski , Abhinav Pandey

Remote sensing image segmentation is crucial for environmental monitoring, disaster assessment, and resource management, but its performance largely depends on the quality of the dataset. Although several high-quality datasets are broadly…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Jianhao Yang , Wenshuo Yu , Yuanchao Lv , Jiance Sun , Bokang Sun , Mingyang Liu

Accurate ground truth estimation in medical screening programs often relies on coalitions of experts and peer second opinions. Algorithms that efficiently aggregate noisy annotations can enhance screening workflows, particularly when data…

机器学习 · 计算机科学 2025-10-07 Tim Bary , Tiffanie Godelaine , Axel Abels , Benoît Macq

Effective document reranking is essential for improving search relevance across diverse applications. While Large Language Models (LLMs) excel at reranking due to their deep semantic understanding and reasoning, their high computational…

计算与语言 · 计算机科学 2025-10-03 Dimitar Peshevski , Kiril Blazhevski , Martin Popovski , Gjorgji Madjarov

Healthcare data suffers from both noise and lack of ground truth. The cost of data increases as it is cleaned and annotated in healthcare. Unlike other data sets, medical data annotation, which is critical to accurate ground truth, requires…

信号处理 · 电气工程与系统科学 2020-10-13 V. Ratna Saripalli , Gopal Avinash , Dibyajyoti Pati , Michael Potter , Charles W. Anderson

We systematically investigate lightweight strategies to adapt large language models (LLMs) for the task of radiology report summarization (RRS). Specifically, we focus on domain adaptation via pretraining (on natural language, biomedical…

An infrastructure for multisite, geographically-distributed creation and collection of diverse, high-quality, curated and labeled radiology image data is crucial for the successful automated development, deployment, monitoring and…

机器学习 · 计算机科学 2020-09-01 Menashe Benjamin , Guy Engelhard , Alex Aisen , Yinon Aradi , Elad Benjamin

Precision medicine has the potential to revolutionize healthcare, but much of the data for patients is locked away in unstructured free-text, limiting research and delivery of effective personalized treatments. Generating large annotated…

计算与语言 · 计算机科学 2020-12-16 Nick Altieri , Briton Park , Mara Olson , John DeNero , Anobel Odisho , Bin Yu

Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation…

计算机视觉与模式识别 · 计算机科学 2021-04-27 Yuan-Hong Liao , Amlan Kar , Sanja Fidler

Unstructured text data annotation is foundational to management research. LLMs offer a cost-effective and scalable alternative to human annotation, but they introduce a novel challenge: the annotator itself can be retired. Proprietary…

计算与语言 · 计算机科学 2026-05-13 Xiang Cheng , Raveesh Mayya , João Sedoc

In chest X-ray (CXR) image analysis, rule-based systems are usually employed to extract labels from reports for dataset releases. However, there is still room for improvement in label quality. These labelers typically output only presence…

图像与视频处理 · 电气工程与系统科学 2025-06-03 Ricardo Bigolin Lanfredi , Pritam Mukherjee , Ronald Summers

Background: Radiography (X-rays) is the dominant modality in orthopedics, and improving the interpretation of radiographs is clinically relevant. Machine learning (ML) has revolutionized data analysis and has been applied to medicine, with…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Jakub Olczak , Max Gordon

Large Language Models (LLMs) have demonstrated their efficacy across a broad spectrum of tasks in healthcare applications. However, often LLMs need to be fine-tuned on task-specific expert annotated data to achieve optimal performance,…

机器学习 · 计算机科学 2024-08-06 Bhawesh Kumar , Jonathan Amar , Eric Yang , Nan Li , Yugang Jia

Weak labeling is a popular weak supervision strategy for Named Entity Recognition (NER) tasks, with the goal of reducing the necessity for hand-crafted annotations. Although there are numerous remarkable annotation tools for NER labeling,…

计算与语言 · 计算机科学 2025-07-04 Mengyang Liu , Haozheng Luo , Leonard Thong , Yinghao Li , Chao Zhang , Le Song

Most natural language tasks in the radiology domain use language models pre-trained on biomedical corpus. There are few pretrained language models trained specifically for radiology, and fewer still that have been trained in a low data…

计算与语言 · 计算机科学 2023-06-06 Rikhiya Ghosh , Sanjeev Kumar Karn , Manuela Daniela Danu , Larisa Micu , Ramya Vunikili , Oladimeji Farri