中文
相关论文

相关论文: Parameter-Efficient VLMs for Gastrointestinal Endo…

200 篇论文

Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We translate the widely known Flickr30k dataset into Romanian…

计算与语言 · 计算机科学 2025-12-18 George-Andrei Dima , Dumitru-Clementin Cercel

Endoscopic procedures such as esophagogastroduodenoscopy (EGD) and colonoscopy play a critical role in diagnosing and managing gastrointestinal (GI) disorders. However, the documentation burden associated with these procedures place…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Evandros Kaklamanos , Kristjana Kristinsdottir , Jonathan Huang , Dustin Carlson , Rajesh Keswani , John Pandolfino , Mozziyar Etemadi

Parameter-efficient fine-tuning (PEFT) has emerged as a promising paradigm for adapting pretrained models under limited data conditions. However, most existing PEFT methods are designed for matrix-structured parameters and are not well…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Guanghua He , Hancan Zhu , Gaohang Yu , An Zhang

Collaborative machine learning across healthcare institutions promises improved diagnostic accuracy by leveraging diverse datasets, yet privacy regulations such as HIPAA prohibit direct patient data sharing. While federated learning (FL)…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Al Amin , Kamrul Hasan , Liang Hong , Sharif Ullah

We introduce Kvasir-VQA, an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations to facilitate advanced machine learning tasks in Gastrointestinal (GI) diagnostics.…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sushant Gautam , Andrea Storås , Cise Midoglu , Steven A. Hicks , Vajira Thambawita , Pål Halvorsen , Michael A. Riegler

The recent development of deep learning large models in medicine shows remarkable performance in medical image analysis and diagnosis, but their large number of parameters causes memory and inference latency challenges. Knowledge…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Shaojie Li , Zhaoshuo Diao

The lack, due to privacy concerns, of large public databases of medical pathologies is a well-known and major problem, substantially hindering the application of deep learning techniques in this field. In this article, we investigate the…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Andrea Asperti , Claudio Mastronardo

The rapid progress of large language models (LLMs) has transformed natural language processing, yet the challenge of efficient adaptation remains unresolved. Full fine-tuning achieves strong performance but imposes prohibitive computational…

量子物理 · 物理学 2025-09-23 Emily Jimin Roh , Hyojun Ahn , Samuel Yen-Chi Chen , Soohyun Park , Joongheon Kim

Colonoscopic polyp diagnosis is pivotal for early colorectal cancer detection, yet traditional automated reporting suffers from inconsistencies and hallucinations due to the scarcity of high-quality multimodal medical data. To bridge this…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Tianyu Zhou , Junyi Tang , Zehui Li , Dahong Qian , Suncheng Xiang

Solving tough clinical questions that require both image and text understanding is still a major challenge in healthcare AI. In this work, we propose Q-FSRU, a new model that combines Frequency Spectrum Representation and Fusion (FSRU) with…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Rakesh Thakur , Yusra Tariq , Rakesh Chandra Joshi

Visual Question-Answering (VQA) has become key to user experience, particularly after improved generalization capabilities of Vision-Language Models (VLMs). But evaluating VLMs for an application requirement using a standardized framework…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Neelabh Sinha , Vinija Jain , Aman Chadha

De-identification of medical images is a critical step to ensure privacy during data sharing in research and clinical settings. The initial step in this process involves detecting Protected Health Information (PHI), which can be found in…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Tuan Truong , Ivo M. Baltruschat , Mark Klemens , Grit Werner , Matthias Lenga

Reliable artificial intelligence (AI) models for medical image analysis often depend on large and diverse labeled datasets. Federated learning (FL) offers a decentralized and privacy-preserving approach to training but struggles in highly…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Mahshad Lotfinia , Arash Tayebiarasteh , Samaneh Samiei , Mehdi Joodaki , Soroosh Tayebi Arasteh

Accurate segmentation of organs and lesions in medical images is essential for clinical applications including diagnosis, prognosis, and treatment planning. While Vision Transformers (ViTs) have shown impressive segmentation performance,…

图像与视频处理 · 电气工程与系统科学 2026-05-13 Jin Yang , Xiaobing Yu , Peijie Qiu

Large-scale generative models like DeepSeek-R1 and OpenAI-O1 benefit substantially from chain-of-thought (CoT) reasoning, yet pushing their performance typically requires vast data, large model sizes, and full-parameter fine-tuning. While…

机器学习 · 计算机科学 2025-09-17 Yining Huang , Bin Li , Keke Tang , Meilian Chen

Frontier artificial intelligence (AI) models, such as OpenAI's GPT-5 and Meta's DINOv3, have advanced rapidly through training on internet-scale public data, yet such systems lack access to private clinical data. Neuroimaging, in…

Biomedical visual question answering (VQA) has been widely studied and has demonstrated significant application value and potential in fields such as assistive medical diagnosis. Despite their success, current biomedical VQA models perform…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Zhengyang Ji , Shang Gao , Li Liu , Yifan Jia , Yutao Yue

Although generative adversarial networks (GANs) have shown promise in medical imaging, they have four main limitations that impeded their utility: computational cost, data requirements, reliable evaluation measures, and training complexity.…

Background: Advances in artificial intelligence, particularly large language models (LLMs), have the potential to enhance technical expertise in magnetic resonance imaging (MRI), regardless of operator skill or geographic location. Methods:…

医学物理 · 物理学 2024-11-20 Alan B McMillan

The "pre-training then fine-tuning (FT)" paradigm is widely adopted to boost the model performance of deep learning-based methods for medical volumetric segmentation. However, conventional full FT incurs high computational and memory costs.…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Jiachen Shen , Wenxuan Wang , Chen Chen , Jianbo Jiao , Jing Liu , Yan Zhang , Shanshan Song , Jiangyun Li