中文
相关论文

相关论文: LEMON: A Large Endoscopic MONocular Dataset and Fo…

200 篇论文

Medical image segmentation plays a pivotal role in disease diagnosis and treatment planning, particularly in resource-constrained clinical settings where lightweight and generalizable models are urgently needed. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Chengqi Dong , Fenghe Tang , Rongge Mao , Xinpei Gao , S. Kevin Zhou

Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yongcheng Yao , Yongshuo Zong , Raman Dutt , Yongxin Yang , Sotirios A Tsaftaris , Timothy Hospedales

We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality…

The growing popularity of robotic minimally invasive surgeries has made deep learning-based surgical training a key area of research. A thorough understanding of the surgical scene components is crucial, which semantic segmentation models…

图像与视频处理 · 电气工程与系统科学 2025-10-31 Muraam Abdel-Ghani , Mahmoud Ali , Mohamed Ali , Fatmaelzahraa Ahmed , Muhammad Arsalan , Abdulaziz Al-Ali , Shidin Balakrishnan

General Multimodal Large Language Models (MLLMs) often underperform in capturing domain-specific nuances in medical diagnosis, trailing behind fully supervised baselines. Although fine-tuning provides a remedy, the high costs of expert…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Wenkai Zhao , Zipei Wang , Mengjie Fang , Di Dong , Jie Tian , Lingwei Zhang

Despite the remarkable performance of deep models in medical imaging, they still require source data for training, which limits their potential in light of privacy concerns. Federated learning (FL), as a decentralized learning framework…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Yihang Wu , Ahmad Chaddad

Multimodal foundation models can process several modalities. However, since the space of possible modalities is large and evolving over time, training a model from scratch to encompass all modalities is unfeasible. Moreover, integrating a…

计算与语言 · 计算机科学 2025-09-08 Osman Batur İnce , André F. T. Martins , Oisin Mac Aodha , Edoardo M. Ponti

Clinical decision-making requires synthesizing heterogeneous evidence, including patient histories, clinical guidelines, and trajectories of comparable cases. While large language models (LLMs) offer strong reasoning capabilities, they…

人工智能 · 计算机科学 2026-03-03 Shuheng Chen , Namratha Patil , Haonan Pan , Angel Hsing-Chi Hwang , Yao Du , Ruishan Liu , Jieyu Zhao

While the Segment Anything Model (SAM) excels in semantic segmentation for general-purpose images, its performance significantly deteriorates when applied to medical images, primarily attributable to insufficient representation of medical…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Yiming Zhang , Tianang Leng , Kun Han , Xiaohui Xie

Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse…

图像与视频处理 · 电气工程与系统科学 2022-09-05 Yuanfeng Ji , Haotian Bai , Jie Yang , Chongjian Ge , Ye Zhu , Ruimao Zhang , Zhen Li , Lingyan Zhang , Wanling Ma , Xiang Wan , Ping Luo

Multimodal Deep Learning enhances decision-making by integrating diverse information sources, such as texts, images, audio, and videos. To develop trustworthy multimodal approaches, it is essential to understand how uncertainty impacts…

机器学习 · 计算机科学 2025-08-14 Grigor Bezirganyan , Sana Sellami , Laure Berti-Équille , Sébastien Fournier

In biomedical vision-language modeling, datasets are typically mined from scientific literature, pairing compound figures with captions that are short, context-dependent, and oftern partially informative. Prior work on subfigure extraction…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Negin Baghbanzadeh , Mohammed Saidul Islam , Sajad Ashkezari , Elham Dolatabadi , Arash Afkanpour

We present RASO, a foundation model designed to Recognize Any Surgical Object, offering robust open-set recognition capabilities across a broad range of surgical procedures and object classes, in both surgical images and videos. RASO…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Jiajie Li , Brian R Quaranto , Chenhui Xu , Ishan Mishra , Ruiyang Qin , Dancheng Liu , Peter C W Kim , Jinjun Xiong

The development of artificial intelligence systems for colonoscopy analysis often necessitates expert-annotated image datasets. However, limitations in dataset size and diversity impede model performance and generalisation. Image-text…

计算机视觉与模式识别 · 计算机科学 2023-10-18 Shuo Wang , Yan Zhu , Xiaoyuan Luo , Zhiwei Yang , Yizhe Zhang , Peiyao Fu , Manning Wang , Zhijian Song , Quanlin Li , Pinghong Zhou , Yike Guo

The deployment of foundation models for medical imaging has demonstrated considerable success. However, their training overheads associated with downstream tasks remain substantial due to the size of the image encoders employed, and the…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Chengxi Zeng , Yuxuan Jiang , Fan Zhang , Alberto Gambaruto , Tilo Burghardt

Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgical actions (BSA), the fundamental unit of operation in any…

Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Koki Maeda , Tosho Hirasawa , Atsushi Hashimoto , Jun Harashima , Leszek Rybicki , Yusuke Fukasawa , Yoshitaka Ushiku

Purpose: We propose a formal framework for the modeling and segmentation of minimally-invasive surgical tasks using a unified set of motion primitives (MPs) to enable more objective labeling and the aggregation of different datasets.…

机器人学 · 计算机科学 2023-05-16 Kay Hutchinson , Ian Reyes , Zongyu Li , Homa Alemzadeh

Colorectal cancer screening critically depends on colonoscopy, yet existing platforms offer limited support for systematically studying the coupled dynamics of operator control, instrument motion, and visual feedback. This gap restricts…

Continuum robots are advancing bronchoscopy procedures by accessing complex lung airways and enabling targeted interventions. However, their development is limited by the lack of realistic training and test environments: Real data is…