中文
相关论文

相关论文: FreeVPS: Repurposing Training-Free SAM2 for Genera…

200 篇论文

Detection of colon polyps has become a trending topic in the intersecting fields of machine learning and gastrointestinal endoscopy. The focus has mainly been on per-frame classification. More recently, polyp segmentation has gained…

图像与视频处理 · 电气工程与系统科学 2021-07-02 Vajira Thambawita , Steven A. Hicks , Pål Halvorsen , Michael A. Riegler

Semantic Segmentation combines two sub-tasks: the identification of pixel-level image masks and the application of semantic labels to those masks. Recently, so-called Foundation Models have been introduced; general models trained on very…

计算机视觉与模式识别 · 计算机科学 2023-10-03 David Balaban , Justin Medich , Pranay Gosar , Justin Hart

Self-supervised video denoising methods typically extend image-based frameworks into the temporal dimension, yet they often struggle to integrate inter-frame temporal consistency with intra-frame spatial specificity. Existing Video…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Mingjie Ji , Zhan Shi , Kailai Zhou , Zixuan Fu , Xun Cao

Interactive medical image segmentation (IMIS) has shown significant potential in enhancing segmentation accuracy by integrating iterative feedback from medical professionals. However, the limited availability of enough 3D medical data…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Chuyun Shen , Wenhao Li , Yuhang Shi , Xiangfeng Wang

Purpose: Foundation models, trained on multitudes of public datasets, often require additional fine-tuning or re-prompting mechanisms to be applied to visually distinct target domains such as surgical videos. Further, without domain…

图像与视频处理 · 电气工程与系统科学 2025-07-02 Ssharvien Kumar Sivakumar , Yannik Frisch , Amin Ranem , Anirban Mukhopadhyay

Background: We evaluate SAM 2 for surgical scene understanding by examining its semantic segmentation capabilities for organs/tissues both in zero-shot scenarios and after fine-tuning. Methods: We utilized five public datasets to evaluate…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Devanish N. Kamtam , Joseph B. Shrager , Satya Deepya Malla , Xiaohan Wang , Nicole Lin , Juan J. Cardona , Serena Yeung-Levy , Clarence Hu

Early identification of a polyp in the lower gastrointestinal (GI) tract can lead to prevention of life-threatening colorectal cancer. Developing computer-aided diagnosis (CAD) systems to detect polyps can improve detection accuracy and…

图像与视频处理 · 电气工程与系统科学 2022-06-01 Jan Andre Fagereng , Vajira Thambawita , Andrea M. Storås , Sravanthi Parasa , Thomas de Lange , Pål Halvorsen , Michael A. Riegler

The recent wave of foundation models has witnessed tremendous success in computer vision (CV) and beyond, with the segment anything model (SAM) having sparked a passion for exploring task-agnostic visual foundation models. Empowered by its…

计算机视觉与模式识别 · 计算机科学 2024-08-19 Chunhui Zhang , Yawen Cui , Weilin Lin , Guanjie Huang , Yan Rong , Li Liu , Shiguang Shan

SAM2 produces high-quality zero-shot segmentation on natural images, but applying it to large remote sensing scenes exposes two problems: (1) its mask generator faces an inherent quality-coverage trade-off: strict thresholds yield precise…

Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant…

图像与视频处理 · 电气工程与系统科学 2025-03-28 Syed Ariff Syed Hesham , Yun Liu , Guolei Sun , Henghui Ding , Jing Yang , Ender Konukoglu , Xue Geng , Xudong Jiang

Colorectal cancer is a one of the highest causes of cancer-related death, especially in men. Polyps are one of the main causes of colorectal cancer and early diagnosis of polyps by colonoscopy could result in successful treatment. Diagnosis…

图像与视频处理 · 电气工程与系统科学 2018-02-02 Mojtaba Akbari , Majid Mohrekesh , Ebrahim Nasr-Esfahani , S. M. Reza Soroushmehr , Nader Karimi , Shadrokh Samavi , Kayvan Najarian

Convolutional neural network (CNN) and Transformer-based architectures are two dominant deep learning models for polyp segmentation. However, CNNs have limited capability for modeling long-range dependencies, while Transformers incur…

图像与视频处理 · 电气工程与系统科学 2025-05-12 Diego Adame , Jose A. Nunez , Fabian Vazquez , Nayeli Gurrola , Huimin Li , Haoteng Tang , Bin Fu , Pengfei Gu

Visual Object Tracking (VOT) is widely used in applications like autonomous driving to continuously track targets in videos. Existing methods can be roughly categorized into template matching and autoregressive methods, where the former…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Qianxiong Xu , Lanyun Zhu , Chenxi Liu , Guosheng Lin , Cheng Long , Ziyue Li , Rui Zhao

Colonoscopy is the most widely used medical technique for preventing Colorectal Cancer, by detecting and removing polyps before they become malignant. Recent studies show that around one quarter of the existing polyps are routinely missed.…

计算机视觉与模式识别 · 计算机科学 2024-03-14 G. Leifman , I. Kligvasser , R. Goldenberg , M. Elad , E. Rivlin

We introduce SAM2Point, a preliminary exploration adapting Segment Anything Model 2 (SAM 2) for zero-shot and promptable 3D segmentation. SAM2Point interprets any 3D data as a series of multi-directional videos, and leverages SAM 2 for…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Ziyu Guo , Renrui Zhang , Xiangyang Zhu , Chengzhuo Tong , Peng Gao , Chunyuan Li , Pheng-Ann Heng

Breast MRI provides high-resolution volumetric imaging critical for tumor assessment and treatment planning, yet manual interpretation of 3D scans remains labor-intensive and subjective. While AI-powered tools hold promise for accelerating…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Solha Kang , Eugene Kim , Joris Vankerschaver , Utku Ozbulak

This paper investigates the fundamental discontinuity between the latest two Segment Anything Models: SAM2 and SAM3. We explain why the expertise in prompt-based segmentation of SAM2 does not transfer to the multimodal concept-driven…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Ranjan Sapkota , Konstantinos I. Roumeliotis , Manoj Karkee

Semantic segmentation in videos has been a focal point of recent research. However, existing models encounter challenges when faced with unfamiliar categories. To address this, we introduce the Open Vocabulary Video Semantic Segmentation…

多媒体 · 计算机科学 2024-12-13 Xinhao Li , Yun Liu , Guolei Sun , Min Wu , Le Zhang , Ce Zhu

Polyp detectors trained on clean datasets often underperform in real-world endoscopy, where illumination changes, motion blur, and occlusions degrade image quality. Existing approaches struggle with the domain gap between controlled…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Shengkai Xu , Hsiang Lun Kao , Tianxiang Xu , Honghui Zhang , Junqiao Wang , Runmeng Ding , Guanyu Liu , Tianyu Shi , Zhenyu Yu , Guofeng Pan , Ziqian Bi , Yuqi Ouyang

This paper presents MirrorSAM2, the first framework that adapts Segment Anything Model 2 (SAM2) to the task of RGB-D video mirror segmentation. MirrorSAM2 addresses key challenges in mirror detection, such as reflection ambiguity and…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Mingchen Xu , Yukun Lai , Ze Ji , Jing Wu