English
Related papers

Related papers: Foundation Model for Endoscopy Video Analysis via …

200 papers

Electroencephalography (EEG) reflects the brain's functional state, making it a crucial tool for diverse detection applications like seizure detection and sleep stage classification. While deep learning-based approaches have recently shown…

Machine Learning · Computer Science 2025-10-07 Kerui Wu , Ziyue Zhao , Bülent Yener

Recent advancements in foundation models, typically trained with self-supervised learning on large-scale and diverse datasets, have shown great potential in medical image analysis. However, due to the significant spatial heterogeneity of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Lingxiao Luo , Xuanzhong Chen , Bingda Tang , Xinsheng Chen , Rong Han , Chengpeng Hu , Yujiang Li , Ting Chen

Electroencephalography (EEG) has emerged as a cost-effective and efficient tool to support neurologists in the detection of Alzheimer's Disease (AD). However, most existing approaches rely heavily on manual feature engineering or data…

Signal Processing · Electrical Eng. & Systems 2025-08-05 Yihe Wang , Nadia Mammone , Darina Petrovsky , Alexandros T. Tzallas , Francesco C. Morabito , Xiang Zhang

Vision foundation models have demonstrated exceptional generalization capabilities in segmentation tasks for both generic and specialized images. However, a performance gap persists between foundation models and task-specific, specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Chengxi Zeng , David Smithard , Alberto M Gambaruto , Tilo Burghardt

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zhongying Deng , Cheng Tang , Ziyan Huang , Jiashi Lin , Ying Chen , Junzhi Ning , Chenglong Ma , Jiyao Liu , Wei Li , Yinghao Zhu , Shujian Gao , Yanyan Huang , Sibo Ju , Yanzhou Su , Pengcheng Chen , Wenhao Tang , Tianbin Li , Haoyu Wang , Yuanfeng Ji , Hui Sun , Shaobo Min , Liang Peng , Feilong Tang , Haochen Xue , Rulin Zhou , Chaoyang Zhang , Wenjie Li , Shaohao Rui , Weijie Ma , Xingyue Zhao , Yibin Wang , Kun Yuan , Zhaohui Lu , Shujun Wang , Jinjie Wei , Lihao Liu , Dingkang Yang , Lin Wang , Yulong Li , Haolin Yang , Yiqing Shen , Lequan Yu , Xiaowei Hu , Yun Gu , Yicheng Wu , Benyou Wang , Minghui Zhang , Angelica I. Aviles-Rivero , Qi Gao , Hongming Shan , Xiaoyu Ren , Fang Yan , Hongyu Zhou , Haodong Duan , Maosong Cao , Shanshan Wang , Bin Fu , Xiaomeng Li , Zhi Hou , Chunfeng Song , Lei Bai , Yuan Cheng , Yuandong Pu , Xiang Li , Wenhai Wang , Hao Chen , Jiaxin Zhuang , Songyang Zhang , Huiguang He , Mengzhang Li , Bohan Zhuang , Zhian Bai , Rongshan Yu , Liansheng Wang , Yukun Zhou , Xiaosong Wang , Xin Guo , Guanbin Li , Xiangru Lin , Dakai Jin , Mianxin Liu , Wenlong Zhang , Qi Qin , Conghui He , Yuqiang Li , Ye Luo , Nanqing Dong , Jie Xu , Wenqi Shao , Bo Zhang , Qiujuan Yan , Yihao Liu , Jun Ma , Zhi Lu , Yuewen Cao , Zongwei Zhou , Jianming Liang , Shixiang Tang , Qi Duan , Dongzhan Zhou , Chen Jiang , Yuyin Zhou , Yanwu Xu , Jiancheng Yang , Shaoting Zhang , Xiaohong Liu , Siqi Luo , Yi Xin , Chaoyu Liu , Haochen Wen , Xin Chen , Alejandro Lozano , Min Woo Sun , Yuhui Zhang , Yue Yao , Xiaoxiao Sun , Serena Yeung-Levy , Xia Li , Jing Ke , Chunhui Zhang , Zongyuan Ge , Ming Hu , Jin Ye , Zhifeng Li , Yirong Chen , Yu Qiao , Junjun He

Video is an essential imaging modality for diagnostics, e.g. in ultrasound imaging, for endoscopy, or movement assessment. However, video hasn't received a lot of attention in the medical image analysis community. In the clinical practice,…

Computer Vision and Pattern Recognition · Computer Science 2020-05-20 Tianrui Liu , Qingjie Meng , Athanasios Vlontzos , Jeremy Tan , Daniel Rueckert , Bernhard Kainz

While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge when training on 3D-CT volumetric data. In this study, we propose TotalFM, a radiological…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Kohei Yamamoto , Tomohiro Kikuchi

Foundation models (FMs), large neural networks pretrained on extensive and diverse datasets, have revolutionized artificial intelligence and shown significant promise in medical imaging by enabling robust performance with limited labeled…

Image and Video Processing · Electrical Eng. & Systems 2025-06-17 Salah Ghamizi , Georgia Kanli , Yu Deng , Magali Perquin , Olivier Keunen

This work presents a novel approach for the early recognition of the type of a laparoscopic surgery from its video. Early recognition algorithms can be beneficial to the development of 'smart' OR systems that can provide automatic…

Computer Vision and Pattern Recognition · Computer Science 2019-09-06 Siddharth Kannan , Gaurav Yengera , Didier Mutter , Jacques Marescaux , Nicolas Padoy

Video matting aims to predict the alpha mattes for each frame from a given input video sequence. Recent solutions to video matting have been dominated by deep convolutional neural networks (CNN) for the past few years, which have become the…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Jiachen Li , Vidit Goel , Marianna Ohanyan , Shant Navasardyan , Yunchao Wei , Humphrey Shi

Endoscopy is essential in medical imaging, used for diagnosis, prognosis and treatment. Developing a robust dynamic 3D reconstruction pipeline for endoscopic videos could enhance visualization, improve diagnostic accuracy, aid in treatment…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Laura Salort-Benejam , Antonio Agudo

We present a methodology for training foundational transformer models capable of processing collider data with diverse kinematic signatures. Our universal foundation model is designed for simultaneous analysis of all processes involving…

High Energy Physics - Phenomenology · Physics 2025-11-13 E. Abasov , L. Dudko , E. Iudin , A. Markina , P. Volkov , M. Perfilov , A. Zaborenko

Training Deep Neural Networks for tracking individual cells in biomedical videos requires a large amount of annotated data. The annotation of videos for cell tracking is very time consuming and often requires domain expertise; this explains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Francesco Benedetto , Roberto Basla , Luca Magri , Giacomo Boracchi

Endotracheal suctioning (ES) is an invasive yet essential clinical procedure that requires a high degree of skill to minimize patient risk - particularly in home care and educational settings, where consistent supervision may be limited.…

Artificial Intelligence · Computer Science 2026-01-30 Hoang Khang Phan , Quang Vinh Dang , Noriyo Colley , Christina Garcia , Nhat Tan Le

Electroencephalography (EEG) signals provide critical insights for applications in disease diagnosis and healthcare. However, the scarcity of labeled EEG data poses a significant challenge. Foundation models offer a promising solution by…

Machine Learning · Computer Science 2025-02-25 Limin Wang , Toyotaro Suzumura , Hiroki Kanezashi

The demand for high-resolution subsurface imaging and continuous Earth monitoring has driven rapid growth in active and passive seismic data from dense geophone deployments, distributed acoustic sensing (DAS) arrays, and large-scale 2D and…

Geophysics · Physics 2026-05-13 Jiahua Zhao , Umair bin Waheed , Jing Sun , Yang Cui , Nikos Savva , Eric Verschuur

Foundation models trained on web-scale data have revolutionized robotics, but their application to low-level control remains largely limited to behavioral cloning. Drawing inspiration from the success of the reinforcement learning stage in…

Machine Learning · Computer Science 2025-09-19 Seyed Kamyar Seyed Ghasemipour , Ayzaan Wahid , Jonathan Tompson , Pannag Sanketi , Igor Mordatch

Vision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, these models cannot…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Raviteja Vemulapalli , Hadi Pouransari , Fartash Faghri , Sachin Mehta , Mehrdad Farajtabar , Mohammad Rastegari , Oncel Tuzel

This work presents EndoStreamDepth, a monocular depth estimation framework for endoscopic video streams. It provides accurate depth maps with sharp anatomical boundaries for each frame, temporally consistent predictions across frames, and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Hao Li , Daiwei Lu , Jiacheng Wang , Robert J. Webster , Ipek Oguz

We present DINO-world, a powerful generalist video world model trained to predict future frames in the latent space of DINOv2. By leveraging a pre-trained image encoder and training a future predictor on a large-scale uncurated video…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Federico Baldassarre , Marc Szafraniec , Basile Terver , Vasil Khalidov , Francisco Massa , Yann LeCun , Patrick Labatut , Maximilian Seitzer , Piotr Bojanowski
‹ Prev 1 8 9 10 Next ›