中文
相关论文

相关论文: Scaling up self-supervised learning for improved s…

200 篇论文

Self-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data structure or context by solving a pretext task.…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Markus Marks , Manuel Knott , Neehar Kondapaneni , Elijah Cole , Thijs Defraeye , Fernando Perez-Cruz , Pietro Perona

Foundation models have recently emerged as powerful feature extractors in computational pathology, yet they typically omit mechanisms for leveraging the global spatial structure of tissues and the local contextual relationships among…

Objective: To enable context-aware computer assistance in the operating room of the future, cognitive systems need to understand automatically which surgical phase is being performed by the medical team. The primary source of information…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Isabel Funke , Dominik Rivoir , Stefanie Krell , Stefanie Speidel

Foundation models (FMs) are transforming computational pathology by offering new ways to analyze histopathology images. However, FMs typically require weeks of training on large databases, making their creation a resource-intensive process.…

图像与视频处理 · 电气工程与系统科学 2026-01-27 Till Nicke , Daniela Schacherer , Jan Raphael Schäfer , Natalia Artysh , Antje Prasse , André Homeyer , Andrea Schenk , Henning Höfener , Johannes Lotz

Surgical phase recognition from video is a technology that automatically classifies the progress of a surgical procedure and has a wide range of potential applications, including real-time surgical support, optimization of medical…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Satoshi Kondo

Retinal blood vessel segmentation can extract clinically relevant information from fundus images. As manual tracing is cumbersome, algorithms based on Convolution Neural Networks have been developed. Such studies have used small publicly…

图像与视频处理 · 电气工程与系统科学 2024-06-24 Jeremiah Fadugba , Patrick Köhler , Lisa Koch , Petru Manescu , Philipp Berens

Recent success of vision foundation models have shown promising performance for the 2D perception tasks. However, it is difficult to train a 3D foundation network directly due to the limited dataset and it remains under explored whether…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Qingdong He , Jinlong Peng , Zhengkai Jiang , Xiaobin Hu , Jiangning Zhang

The rapid growth of medical imaging has fueled the development of Foundation Models (FMs) to reduce the growing, unsustainable workload on radiologists. While recent FMs have shown the power of large-scale pre-training to CT and MRI…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Antoine Saporta , Baptiste Callard , Corentin Dancette , Julien Khlaut , Charles Corbière , Leo Butsanets , Amaury Prat , Pierre Manceron

Foundation models have achieved remarkable success across video, image, and language domains. By scaling up the number of parameters and training datasets, these models acquire generalizable world knowledge and often surpass task-specific…

机器学习 · 计算机科学 2025-07-16 Tung Nguyen , Arsh Koneru , Shufan Li , Aditya Grover

Foundation models, pre-trained on massive datasets, have achieved unprecedented generalizability. However, is it truly necessary to involve such vast amounts of data in pre-training, consuming extensive computational resources? This paper…

机器学习 · 计算机科学 2024-08-19 Wenxuan Yang , Weimin Tan , Yuqi Sun , Bo Yan

Image-text training like CLIP has dominated the pretraining of vision foundation models in recent years. Subsequent efforts have been made to introduce region-level visual learning into CLIP's pretraining but face scalability challenges due…

计算机视觉与模式识别 · 计算机科学 2024-04-12 Xiaohu Jiang , Yixiao Ge , Yuying Ge , Dachuan Shi , Chun Yuan , Ying Shan

Artificial intelligence (AI) has shown promise in detecting and characterizing musculoskeletal diseases from radiographs. However, most existing models remain task-specific, annotation-dependent, and limited in generalizability across…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Shinn Kim , Soobin Lee , Kyoungseob Shin , Han-Soo Kim , Yongsung Kim , Minsu Kim , Juhong Nam , Somang Ko , Daeheon Kwon , Wook Huh , Ilkyu Han , Sunghoon Kwon

In minimally invasive surgery, clinical decisions depend on real-time visual interpretation, yet intraoperative perception varies substantially across surgeons and procedures. This variability limits consistent assessment, training, and the…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Kanggil Park , Yongjun Jeon , Soyoung Lim , Seonmin Park , Jongmin Shin , Jung Yong Kim , Sehyeon An , Jinsoo Rhu , Jongman Kim , Gyu-Seong Choi , Namkee Oh , Kyu-Hwan Jung

The Vision-Language Foundation model is increasingly investigated in the fields of computer vision and natural language processing, yet its exploration in ophthalmology and broader medical applications remains limited. The challenge is the…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Jiawei Du , Jia Guo , Weihang Zhang , Shengzhu Yang , Hanruo Liu , Huiqi Li , Ningli Wang

Self-Supervised Learning (SSL) is at the core of training modern large machine learning models, providing a scheme for learning powerful representations that can be used in a variety of downstream tasks. However, SSL strategies must be…

高能物理 - 唯象学 · 物理学 2025-02-26 Philip Harris , Michael Kagan , Jeffrey Krupa , Benedikt Maier , Nathaniel Woodward

Large-scale models pre-trained on large-scale datasets have profoundly advanced the development of deep learning. However, the state-of-the-art models for medical image segmentation are still small-scale, with their parameters only in the…

计算机视觉与模式识别 · 计算机科学 2023-04-14 Ziyan Huang , Haoyu Wang , Zhongying Deng , Jin Ye , Yanzhou Su , Hui Sun , Junjun He , Yun Gu , Lixu Gu , Shaoting Zhang , Yu Qiao

Deep learning (DL) has achieved remarkable progress in the field of medical imaging. However, adapting DL models to medical tasks remains a significant challenge, primarily due to two key factors: (1) architecture selection, as different…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Lotfi Abdelkrim Mecharbat , Ibrahim Almakky , Martin Takac , Mohammad Yaqub

Recent advancements for large-scale pre-training with neural signals such as electroencephalogram (EEG) have shown promising results, significantly boosting the development of brain-computer interfaces (BCIs) and healthcare. However, these…

信号处理 · 电气工程与系统科学 2025-03-21 Wei-Bang Jiang , Yansen Wang , Bao-Liang Lu , Dongsheng Li

The integration of deep learning based systems in clinical practice is often impeded by challenges rooted in limited and heterogeneous medical datasets. In addition, the field has increasingly prioritized marginal performance gains on a…

图像与视频处理 · 电气工程与系统科学 2025-03-18 Sebastian Doerrich , Francesco Di Salvo , Julius Brockmann , Christian Ledig

Automated segmentation is a fundamental medical image analysis task, which enjoys significant advances due to the advent of deep learning. While foundation models have been useful in natural language processing and some vision tasks for…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Hanxue Gu , Haoyu Dong , Jichen Yang , Maciej A. Mazurowski