中文
相关论文

相关论文: ViTALS: Vision Transformer for Action Localization…

200 篇论文

Urinary bladder cancer surveillance requires tracking tumor sites across repeated interventions, yet the deformable and hollow bladder lacks stable landmarks for orientation. While blood vessels visible during endoscopy offer a…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Franziska Krauß , Matthias Ege , Zoltan Lovasz , Albrecht Bartz-Schmidt , Igor Tsaur , Oliver Sawodny , Carina Veil

This paper tackles the challenge of point-supervised temporal action detection, wherein only a single frame is annotated for each action instance in the training set. Most of the current methods, hindered by the sparse nature of annotated…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Elahe Vahdani , Yingli Tian

This paper presents a novel framework for processing volumetric medical information using Visual Transformers (ViTs). First, We extend the state-of-the-art Swin Transformer model to the 3D medical domain. Second, we propose a new approach…

图像与视频处理 · 电气工程与系统科学 2024-06-06 Cristhian Forigua , Maria Escobar , Pablo Arbelaez

Surgical video understanding is essential for computer-assisted interventions, yet existing surgical foundation models remain constrained by limited data scale, procedural diversity, and inconsistent evaluation, often lacking a reproducible…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Sicheng Lu , Zikai Xiao , Jianhui Wei , Danyu Sun , Qi Lu , Keli Hu , Yang Feng , Jian Wu , Zongxin Yang , Zuozhu Liu

Image annotation is one of the most essential tasks for guaranteeing proper treatment for patients and tracking progress over the course of therapy in the field of medical imaging and disease diagnosis. However, manually annotating a lot of…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Md Abdul Kadir , Hasan Md Tusfiqur Alam , Pascale Maul , Hans-Jürgen Profitlich , Moritz Wolf , Daniel Sonntag

Lung cancer is highly lethal, emphasizing the critical need for early detection. However, identifying lung nodules poses significant challenges for radiologists, who rely heavily on their expertise for accurate diagnosis. To address this…

图像与视频处理 · 电气工程与系统科学 2023-10-17 Hossein Jafari , Karim Faez , Hamidreza Amindavar

In robot learning, Vision Transformers (ViTs) are standard for visual perception, yet most methods discard valuable information by using only the final layer's features. We argue this provides an insufficient representation and propose the…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Wenhao Li , Chengwei Ma , Weixin Mao

Active learning (AL) can reduce annotation costs in surgical video analysis while maintaining model performance. However, traditional AL methods, developed for images or short video clips, are suboptimal for surgical step recognition due to…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Nisarg A. Shah , Bardia Safaei , Shameema Sikder , S. Swaroop Vedula , Vishal M. Patel

Endoscopic surgery is the gold standard for robotic-assisted minimally invasive surgery, offering significant advantages in early disease detection and precise interventions. However, the complexity of surgical scenes, characterized by high…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Guankun Wang , Rui Tang , Mengya Xu , Long Bai , Huxin Gao , Hongliang Ren

Visual-Language Models (VLMs) have significantly advanced action video recognition. Supervised by the semantics of action labels, recent works adapt the visual branch of VLMs to learn video representations. Despite the effectiveness proved…

计算机视觉与模式识别 · 计算机科学 2023-10-11 Yifei Chen , Dapeng Chen , Ruijin Liu , Hao Li , Wei Peng

Unsupervised video representation learning has made remarkable achievements in recent years. However, most existing methods are designed and optimized for video classification. These pre-trained models can be sub-optimal for temporal…

计算机视觉与模式识别 · 计算机科学 2022-03-28 Can Zhang , Tianyu Yang , Junwu Weng , Meng Cao , Jue Wang , Yuexian Zou

To enable a deep learning-based system to be used in the medical domain as a computer-aided diagnosis system, it is essential to not only classify diseases but also present the locations of the diseases. However, collecting instance-level…

计算机视觉与模式识别 · 计算机科学 2021-01-26 Hyun-Woo Kim , Hong-Gyu Jung , Seong-Whan Lee

Super-Resolution Ultrasound (SRUS) imaging through localising and tracking microbubbles, also known as Ultrasound Localisation Microscopy (ULM), has demonstrated significant potential for reconstructing microvasculature and flows with…

图像与视频处理 · 电气工程与系统科学 2025-03-26 Jipeng Yan , Qingyuan Tan , Shusei Kawara , Jingwen Zhu , Bingxue Wang , Matthieu Toulemonde , Honghai Liu , Ying Tan , Meng-Xing Tang

Multi-organ segmentation of 3D medical images is fundamental with meaningful applications in various clinical automation pipelines. Although deep learning has achieved superior performance, the time and memory consumption of segmenting the…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Xueqi Guo , Halid Ziya Yerebakan , Yoshihisa Shinagawa , Kritika Iyer , Gerardo Hermosillo Valadez

Semantic segmentation in surgical videos is a prerequisite for a broad range of applications towards improving surgical outcomes and surgical video analysis. However, semantic segmentation in surgical videos involves many challenges. In…

图像与视频处理 · 电气工程与系统科学 2021-09-28 Negin Ghamsarian , Mario Taschwer , Doris Putzgruber-Adamitsch , Stephanie Sarny , Yosuf El-Shabrawi , Klaus Schoeffmann

Convolutional operations have two limitations: (1) do not explicitly model where to focus as the same filter is applied to all the positions, and (2) are unsuitable for modeling long-range dependencies as they only operate on a small…

计算机视觉与模式识别 · 计算机科学 2020-08-03 Xiaofang Wang , Xuehan Xiong , Maxim Neumann , AJ Piergiovanni , Michael S. Ryoo , Anelia Angelova , Kris M. Kitani , Wei Hua

Automatic surgical activity recognition enables more intelligent surgical devices and a more efficient workflow. Integration of such technology in new operating rooms has the potential to improve care delivery to patients and decrease…

计算机视觉与模式识别 · 计算机科学 2022-07-08 Ali Mottaghi , Aidean Sharghi , Serena Yeung , Omid Mohareri

Video transformers have recently demonstrated strong potential for echocardiogram (echo) analysis, leveraging self-supervised pre-training and flexible adaptation across diverse tasks. However, like other models operating on videos, they…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Alexander Thorley , Agis Chartsias , Jordan Strom , Jeremy Slivnick , Dipak Kotecha , Alberto Gomez , Jinming Duan

Nucleus segmentation is an important analysis task in digital pathology. However, methods for automatic segmentation often struggle with new data from a different distribution, requiring users to manually annotate nuclei and retrain…

图像与视频处理 · 电气工程与系统科学 2025-06-03 Titus Griebel , Anwai Archit , Constantin Pape

Recent state-of-the-art performances of Vision Transformers (ViT) in computer vision tasks demonstrate that a general-purpose architecture, which implements long-range self-attention, could replace the local feature learning operations of…