English
Related papers

Related papers: ViTALS: Vision Transformer for Action Localization…

200 papers

Surgical phase recognition is crucial to providing surgery understanding in smart operating rooms. Despite great progress in automatic surgical phase recognition, most existing methods are still restricted by two problems. First, these…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Xingjian Luo , You Pang , Zhen Chen , Jinlin Wu , Zongmin Zhang , Zhen Lei , Hongbin Liu

As of today, state-of-the-art activity recognition from wearable sensors relies on algorithms being trained to classify fixed windows of data. In contrast, video-based Human Activity Recognition, known as Temporal Action Localization (TAL),…

Machine Learning · Computer Science 2024-10-15 Marius Bock , Michael Moeller , Kristof Van Laerhoven

Understanding surgical tasks represents an important challenge for autonomy in surgical robotic systems. To achieve this, we propose an online task segmentation framework that uses hierarchical transition state clustering to activate…

Robotics · Computer Science 2024-06-17 Yutaro Yamada , Jacinto Colan , Ana Davila , Yasuhisa Hasegawa

Vital sign (breathing and heartbeat) monitoring is essential for patient care and sleep disease prevention. Most current solutions are based on wearable sensors or cameras; however, the former could affect sleep quality, while the latter…

Human-Computer Interaction · Computer Science 2023-05-25 Xiang Zhang , Yu Gu , Huan Yan , Yantong Wang , Mianxiong Dong , Kaoru Ota , Fuji Ren , Yusheng Ji

While large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos is still a relatively unexplored task. We design a new…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Shen Yan , Xuehan Xiong , Arsha Nagrani , Anurag Arnab , Zhonghao Wang , Weina Ge , David Ross , Cordelia Schmid

Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Sifeng He , Yue He , Minlong Lu , Chen Jiang , Xudong Yang , Feng Qian , Xiaobo Zhang , Lei Yang , Jiandong Zhang

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

Robotics · Computer Science 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

Open, or non-laparoscopic surgery, represents the vast majority of all operating room procedures, but few tools exist to objectively evaluate these techniques at scale. Current efforts involve human expert-based visual assessment. We…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Michael Zhang , Xiaotian Cheng , Daniel Copeland , Arjun Desai , Melody Y. Guan , Gabriel A. Brat , Serena Yeung

In minimally invasive surgery, surgical instrument localization is a crucial task for endoscopic videos, which enables various applications for improving surgical outcomes. However, annotating the instrument localization in endoscopic…

Image and Video Processing · Electrical Eng. & Systems 2024-06-24 Rongfeng Wei , Jinlin Wu , Xuexue Bai , Ming Feng , Zhen Lei , Hongbin Liu , Zhen Chen

We present CataractSAM-2, a domain-adapted extension of Meta's Segment Anything Model 2, designed for real-time semantic segmentation of cataract ophthalmic surgery videos with high accuracy. Positioned at the intersection of computer…

Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per action instance) to train a model to effectively locate action instances within untrimmed…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Yunchuan Ma , Laiyun Qing , Guorong Li , Yuqing Liu , Yuankai Qi , Qingming Huang

The automatic analysis of the surgical process, from videos recorded during surgeries, could be very useful to surgeons, both for training and for acquiring new techniques. The training process could be optimized by automatically providing…

Computer Vision and Pattern Recognition · Computer Science 2016-10-19 Katia Charrière , Gwenolé Quellec , Mathieu Lamard , David Martiano , Guy Cazuguel , Gouenou Coatrieux , Béatrice Cochener

Medical image segmentation plays a crucial role in various healthcare applications, enabling accurate diagnosis, treatment planning, and disease monitoring. Traditionally, convolutional neural networks (CNNs) dominated this domain,…

Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to…

Human-Computer Interaction · Computer Science 2025-05-26 Yuchen He , Jianbing Lv , Liqi Cheng , Lingyu Meng , Dazhen Deng , Yingcai Wu

Zero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model on a large amount of annotated training data. While…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Benedetta Liberatori , Alessandro Conti , Paolo Rota , Yiming Wang , Elisa Ricci

Weakly supervised object localization (WSOL) aims at predicting object locations in an image using only image-level category labels. Common challenges that image classification models encounter when localizing objects are, (a) they tend to…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Saurav Gupta , Sourav Lakhotia , Abhay Rawat , Rahul Tallamraju

Introduction: Computer vision (CV) has had a transformative impact in biomedical fields such as radiology, dermatology, and pathology. Its real-world adoption in surgical applications, however, remains limited. We review the current…

Image and Video Processing · Electrical Eng. & Systems 2025-02-25 Devanish N. Kamtam , Joseph B. Shrager , Satya Deepya Malla , Nicole Lin , Juan J. Cardona , Jake J. Kim , Clarence Hu

In visual exploration and analysis of data, determining how to select and transform the data for visualization is a challenge for data-unfamiliar or inexperienced users. Our main hypothesis is that for many data sets and common analysis…

This paper considers the problem of localizing actions in videos as a sequences of bounding boxes. The objective is to generate action proposals that are likely to include the action of interest, ideally achieving high recall with few…

Computer Vision and Pattern Recognition · Computer Science 2016-07-08 Mihir Jain , Jan van Gemert , Hervé Jégou , Patrick Bouthemy , Cees G. M. Snoek