中文
相关论文

相关论文: PEg TRAnsfer Workflow recognition challenge report…

200 篇论文

Automatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical…

计算机视觉与模式识别 · 计算机科学 2021-06-30 Yonghao Long , Jie Ying Wu , Bo Lu , Yueming Jin , Mathias Unberath , Yun-Hui Liu , Pheng Ann Heng , Qi Dou

With the advent of large pre-trained transformer models, fine-tuning these models for various downstream tasks is a critical problem. Paucity of training data, the existence of data silos, and stringent privacy constraints exacerbate this…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Naif Alkhunaizi , Faris Almalik , Rouqaiah Al-Refai , Muzammal Naseer , Karthik Nandakumar

Automated personality and soft skill assessment from multimodal behavioral data remains challenging due to limited datasets and methods that fail to capture geometric structure inherent in human traits. We introduce RecruitView, a dataset…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Amit Kumar Gupta , Farhan Sheth , Hammad Shaikh , Dheeraj Kumar , Angkul Puniya , Deepak Panwar , Sandeep Chaurasia , Priya Mathur

Multi-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learning offers excellent recognition results, its computational…

计算机视觉与模式识别 · 计算机科学 2021-05-13 Rameswar Panda , Chun-Fu Chen , Quanfu Fan , Ximeng Sun , Kate Saenko , Aude Oliva , Rogerio Feris

The recent advances in deep learning (DL) have been accelerated by access to large-scale data and compute. These large-scale resources have been used to train progressively larger models which are resource intensive in terms of compute,…

机器学习 · 计算机科学 2024-12-06 Raghavendra Selvan , Bob Pepin , Christian Igel , Gabrielle Samuel , Erik B Dam

Multimodal information (e.g., visual, acoustic, and textual) has been widely used to enhance representation learning for micro-video recommendation. For integrating multimodal information into a joint representation of micro-video,…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Han Liu , Yinwei Wei , Fan Liu , Wenjie Wang , Liqiang Nie , Tat-Seng Chua

Pre-trained vision models (PVMs) have demonstrated remarkable adaptability across a wide range of downstream vision tasks, showcasing exceptional performance. However, as these models scale to billions or even trillions of parameters,…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Yi Xin , Jianjiang Yang , Siqi Luo , Yuntao Du , Qi Qin , Kangrui Cen , Yangfan He , Zhiwei Zhang , Bin Fu , Xiaokang Yang , Guangtao Zhai , Ming-Hsuan Yang , Xiaohong Liu

Data-driven approaches to assist operating room (OR) workflow analysis depend on large curated datasets that are time consuming and expensive to collect. On the other hand, we see a recent paradigm shift from supervised learning to…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Muhammad Abdullah Jamal , Omid Mohareri

Automated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance…

计算机视觉与模式识别 · 计算机科学 2021-07-20 Xueying Shi , Yueming Jin , Qi Dou , Jing Qin , Pheng-Ann Heng

Accurate surgical phase recognition is essential for analyzing procedural workflows, supporting intraoperative decision-making, and enabling data-driven improvements in surgical education and performance evaluation. In this work, we present…

Capitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Shu Yang , Zhiyuan Cai , Luyang Luo , Ning Ma , Shuchang Xu , Hao Chen

The lack of data and the difficulty of multimodal fusion have always been challenges for multimodal emotion recognition (MER). In this paper, we propose to use pretrained models as upstream network, wav2vec 2.0 for audio modality and BERT…

计算与语言 · 计算机科学 2023-02-28 Dekai Sun , Yancheng He , Jiqing Han

Ambivalence/hesitancy recognition in unconstrained videos is a challenging problem due to the subtle, multimodal, and context-dependent nature of this behavioral state. In this paper, a multimodal approach for video-level…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Elena Ryumina , Alexandr Axyonov , Dmitry Sysoev , Timur Abdulkadirov , Kirill Almetov , Yulia Morozova , Dmitry Ryumin

Pre-training on large-scale datasets and utilizing margin-based loss functions have been highly successful in training models for high-resolution face recognition. However, these models struggle with low-resolution face datasets, in which…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Kartik Narayan , Nithin Gopalakrishnan Nair , Jennifer Xu , Rama Chellappa , Vishal M. Patel

Real-time computational speed and a high degree of precision are requirements for computer-assisted interventions. Applying a segmentation network to a medical video processing task can introduce significant inter-frame prediction noise.…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Robert Mendel , Tobias Rueckert , Dirk Wilhelm , Daniel Rueckert , Christoph Palm

Aligning features from different modalities, is one of the most fundamental challenges for cross-modal tasks. Although pre-trained vision-language models can achieve a general alignment between image and text, they often require…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Ziqi Jiang , Yanghao Wang , Long Chen

Recognizing surgical gestures in real-time is a stepping stone towards automated activity recognition, skill assessment, intra-operative assistance, and eventually surgical automation. The current robotic surgical systems provide us with…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Jumanh Atoum , Garrison L. H. Johnston , Nabil Simaan , Jie Ying Wu

Out of all existing frameworks for surgical workflow analysis in endoscopic videos, action triplet recognition stands out as the only one aiming to provide truly fine-grained and comprehensive information on surgical activities. This…

计算机视觉与模式识别 · 计算机科学 2022-04-04 Chinedu Innocent Nwoye , Tong Yu , Cristians Gonzalez , Barbara Seeliger , Pietro Mascagni , Didier Mutter , Jacques Marescaux , Nicolas Padoy