English
Related papers

Related papers: SkillFormer: Unified Multi-View Video Understandin…

200 papers

Can Video-LLMs achieve consistent temporal understanding when videos capture the same event from different viewpoints? To study this, we introduce EgoExo-Con (Consistency), a benchmark of comprehensively synchronized egocentric and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Minjoon Jung , Junbin Xiao , Junghyun Kim , Byoung-Tak Zhang , Angela Yao

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Rawal Khirodkar , Aayush Bansal , Lingni Ma , Richard Newcombe , Minh Vo , Kris Kitani

This paper presents an investigation of vision transformer learning for multi-view geometry tasks, such as optical flow estimation, by fine-tuning video foundation models. Unlike previous methods that involve custom architectural designs…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Huimin Wu , Kwang-Ting Cheng , Stephen Lin , Zhirong Wu

Few-shot classification which aims to recognize unseen classes using very limited samples has attracted more and more attention. Usually, it is formulated as a metric learning problem. The core issue of few-shot classification is how to…

Computer Vision and Pattern Recognition · Computer Science 2022-08-29 Xixi Wang , Xiao Wang , Bo Jiang , Bin Luo

Multi-person 3D mesh recovery from videos is a critical first step towards automatic perception of group behavior in virtual reality, physical therapy and beyond. However, existing approaches rely on multi-stage paradigms, where the person…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Haoyuan Li , Haoye Dong , Hanchao Jia , Dong Huang , Michael C. Kampffmeyer , Liang Lin , Xiaodan Liang

Egocentric video gaze estimation requires models to capture individual gaze patterns while adapting to diverse user data. Our approach leverages a transformer-based architecture, integrating it into a PFL framework where only the most…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Yuhu Feng , Keisuke Maeda , Takahiro Ogawa , Miki Haseyama

Efficient video action recognition remains a challenging problem. One large model after another takes the place of the state-of-the-art on the Kinetics dataset, but real-world efficiency evaluations are often lacking. In this work, we fill…

Computer Vision and Pattern Recognition · Computer Science 2021-11-19 Raivo Koot , Haiping Lu

Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 G. Thomas Hudson , Dean Slack , Thomas Winterbottom , Jamie Sterling , Chenghao Xiao , Junjie Shentu , Noura Al Moubayed

We present WidthFormer, a novel transformer-based module to compute Bird's-Eye-View (BEV) representations from multi-view cameras for real-time autonomous-driving applications. WidthFormer is computationally efficient, robust and does not…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Chenhongyi Yang , Tianwei Lin , Lichao Huang , Elliot J. Crowley

In this paper, we propose self-supervised training for video transformers using unlabeled video data. From a given video, we create local and global spatiotemporal views with varying spatial sizes and frame rates. Our self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Kanchana Ranasinghe , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan , Michael Ryoo

The development of effective training and evaluation strategies is critical. Conventional methods for assessing surgical proficiency typically rely on expert supervision, either through onsite observation or retrospective analysis of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yan Meng , Daniel A. Donoho , Marcelle Altshuler , Omar Arnaout

While large-scale diffusion models have revolutionized video synthesis, achieving precise control over both multi-subject identity and multi-granularity motion remains a significant challenge. Recent attempts to bridge this gap often suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yujie Wei , Xinyu Liu , Shiwei Zhang , Hangjie Yuan , Jinbo Xing , Zhekai Chen , Xiang Wang , Haonan Qiu , Rui Zhao , Yutong Feng , Ruihang Chu , Yingya Zhang , Yike Guo , Xihui Liu , Hongming Shan

Tabular data from different tables exhibit significant diversity due to varied definitions and types of features, as well as complex inter-feature and feature-target relationships. Cross-dataset pretraining, which learns reusable patterns…

Machine Learning · Computer Science 2024-06-04 Jintai Chen , Zhen Lin , Qiyuan Chen , Jimeng Sun

3D human pose estimation can be handled by encoding the geometric dependencies between the body parts and enforcing the kinematic constraints. Recently, Transformer has been adopted to encode the long-range dependencies between the joints…

Computer Vision and Pattern Recognition · Computer Science 2022-03-28 Mohammed Hassanin , Abdelwahed Khamiss , Mohammed Bennamoun , Farid Boussaid , Ibrahim Radwan

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns unified BEV…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Zhiqi Li , Wenhai Wang , Hongyang Li , Enze Xie , Chonghao Sima , Tong Lu , Qiao Yu , Jifeng Dai

The recent success of Transformers in the language domain has motivated adapting it to a multimodal setting, where a new visual model is trained in tandem with an already pretrained language model. However, due to the excessive memory…

Computer Vision and Pattern Recognition · Computer Science 2021-09-23 Sangho Lee , Youngjae Yu , Gunhee Kim , Thomas Breuel , Jan Kautz , Yale Song

Self-supervised methods have showed promising results on depth estimation task. However, previous methods estimate the target depth map and camera ego-motion simultaneously, underusing multi-frame correlation information and ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Songchun Zhang , Chunhui Zhao

View-based methods have demonstrated promising performance in 3D shape understanding. However, they tend to make strong assumptions about the relations between views or learn the multi-view correlations indirectly, which limits the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Hongyu Sun , Yongcai Wang , Peng Wang , Haoran Deng , Xudong Cai , Deying Li

This paper presents a novel method for face clustering in videos using a video-centralised transformer. Previous works often employed contrastive learning to learn frame-level representation and used average pooling to aggregate the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-16 Yujiang Wang , Mingzhi Dong , Jie Shen , Yiming Luo , Yiming Lin , Pingchuan Ma , Stavros Petridis , Maja Pantic

Automated surgical step recognition is an important task that can significantly improve patient safety and decision-making during surgeries. Existing state-of-the-art methods for surgical step recognition either rely on separate,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-22 Nisarg A. Shah , Shameema Sikder , S. Swaroop Vedula , Vishal M. Patel