中文
相关论文

相关论文: SViTT-Ego: A Sparse Video-Text Transformer for Ego…

200 篇论文

Due to its deficiency in prior knowledge (inductive bias), Vision Transformer (ViT) requires pre-training on large-scale datasets to perform well. Moreover, the growing layers and parameters in ViT models impede their applicability to…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Chenhao Xu , Chang-Tsun Li , Chee Peng Lim , Douglas Creighton

Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining architectures and objectives within the masked autoencoding…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Mariana-Iuliana Georgescu , Eduardo Fonseca , Radu Tudor Ionescu , Mario Lucic , Cordelia Schmid , Anurag Arnab

Learning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning the meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned…

In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Heqian Qiu , Zhaofeng Shi , Lanxiao Wang , Huiyu Xiong , Xiang Li , Hongliang Li

Large pre-trained vision models achieve impressive success in computer vision. However, fully fine-tuning large models for downstream tasks, particularly in video understanding, can be prohibitively computationally expensive. Recent studies…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Huanjin Yao , Wenhao Wu , Zhiheng Li

AI assistants that support humans in daily life are becoming increasingly feasible, driven by the rapid advancements in multimodal language models. A key challenge lies in overcoming the generic nature of these models to deliver…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Soroush Seifi , Simon Gardier , Vaggelis Dorovatas , Daniel Olmeda Reino , Rahaf Aljundi

Vision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities, with the majority focusing on the third-person perspective, and only a few…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Sijie Cheng , Zhicheng Guo , Jingwen Wu , Kechen Fang , Peng Li , Huaping Liu , Yang Liu

While exocentric video synthesis has achieved great progress, egocentric video generation remains largely underexplored, which requires modeling first-person view content along with camera motion patterns induced by the wearer's body…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Jingqiao Xiu , Fangzhou Hong , Yicong Li , Mengze Li , Wentao Wang , Sirui Han , Liang Pan , Ziwei Liu

We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Yutong Bai , Danny Tran , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Taewoong Kang , Kinam Kim , Dohyeon Kim , Minho Park , Junha Hyung , Jaegul Choo

Understanding visual inputs for a given task amidst varied changes is a key challenge posed by visual reinforcement learning agents. We propose \textit{Value Explicit Pretraining} (VEP), a method that learns generalizable representations…

机器学习 · 计算机科学 2026-05-04 Kiran Lekkala , Henghui Bao , Sumedh A. Sontakke , Erdem Biyik , Laurent Itti

We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Jathushan Rajasegaran , Ilija Radosavovic , Rahul Ravishankar , Yossi Gandelsman , Christoph Feichtenhofer , Jitendra Malik

Dynamic head pruning in Vision Transformers (ViTs) improves efficiency by removing redundant attention heads, but existing pruning policies are often difficult to interpret and control. In this work, we propose a novel framework by…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yousung Lee , Dongsoo Har

This paper addresses key challenges in object-centric representation learning of video. While existing approaches struggle with complex scenes, we propose a novel weakly-supervised framework that emphasises geometric understanding and…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Phúc H. Le Khac , Graham Healy , Alan F. Smeaton

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

计算机视觉与模式识别 · 计算机科学 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

In this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior understanding. The model aims to capture and extract useful…

多媒体 · 计算机科学 2022-01-25 Minh Tran , Mohammad Soleymani

Predicting turn-taking in multiparty conversations has many practical applications in human-computer/robot interaction. However, the complexity of human communication makes it a challenging task. Recent advances have shown that synchronous…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Mehdi Fatan , Emanuele Mincato , Dimitra Pintzou , Mariella Dimiccoli

While transformers have surpassed convolutional neural networks (CNNs) in various computer vision tasks, microelectronics defect detection still largely relies on CNNs. We hypothesize that this gap is due to the fact that a) transformers…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Nikolai Röhrich , Alwin Hoffmann , Richard Nordsieck , Emilio Zarbali , Alireza Javanmardi

Estimating 3D human motion from an egocentric video sequence plays a critical role in human behavior understanding and has various applications in VR/AR. However, naively learning a mapping between egocentric videos and human motions is…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Jiaman Li , C. Karen Liu , Jiajun Wu

While many visual ego-motion algorithm variants have been proposed in the past decade, learning based ego-motion estimation methods have seen an increasing attention because of its desirable properties of robustness to image noise and…

计算机视觉与模式识别 · 计算机科学 2019-06-20 Guangyao Zhai , Liang Liu , Linjian Zhang , Yong Liu