中文
相关论文

相关论文: SViTT-Ego: A Sparse Video-Text Transformer for Ego…

200 篇论文

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding,…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Hanrong Ye , Haotian Zhang , Erik Daxberger , Lin Chen , Zongyu Lin , Yanghao Li , Bowen Zhang , Haoxuan You , Dan Xu , Zhe Gan , Jiasen Lu , Yinfei Yang

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we…

We present Ego-Only, the first approach that enables state-of-the-art action detection on egocentric (first-person) videos without any form of exocentric (third-person) transferring. Despite the content and appearance gap separating the two…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Huiyu Wang , Mitesh Kumar Singh , Lorenzo Torresani

Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yogesh Kulkarni , Pooyan Fazli

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Thomas Hummel , Shyamgopal Karthik , Mariana-Iuliana Georgescu , Zeynep Akata

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent…

Egocentric temporal action segmentation in videos is a crucial task in computer vision with applications in various fields such as mixed reality, human behavior analysis, and robotics. Although recent research has utilized advanced…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Sakib Reza , Balaji Sundareshan , Mohsen Moghaddam , Octavia Camps

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Vishakha Lall , Yisi Liu

In this report, we present the transferring pretrained video mask autoencoders(VideoMAE) to egocentric tasks for Ego4d Looking at me Challenge. VideoMAE is the data-efficient pretraining model for self-supervised video pre-training and can…

计算机视觉与模式识别 · 计算机科学 2022-11-30 Yinan He , Guo Chen

Conventional wisdom suggests that pre-training Vision Transformers (ViT) improves downstream performance by learning useful representations. Is this actually true? We investigate this question and find that the features and representations…

机器学习 · 计算机科学 2024-11-15 Alexander C. Li , Yuandong Tian , Beidi Chen , Deepak Pathak , Xinlei Chen

We present Vinci, a vision-language system designed to provide real-time, comprehensive AI assistance on portable devices. At its core, Vinci leverages EgoVideo-VL, a novel model that integrates an egocentric vision foundation model with a…

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Junbin Xiao , Shenglang Zhang , Pengxiang Zhu , Angela Yao

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Yuki Sakai , Ryosuke Furuta , Juichun Yen , Yoichi Sato

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

计算机视觉与模式识别 · 计算机科学 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Shuo Liang , Yiwu Zhong , Zi-Yuan Hu , Yeyao Tao , Liwei Wang

Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans'…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Danny Tran , Roberto Martín-Martín , Kristen Grauman

Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to egocentric videos has…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Anna Kukleva , Fadime Sener , Edoardo Remelli , Bugra Tekin , Eric Sauser , Bernt Schiele , Shugao Ma

This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker's view of the…

声音 · 计算机科学 2023-05-04 Roshan Sharma , Weipeng He , Ju Lin , Egor Lakomkin , Yang Liu , Kaustubh Kalgaonkar

Vision Transformer (ViT), a radically different architecture than convolutional neural networks offers multiple advantages including design simplicity, robustness and state-of-the-art performance on many vision tasks. However, in contrast…

计算机视觉与模式识别 · 计算机科学 2022-10-14 Hanan Gani , Muzammal Naseer , Mohammad Yaqub

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Baoqi Pei , Yifei Huang , Jilan Xu , Yuping He , Guo Chen , Fei Wu , Yu Qiao , Jiangmiao Pang