中文
相关论文

相关论文: HD-EPIC: A Highly-Detailed Egocentric Video Datase…

200 篇论文

Egocentric, or first-person vision which became popular in recent years with an emerge in wearable technology, is different than exocentric (third-person) vision in some distinguishable ways, one of which being that the camera wearer is…

计算机视觉与模式识别 · 计算机科学 2016-10-11 Jessica Finocchiaro , Aisha Urooj Khan , Ali Borji

Humans develop visual intelligence through perceiving and interacting with their environment - a self-supervised learning process grounded in egocentric experience. Inspired by this, we ask how can artificial systems learn stable object…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yuting Tan , Xilong Cheng , Yunxiao Qin , Zhengnan Li , Jingjing Zhang

Egocentric vision is an emerging field of computer vision that is characterized by the acquisition of images and video from the first person perspective. In this paper we address the challenge of egocentric human action recognition by…

计算机视觉与模式识别 · 计算机科学 2019-05-03 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas P. J. J. Noldus , Remco C. Veltkamp

Understanding human behavior from complementary egocentric (ego) and exocentric (exo) points of view enables the development of systems that can support workers in industrial environments and enhance their safety. However, progress in this…

In this paper, we propose a novel approach to enhance the 3D body pose estimation of a person computed from videos captured from a single wearable camera. The key idea is to leverage high-level features linking first- and third-views in a…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Ameya Dhamanaskar , Mariella Dimiccoli , Enric Corona , Albert Pumarola , Francesc Moreno-Noguer

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Junbin Xiao , Nanxin Huang , Hao Qiu , Zhulin Tao , Xun Yang , Richang Hong , Meng Wang , Angela Yao

Recent Multi-modal Large Language Models (MLLMs) have made great progress in video understanding. However, their performance on videos involving human actions is still limited by the lack of high-quality data. To address this, we introduce…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Xiao Wang , Jingyun Hua , Weihong Lin , Yuanxing Zhang , Fuzheng Zhang , Jianlong Wu , Di Zhang , Liqiang Nie

Egocentric action anticipation consists in understanding which objects the camera wearer will interact with in the near future and which actions they will perform. We tackle the problem proposing an architecture able to anticipate actions…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Antonino Furnari , Giovanni Maria Farinella

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment,…

Humans exhibit adaptive, context-sensitive responses to egocentric visual input. However, faithfully modeling such reactions from egocentric video remains challenging due to the dual requirements of strictly causal generation and precise 3D…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Libo Zhang , Zekun Li , Tianyu Li , Zeyu Cao , Rui Xu , Xiaoxiao Long , Wenjia Wang , Jingbo Wang , Yuan Liu , Wenping Wang , Daquan Zhou , Taku Komura , Zhiyang Dou

We introduce FEEL (Force-Enhanced Egocentric Learning), the first large-scale dataset pairing force measurements gathered from custom piezoresistive gloves with egocentric video. Our gloves enable scalable data collection, and FEEL contains…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Eadom Dessalene , Botao He , Michael Maynord , Yonatan Tussa , Pavan Mantripragada , Yianni Karabati , Nirupam Roy , Yiannis Aloimonos

Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecast what will happen…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Tatiana Tommasi , Giuseppe Averta

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we…

计算机视觉与模式识别 · 计算机科学 2019-08-01 Antoine Miech , Dimitri Zhukov , Jean-Baptiste Alayrac , Makarand Tapaswi , Ivan Laptev , Josef Sivic

Surgical tool detection is a fundamental task for understanding egocentric open surgery videos. However, detecting surgical tools presents significant challenges due to their highly imbalanced class distribution, similar shapes and similar…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Ryo Fujii , Hideo Saito , Hiroki Kajita

Accurate 3D understanding of human hands and objects during manipulation remains a significant challenge for egocentric computer vision. Existing hand-object interaction datasets are predominantly captured in controlled studio settings,…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Patrick Rim , Kevin Harris , Braden Copple , Shangchen Han , Xu Xie , Ivan Shugurov , Sizhe An , He Wen , Alex Wong , Tomas Hodan , Kun He

Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view data and poor generalization to new viewpoints. We introduce…

3D hand pose estimation in everyday egocentric images is challenging for several reasons: poor visual signal (occlusion from the object of interaction, low resolution & motion blur), large perspective distortion (hands are close to the…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Aditya Prakash , Ruisen Tu , Matthew Chang , Saurabh Gupta

Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for minutes or hours…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Md Mohaiminul Islam , Ngan Ho , Xitong Yang , Tushar Nagarajan , Lorenzo Torresani , Gedas Bertasius

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera movement.…

Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Alkesh Patel , Vibhav Chitalia , Yinfei Yang
‹ 上一页 1 8 9 10 下一页 ›