English
Related papers

Related papers: Developing Vision-Language-Action Model from Egoce…

200 papers

This work adapts a deep neural model for image saliency prediction to the temporal domain of egocentric video. We compute the saliency map for each video frame, firstly with an off-the-shelf model trained from static images, secondly by…

Computer Vision and Pattern Recognition · Computer Science 2018-09-06 Panagiotis Linardos , Eva Mohedano , Monica Cherto , Cathal Gurrin , Xavier Giro-i-Nieto

Understanding human activity is a crucial yet intricate task in egocentric vision, a field that focuses on capturing visual perspectives from the camera wearer's viewpoint. Traditional methods heavily rely on representation learning that is…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Sanghwan Kim , Daoji Huang , Yongqin Xian , Otmar Hilliges , Luc Van Gool , Xi Wang

Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often suffers from a substantial visual domain gap and limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Chenyu Hui , Xiaodi Huang , Siyu Xu , Yunke Wang , Shan You , Fei Wang , Tao Huang , Chang Xu

Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visual data. However, capturing such scenarios in the real world is often difficult, costly,…

Computation and Language · Computer Science 2026-05-12 Yu-Hsiang Liu , Yu-Chien Tang , An-Zi Yen

Egocentric, or first-person vision which became popular in recent years with an emerge in wearable technology, is different than exocentric (third-person) vision in some distinguishable ways, one of which being that the camera wearer is…

Computer Vision and Pattern Recognition · Computer Science 2016-10-11 Jessica Finocchiaro , Aisha Urooj Khan , Ali Borji

In egocentric videos, actions occur in quick succession. We capitalise on the action's temporal context and propose a method that learns to attend to surrounding actions in order to improve recognition performance. To incorporate the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Evangelos Kazakos , Jaesung Huh , Arsha Nagrani , Andrew Zisserman , Dima Damen

Recently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-language transformer models do not explicitly fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2022-05-19 Alex Jinpeng Wang , Yixiao Ge , Guanyu Cai , Rui Yan , Xudong Lin , Ying Shan , Xiaohu Qie , Mike Zheng Shou

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower model and leveraging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Baoqi Pei , Guo Chen , Jilan Xu , Yuping He , Yicheng Liu , Kanghua Pan , Yifei Huang , Yali Wang , Tong Lu , Limin Wang , Yu Qiao

Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Tao Lin , Yuxin Du , Jiting Liu , Nuobei Zhu , Yunhe Li , Yuqian Fu , Yinxinyu Chen , Hongyi Cai , Zewei Ye , Bing Cheng , Kai Ye , Yiran Mao , Yilei Zhong , MingKang Dong , Junchi Yan , Gen Li , Bo Zhao

Vision-Language-Action (VLA) models have gained popularity for learning robotic manipulation tasks that follow language instructions. State-of-the-art VLAs, such as OpenVLA and $\pi_{0}$, were trained on large-scale, manually labeled action…

Robotics · Computer Science 2025-09-24 Bahey Tharwat , Yara Nasser , Ali Abouzeid , Ian Reid

Long-horizon robotic manipulation remains challenging for Vision-Language-Action (VLA) models despite recent progress in zero-shot generalization and simulation-to-real-world transfer. Current VLA models suffer from stage hallucination,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Zeting Liu , Zida Yang , Zeyu Zhang , Hao Tang

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera movement.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Ashish Seth , Xinhao Mei , Changsheng Zhao , Varun Nagaraja , Ernie Chang , Gregory P. Meyer , Gael Le Lan , Yunyang Xiong , Vikas Chandra , Yangyang Shi , Dinesh Manocha , Zhipeng Cai

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid egomotion and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Runjia Li , Moayed Haji-Ali , Ashkan Mirzaei , Chaoyang Wang , Arpit Sahni , Ivan Skorokhodov , Aliaksandr Siarohin , Tomas Jakab , Junlin Han , Sergey Tulyakov , Philip Torr , Willi Menapace

We introduce FEEL (Force-Enhanced Egocentric Learning), the first large-scale dataset pairing force measurements gathered from custom piezoresistive gloves with egocentric video. Our gloves enable scalable data collection, and FEEL contains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Eadom Dessalene , Botao He , Michael Maynord , Yonatan Tussa , Pavan Mantripragada , Yianni Karabati , Nirupam Roy , Yiannis Aloimonos

In this paper we propose an end-to-end trainable deep neural network model for egocentric activity recognition. Our model is built on the observation that egocentric activities are highly characterized by the objects and their locations in…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Swathikiran Sudhakaran , Oswald Lanz

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained…

Robotics · Computer Science 2025-12-09 Yichao Shen , Fangyun Wei , Zhiying Du , Yaobo Liang , Yan Lu , Jiaolong Yang , Nanning Zheng , Baining Guo

Wearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in egocentric settings and…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Francesco Ragusa , Antonino Furnari , Salvatore Livatino , Giovanni Maria Farinella

The possibility of sharing one's point of view makes use of wearable cameras compelling. These videos are often long, boring and coupled with extreme shake, as the camera is worn on a moving person. Fast forwarding (i.e. frame sampling) is…

Computer Vision and Pattern Recognition · Computer Science 2017-01-13 Tavi Halperin , Yair Poleg , Chetan Arora , Shmuel Peleg

Egocentric cameras are becoming increasingly popular and provide us with large amounts of videos, captured from the first person perspective. At the same time, surveillance cameras and drones offer an abundance of visual information, often…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Shervin Ardeshir , Ali Borji

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs…

‹ Prev 1 8 9 10 Next ›