中文
相关论文

相关论文: Domain Generalization using Action Sequences for E…

200 篇论文

Visual reinforcement learning policies trained on pixel observations often struggle to generalize when visual conditions change at test time. Object-centric representations are a promising alternative, but most approaches use fixed-size…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Alexandre Brown , Glen Berseth

Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shifts such as viewpoint distortion, scale variation, illumination changes, and environmental…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Ze Dong , Hao Shi , Zejia Gao , Zhonghua Yi , Kaiwei Wang , Lin Wang

What if a video generation model could not only imagine a plausible future, but the correct one, accurately reflecting how the world changes with each action? We address this question by presenting the Egocentric World Model (EgoWM), a…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Anurag Bagchi , Zhipeng Bao , Homanga Bharadhwaj , Yu-Xiong Wang , Pavel Tokmakov , Martial Hebert

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

计算机视觉与模式识别 · 计算机科学 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

Human-object interaction segmentation is a fundamental task of daily activity understanding, which plays a crucial role in applications such as assistive robotics, healthcare, and autonomous systems. Most existing learning-based methods…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Hao Xing , Kai Zhe Boey , Gordon Cheng

Given a video captured from a first person perspective and the environment context of where the video is recorded, can we recognize what the person is doing and identify where the action occurs in the 3D space? We address this challenging…

计算机视觉与模式识别 · 计算机科学 2022-08-16 Miao Liu , Lingni Ma , Kiran Somasundaram , Yin Li , Kristen Grauman , James M. Rehg , Chao Li

In this paper we propose a sequential learning framework for Domain Generalization (DG), the problem of training a model that is robust to domain shift by design. Various DG approaches have been proposed with different motivating…

计算机视觉与模式识别 · 计算机科学 2020-04-06 Da Li , Yongxin Yang , Yi-Zhe Song , Timothy Hospedales

Unsupervised segmentation of action segments in egocentric videos is a desirable feature in tasks such as activity recognition and content-based video retrieval. Reducing the search space into a finite set of action segments facilitates a…

计算机视觉与模式识别 · 计算机科学 2021-06-24 I. Hipiny , H. Ujir , J. L. Minoi , S. F. Samson Juan , M. A. Khairuddin , M. S. Sunar

Electroencephalography (EEG) based emotion recognition has demonstrated tremendous improvement in recent years. Specifically, numerous domain adaptation (DA) algorithms have been exploited in the past five years to enhance the…

信号处理 · 电气工程与系统科学 2022-04-20 Yan Li , Hao Chen , Jake Zhao , Haolan Zhang , Jinpeng Li

Person re-identification (re-ID) in first-person (egocentric) vision is a fairly new and unexplored problem. With the increase of wearable video recording devices, egocentric data becomes readily available, and person re-identification has…

计算机视觉与模式识别 · 计算机科学 2021-03-09 Ankit Choudhary , Deepak Mishra , Arnab Karmakar

We study the domain adaptation task for action recognition, namely domain adaptive action recognition, which aims to effectively transfer action recognition power from a label-sufficient source domain to a label-free target domain. Since…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Kun-Yu Lin , Jiaming Zhou , Wei-Shi Zheng

Unlike traditional third-person cameras mounted on robots, a first-person camera, captures a person's visual sensorimotor object interactions from up close. In this paper, we study the tight interplay between our momentary visual attention…

计算机视觉与模式识别 · 计算机科学 2017-06-13 Gedas Bertasius , Hyun Soo Park , Stella X. Yu , Jianbo Shi

"Looking for things" is a mundane but critical task we repeatedly carry on in our daily life. We introduce a method to develop a human character capable of searching for a randomly located target object in a detailed 3D scene using its…

机器人学 · 计算机科学 2021-09-16 Maks Sorokin , Wenhao Yu , Sehoon Ha , C. Karen Liu

Recent progress of self-supervised visual representation learning has achieved remarkable success on many challenging computer vision benchmarks. However, whether these techniques can be used for domain adaptation has not been explored. In…

计算机视觉与模式识别 · 计算机科学 2019-12-12 Jiaolong Xu , Liang Xiao , Antonio M. Lopez

Purpose: Advances in deep learning have resulted in effective models for surgical video analysis; however, these models often fail to generalize across medical centers due to domain shift caused by variations in surgical workflow, camera…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Siddhant Satyanaik , Aditya Murali , Deepak Alapatt , Xin Wang , Pietro Mascagni , Nicolas Padoy

In this paper we propose an end-to-end trainable deep neural network model for egocentric activity recognition. Our model is built on the observation that egocentric activities are highly characterized by the objects and their locations in…

计算机视觉与模式识别 · 计算机科学 2018-08-01 Swathikiran Sudhakaran , Oswald Lanz

Activity recognition from long unstructured egocentric photo-streams has several applications in assistive technology such as health monitoring and frailty detection, just to name a few. However, one of its main technical challenges is to…

计算机视觉与模式识别 · 计算机科学 2017-08-29 Alejandro Cartas , Mariella Dimiccoli , Petia Radeva

Advancements in egocentric video datasets like Ego4D, EPIC-Kitchens, and Ego-Exo4D have enriched the study of first-person human interactions, which is crucial for applications in augmented reality and assisted living. Despite these…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Joungbin An , Yunsu Park , Hyolim Kang , Seon Joo Kim

This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length videos, capturing…

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-body action…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Yuan-Ming Li , Wei-Jin Huang , An-Lan Wang , Ling-An Zeng , Jing-Ke Meng , Wei-Shi Zheng