中文
相关论文

相关论文: LSTA: Long Short-Term Attention for Egocentric Act…

200 篇论文

Understanding instructional videos requires recognizing fine-grained actions and modeling their temporal relations, which remains challenging for current Video Foundation Models (VFMs). This difficulty stems from noisy web supervision and a…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Zhuoyi Yang , Jiapeng Yu , Reuben Tan , Boyang Li , Huijuan Xu

Video prediction aims to predict future frames by modeling the complex spatiotemporal dynamics in videos. However, most of the existing methods only model the temporal information and the spatial information for videos in an independent…

计算机视觉与模式识别 · 计算机科学 2022-04-21 Zheng Chang , Xinfeng Zhang , Shanshe Wang , Siwei Ma , Wen Gao

In the research field of few-shot learning, the main difference between image-based and video-based is the additional temporal dimension. In recent years, some works have used the Transformer to deal with frames, then get the attention…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Fei Guo , Li Zhu , YiWang Wang , Jing Sun

The design of deep learning methods for low light video enhancement remains a challenging problem owing to the difficulty in capturing low light and ground truth video pairs. This is particularly hard in the context of dynamic scenes or…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Shivam Chhirolya , Sameer Malik , Rajiv Soundararajan

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Yuki Sakai , Ryosuke Furuta , Juichun Yen , Yoichi Sato

Generating long, coherent egocentric videos is difficult, as hand-object interactions and procedural tasks require reliable long-term memory. Existing autoregressive models suffer from content drift, where object identity and scene…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Liuzhou Zhang , Jiarui Ye , Yuanlei Wang , Ming Zhong , Mingju Cao , Wanke Xia , Bowen Zeng , Zeyu Zhang , Hao Tang

This paper presents a novel spatiotemporal transformer network that introduces several original components to detect actions in untrimmed videos. First, the multi-feature selective semantic attention model calculates the correlations…

计算机视觉与模式识别 · 计算机科学 2024-05-15 Matthew Korban , Peter Youngs , Scott T. Acton

Action detection and recognition tasks have been the target of much focus in the computer vision community due to their many applications, namely, security, robotics and recommendation systems. Recently, datasets like AVA, provide…

计算机视觉与模式识别 · 计算机科学 2019-07-31 João Antunes , Pedro Abreu , Alexandre Bernardino , Asim Smailagic , Daniel Siewiorek

This technical report analyzes an egocentric video action detection method we used in the 2021 EPIC-KITCHENS-100 competition hosted in CVPR2021 Workshop. The goal of our task is to locate the start time and the end time of the action in the…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Zhiwu Qing , Ziyuan Huang , Xiang Wang , Yutong Feng , Shiwei Zhang , Jianwen Jiang , Mingqian Tang , Changxin Gao , Marcelo H. Ang , Nong Sang

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

The quadratic complexity of standard attention mechanisms poses a significant scalability bottleneck for large language models (LLMs) in long-context scenarios. While hybrid attention strategies that combine sparse and full attention within…

计算与语言 · 计算机科学 2026-01-29 Zecheng Tang , Quantong Qiu , Yi Yang , Zhiyi Hong , Haiya Xiang , Kebin Liu , Qingqing Dang , Juntao Li , Min Zhang

Joint visual attention (JVA) provides informative cues on human behavior during social interactions. The ubiquity of egocentric eye-trackers and large-scale datasets on everyday interactions offer research opportunities in identifying JVA…

人机交互 · 计算机科学 2025-09-17 Kumushini Thennakoon , Yasasi Abeysinghe , Bhanuka Mahanama , Vikas Ashok , Sampath Jayarathna

The dominant paradigm in spatiotemporal action detection is to classify actions using spatiotemporal features learned by 2D or 3D Convolutional Networks. We argue that several actions are characterized by their context, such as relevant…

机器学习 · 计算机科学 2021-07-30 Michail Tsiaousis , Gertjan Burghouts , Fieke Hillerström , Peter van der Putten

Person re-identification (re-ID) in first-person (egocentric) vision is a fairly new and unexplored problem. With the increase of wearable video recording devices, egocentric data becomes readily available, and person re-identification has…

计算机视觉与模式识别 · 计算机科学 2021-03-09 Ankit Choudhary , Deepak Mishra , Arnab Karmakar

Skeleton-based action recognition is an important task that requires the adequate understanding of movement characteristics of a human action from the given skeleton sequence. Recent studies have shown that exploring spatial and temporal…

计算机视觉与模式识别 · 计算机科学 2019-04-01 Chenyang Si , Wentao Chen , Wei Wang , Liang Wang , Tieniu Tan

Despite the recent progress of fully-supervised action segmentation techniques, the performance is still not fully satisfactory. One main challenge is the problem of spatiotemporal variations (e.g. different people may perform the same…

计算机视觉与模式识别 · 计算机科学 2020-03-20 Min-Hung Chen , Baopu Li , Yingze Bao , Ghassan AlRegib , Zsolt Kira

Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a two-branch structure for action recognition, ie, one…

计算机视觉与模式识别 · 计算机科学 2020-02-11 Xiaohan Wang , Yu Wu , Linchao Zhu , Yi Yang

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we…

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

计算机视觉与模式识别 · 计算机科学 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Event boundaries play a crucial role as a pre-processing step for detection, localization, and recognition tasks of human activities in videos. Typically, although their intrinsic subjectiveness, temporal bounds are provided manually as…

计算机视觉与模式识别 · 计算机科学 2018-09-07 Alejandro Cartas , Estefania Talavera , Petia Radeva , Mariella Dimiccoli