中文
相关论文

相关论文: Transformed ROIs for Capturing Visual Transformati…

200 篇论文

Standard methods for video recognition use large CNNs designed to capture spatio-temporal data. However, training these models requires a large amount of labeled training data, containing a wide variety of actions, scenes, settings and…

计算机视觉与模式识别 · 计算机科学 2021-03-31 AJ Piergiovanni , Michael S. Ryoo

Wearable cameras are increasingly used as an observational and interventional tool for human behaviors by providing detailed visual data of hand-related activities. This data can be leveraged to facilitate memory recall for logging of…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Soroush Shahi , Farzad Shahabi , Rama Nabulsi , Glenn Fernandes , Aggelos Katsaggelos , Nabil Alshurafa

This work targets human action recognition in video. While recent methods typically represent actions by statistics of local video features, here we argue for the importance of a representation derived from human pose. To this end we…

计算机视觉与模式识别 · 计算机科学 2015-09-24 Guilhem Chéron , Ivan Laptev , Cordelia Schmid

Video object detection is challenging in the presence of appearance deterioration in certain video frames. Therefore, it is a natural choice to aggregate temporal information from other frames of the same video into the current frame.…

计算机视觉与模式识别 · 计算机科学 2021-09-14 Tao Gong , Kai Chen , Xinjiang Wang , Qi Chu , Feng Zhu , Dahua Lin , Nenghai Yu , Huamin Feng

In the field of action recognition, video clips are always treated as ordered frames for subsequent processing. To achieve spatio-temporal perception, existing approaches propose to embed adjacent temporal interaction in the convolutional…

计算机视觉与模式识别 · 计算机科学 2022-02-01 Rongchang Li , Xiao-Jun Wu , Tianyang Xu

In this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks. Instead of relying on hand-designed splitting…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Michael S. Ryoo , AJ Piergiovanni , Anurag Arnab , Mostafa Dehghani , Anelia Angelova

Human-object interactions (HOI) recognition and pose estimation are two closely related tasks. Human pose is an essential cue for recognizing actions and localizing the interacted objects. Meanwhile, human action and their interacted…

计算机视觉与模式识别 · 计算机科学 2019-03-18 Wei Feng , Wentao Liu , Tong Li , Jing Peng , Chen Qian , Xiaolin Hu

Hand-Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across frames. Although recent HOI reenactment approaches have…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Xuan Huang , Mochu Xiang , Zhelun Shen , Jinbo Wu , Chenming Wu , Chen Zhao , Kaisiyuan Wang , Hang Zhou , Shanshan Liu , Haocheng Feng , Wei He , Jingdong Wang

Motivated by the success of Transformers in natural language processing (NLP) tasks, there emerge some attempts (e.g., ViT and DeiT) to apply Transformers to the vision domain. However, pure Transformer architectures often require a large…

计算机视觉与模式识别 · 计算机科学 2021-04-21 Kun Yuan , Shaopeng Guo , Ziwei Liu , Aojun Zhou , Fengwei Yu , Wei Wu

This paper strives to localize the temporal extent of an action in a long untrimmed video. Where existing work leverages many examples with their start, their ending, and/or the class of the action during training time, we propose few-shot…

计算机视觉与模式识别 · 计算机科学 2020-08-26 Pengwan Yang , Vincent Tao Hu , Pascal Mettes , Cees G. M. Snoek

The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification. To this end, we introduce a…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Chuhan Zhang , Ankush Gupta , Andrew Zisserman

AI-assisted lesion detection models play a crucial role in the early screening of cancer. However, previous image-based models ignore the inter-frame contextual information present in videos. On the other hand, video-based models capture…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Yuncheng Jiang , Zixun Zhang , Jun Wei , Chun-Mei Feng , Guanbin Li , Xiang Wan , Shuguang Cui , Zhen Li

Recent studies have demonstrated the power of recurrent neural networks for machine translation, image captioning and speech recognition. For the task of capturing temporal structure in video, however, there still remain numerous open…

计算机视觉与模式识别 · 计算机科学 2016-02-11 Lionel Pigou , Aäron van den Oord , Sander Dieleman , Mieke Van Herreweghe , Joni Dambre

Humans have the natural ability to recognize actions even if the objects involved in the action or the background are changed. Humans can abstract away the action from the appearance of the objects which is referred to as compositionality…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Ramanathan Rajendiran , Debaditya Roy , Basura Fernando

Detecting human-object interactions (HOI) is an important step toward a comprehensive visual understanding of machines. While detecting non-temporal HOIs (e.g., sitting on a chair) from static images is feasible, it is unlikely even for…

计算机视觉与模式识别 · 计算机科学 2021-06-25 Meng-Jiun Chiou , Chun-Yu Liao , Li-Wei Wang , Roger Zimmermann , Jiashi Feng

The interaction decoder utilized in prevalent Transformer-based HOI detectors typically accepts pre-composed human-object pairs as inputs. Though achieving remarkable performance, such paradigm lacks feasibility and cannot explore novel…

计算机视觉与模式识别 · 计算机科学 2023-11-17 Liulei Li , Jianan Wei , Wenguan Wang , Yi Yang

Convolutional Neural Networks (CNNs) are known to be brittle under various image transformations, including rotations, scalings, and changes of lighting conditions. We observe that the features of a transformed image are drastically…

计算机视觉与模式识别 · 计算机科学 2020-04-14 Shaohua Li , Xiuchao Sui , Jie Fu , Yong Liu , Rick Siow Mong Goh

Animals (especially humans) have an amazing ability to learn new tasks quickly, and switch between them flexibly. How brains support this ability is largely unknown, both neuroscientifically and algorithmically. One reasonable supposition…

机器学习 · 计算机科学 2017-06-23 Kevin T. Feigelis , Daniel L. K. Yamins

Tracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose a solution named TransMOT, which leverages powerful graph transformers to efficiently model the spatial and…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Peng Chu , Jiang Wang , Quanzeng You , Haibin Ling , Zicheng Liu

Human-Object Interaction (HOI) detection is the task of identifying a set of <human, object, interaction> triplets from an image. Recent work proposed transformer encoder-decoder architectures that successfully eliminated the need for many…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Bumsoo Kim , Jonghwan Mun , Kyoung-Woon On , Minchul Shin , Junhyun Lee , Eun-Sol Kim