English
Related papers

Related papers: Modelling Spatio-Temporal Interactions For Composi…

200 papers

Understanding and reasoning about objects' physical properties in the natural world is a fundamental challenge in artificial intelligence. While some properties like colors and shapes can be directly observed, others, such as mass and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Zhenfang Chen , Shilong Dong , Kexin Yi , Yunzhu Li , Mingyu Ding , Antonio Torralba , Joshua B. Tenenbaum , Chuang Gan

Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness should be understood through the compositional structure of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zijie Wang , Wei Zhang , Weiming Zhang , Fanqi Zhang , Xiao Tan , Yipeng Qin , Guanbin Li

Human activity videos involve rich, varied interactions between people and objects. In this paper we develop methods for generating such videos -- making progress toward addressing the important, open problem of video generation in complex…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Megha Nawhal , Mengyao Zhai , Andreas Lehrmann , Leonid Sigal , Greg Mori

As the success of deep models has led to their deployment in all areas of computer vision, it is increasingly important to understand how these representations work and what they are capturing. In this paper, we shed light on deep…

Computer Vision and Pattern Recognition · Computer Science 2018-01-08 Christoph Feichtenhofer , Axel Pinz , Richard P. Wildes , Andrew Zisserman

In this paper, a novel human action recognition technique from video is presented. Any action of human is a combination of several micro action sequences performed by one or more body parts of the human. The proposed approach uses…

Computer Vision and Pattern Recognition · Computer Science 2015-10-16 Satyabrata Maity , Debotosh Bhattacharjee , Amlan Chakrabarti

Learning structured representations of visual scenes is currently a major bottleneck to bridging perception with reasoning. While there has been exciting progress with slot-based models, which learn to segment scenes into sets of objects,…

Machine Learning · Computer Science 2021-07-26 James C. R. Whittington , Rishabh Kabra , Loic Matthey , Christopher P. Burgess , Alexander Lerchner

In this work we propose to utilize information about human actions to improve pose estimation in monocular videos. To this end, we present a pictorial structure model that exploits high-level information about activities to incorporate…

Computer Vision and Pattern Recognition · Computer Science 2017-02-13 Umar Iqbal , Martin Garbade , Juergen Gall

The design of a complex system warrants a compositional methodology, i.e., composing simple components to obtain a larger system that exhibits their collective behavior in a meaningful way. We propose an automaton-based paradigm for…

Logic in Computer Science · Computer Science 2023-02-03 Tobias Kappé , Farhad Arbab , Carolyn Talcott

In this paper, we propose a new approach to under-stand actions in egocentric videos that exploits the semantics of object interactions at both frame and temporal levels. At the frame level, we use a region-based approach that takes as…

Computer Vision and Pattern Recognition · Computer Science 2021-04-26 Alejandro Cartas , Petia Radeva , Mariella Dimiccoli

While Human-Object Interaction(HOI) Detection has achieved tremendous advances in recent, it still remains challenging due to complex interactions with multiple humans and objects occurring in images, which would inevitably lead to…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Kunlun Xu , Zhimin Li , Zhijun Zhang , Leizhen Dong , Wenhui Xu , Luxin Yan , Sheng Zhong , Xu Zou

We present a new architecture for human action forecasting from videos. A temporal recurrent encoder captures temporal information of input videos while a self-attention model is used to attend on relevant feature dimensions of the input…

Computer Vision and Pattern Recognition · Computer Science 2021-07-20 Yan Bin Ng , Basura Fernando

Compositional generalization is a basic and essential intellective capability of human beings, which allows us to recombine known parts readily. However, existing neural network based models have been proven to be extremely deficient in…

Artificial Intelligence · Computer Science 2020-10-27 Qian Liu , Shengnan An , Jian-Guang Lou , Bei Chen , Zeqi Lin , Yan Gao , Bin Zhou , Nanning Zheng , Dongmei Zhang

Video understanding is to recognize and classify different actions or activities appearing in the video. A lot of previous work, such as video captioning, has shown promising performance in producing general video understanding. However, it…

Computer Vision and Pattern Recognition · Computer Science 2023-11-23 Zijian Kuang , Xinran Tie

People often use physical intuition when manipulating articulated objects, irrespective of object semantics. Motivated by this observation, we identify an important embodied task where an agent must play with objects to recover their parts.…

Computer Vision and Pattern Recognition · Computer Science 2021-05-04 Samir Yitzhak Gadre , Kiana Ehsani , Shuran Song

Analyzing the interactions between humans and objects from a video includes identification of the relationships between humans and the objects present in the video. It can be thought of as a specialized version of Visual Relationship…

Computer Vision and Pattern Recognition · Computer Science 2020-12-18 Sai Praneeth Reddy Sunkesula , Rishabh Dabral , Ganesh Ramakrishnan

Interactive object understanding, or what we can do to objects and how is a long-standing goal of computer vision. In this paper, we tackle this problem through observation of human hands in in-the-wild egocentric videos. We demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2022-04-11 Mohit Goyal , Sahil Modi , Rishabh Goyal , Saurabh Gupta

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations - sparsity of action…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Dibyadip Chatterjee , Fadime Sener , Shugao Ma , Angela Yao

We introduce a novel self-supervised learning approach to learn representations of videos that are responsive to changes in the motion dynamics. Our representations can be learned from data without human annotation and provide a substantial…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Simon Jenni , Givi Meishvili , Paolo Favaro

Human action analysis and understanding in videos is an important and challenging task. Although substantial progress has been made in past years, the explainability of existing methods is still limited. In this work, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2019-08-29 Tao Zhuo , Zhiyong Cheng , Peng Zhang , Yongkang Wong , Mohan Kankanhalli

Image-to-video adaptation seeks to efficiently adapt image models for use in the video domain. Instead of finetuning the entire image backbone, many image-to-video adaptation paradigms use lightweight adapters for temporal modeling on top…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Rui Qian , Shuangrui Ding , Dahua Lin
‹ Prev 1 8 9 10 Next ›