English
Related papers

Related papers: Video Relationship Reasoning using Gated Spatio-Te…

200 papers

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understanding. Existing spatio-temporal methods primarily focus on the…

Artificial Intelligence · Computer Science 2025-10-14 Wentao Wang , Heqing Zou , Tianze Luo , Rui Huang , Yutian Zhao , Zhuochen Wang , Hansheng Zhang , Chengwei Qin , Yan Wang , Lin Zhao , Huaijian Zhang

The task of video grounding, which temporally localizes a natural language description in a video, plays an important role in understanding videos. Existing studies have adopted strategies of sliding window over the entire video or…

Computer Vision and Pattern Recognition · Computer Science 2019-01-23 Dongliang He , Xiang Zhao , Jizhou Huang , Fu Li , Xiao Liu , Shilei Wen

In this work we introduce a time- and memory-efficient method for structured prediction that couples neuron decisions across both space at time. We show that we are able to perform exact and efficient inference on a densely connected…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Siddhartha Chandra , Camille Couprie , Iasonas Kokkinos

Advances in GPS telemetry technology have enabled analysis of animal movement in open areas. Ecologists today are utilizing modern analytic tools to study animal behaviors from large quantity of GPS coordinates. Analytic tools with…

Human-Computer Interaction · Computer Science 2020-01-31 Wei Li , Mathias Funk , Jasper Eikelboom , Aarnout Brombacher

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Mingfang Zhang , Jingjing Pan , Ashutosh Kumar , Rajat Saini , Mustafa Erdogan , Hsuan-Kung Yang , Caixin Kang , Yifei Huang , Yoichi Sato , Quan Kong

Spatio-temporal information is key to resolve occlusion and depth ambiguity in 3D pose estimation. Previous methods have focused on either temporal contexts or local-to-global architectures that embed fixed-length spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2020-10-21 Junfa Liu , Juan Rojas , Zhijun Liang , Yihui Li , Yisheng Guan

Building a robot that can understand and learn to interact by watching humans has inspired several vision problems. However, despite some successful results on static datasets, it remains unclear how current models can be used on a robot…

Robotics · Computer Science 2023-04-18 Shikhar Bahl , Russell Mendonca , Lili Chen , Unnat Jain , Deepak Pathak

We present the task of Spatio-Temporal Video Question Answering, which requires intelligent systems to simultaneously retrieve relevant moments and detect referenced visual concepts (people and objects) to answer natural language questions…

Computer Vision and Pattern Recognition · Computer Science 2020-05-13 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

Humans are able to perceive, understand and reason about causal events. Developing models with similar physical and causal understanding capabilities is a long-standing goal of artificial intelligence. As a step towards this direction, we…

Artificial Intelligence · Computer Science 2022-03-02 Tayfun Ates , M. Samil Atesoglu , Cagatay Yigit , Ilker Kesen , Mert Kobas , Erkut Erdem , Aykut Erdem , Tilbe Goksun , Deniz Yuret

Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently requires visual reasoning on top of the relationships among…

Computer Vision and Pattern Recognition · Computer Science 2019-09-19 Sibei Yang , Guanbin Li , Yizhou Yu

We propose a novel model to address the task of Visual Dialog which exhibits complex dialog structures. To obtain a reasonable answer based on the current question and the dialog history, the underlying semantic dependencies between dialog…

Computer Vision and Pattern Recognition · Computer Science 2019-05-30 Zilong Zheng , Wenguan Wang , Siyuan Qi , Song-Chun Zhu

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Baoxiong Jia , Ting Lei , Song-Chun Zhu , Siyuan Huang

Dense video captioning is an extremely challenging task since accurate and coherent description of events in a video requires holistic understanding of video contents as well as contextual reasoning of individual events. Most existing…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Jonghwan Mun , Linjie Yang , Zhou Ren , Ning Xu , Bohyung Han

Dramatic progress has been made in animating individual characters. However, we still lack automatic control over activities between characters, especially those involving interactions. In this paper, we present a novel energy-based…

Graphics · Computer Science 2022-03-10 Yizhou Zhao , Liang Qiu , Wensi Ai , Pan Lu , Song-Chun Zhu

Social relationships (e.g., friends, couple etc.) form the basis of the social network in our daily life. Automatically interpreting such relationships bears a great potential for the intelligent systems to understand human behavior in…

Computer Vision and Pattern Recognition · Computer Science 2018-07-03 Zhouxia Wang , Tianshui Chen , Jimmy Ren , Weihao Yu , Hui Cheng , Liang Lin

Anticipating human motion in crowded scenarios is essential for developing intelligent transportation systems, social-aware robots and advanced video surveillance applications. A key component of this task is represented by the inherently…

Computer Vision and Pattern Recognition · Computer Science 2021-07-09 Alessia Bertugli , Simone Calderara , Pasquale Coscia , Lamberto Ballan , Rita Cucchiara

Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Oleh Kolner , Thomas Ortner , Stanisław Woźniak , Angeliki Pantazi

We propose a novel approach for modeling semantic contextual relationships in videos. This graph-based model enables the learning and propagation of higher-level spatial-temporal contexts to facilitate the semantic labeling of local…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Tinghuai Wang , Huiling Wang

Predicting an interaction before it is fully executed is very important in applications such as human-robot interaction and video surveillance. In a two-human interaction scenario, there often contextual dependency structure between the…

Computer Vision and Pattern Recognition · Computer Science 2018-06-13 Qiuhong Ke , Mohammed Bennamoun , Senjian An , Farid Bossaid , Ferdous Sohel

Video question answering (VideoQA) is challenging as it requires modeling capacity to distill dynamic visual artifacts and distant relations and to associate them with linguistic concepts. We introduce a general-purpose reusable neural unit…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Thao Minh Le , Vuong Le , Svetha Venkatesh , Truyen Tran