中文
相关论文

相关论文: Video Relationship Reasoning using Gated Spatio-Te…

200 篇论文

As the intermediate-level representations bridging the two levels, structured representations of visual scenes, such as visual relationships between pairwise objects, have been shown to not only benefit compositional models in learning to…

计算机视觉与模式识别 · 计算机科学 2022-07-12 Meng-Jiun Chiou

Temporal relational reasoning, the ability to link meaningful transformations of objects or entities over time, is a fundamental property of intelligent species. In this paper, we introduce an effective and interpretable network module, the…

计算机视觉与模式识别 · 计算机科学 2018-07-26 Bolei Zhou , Alex Andonian , Aude Oliva , Antonio Torralba

Graph based representation has been widely used in modelling spatio-temporal relationships in video understanding. Although effective, existing graph-based approaches focus on capturing the human-object relationships while ignoring…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Chinthani Sugandhika , Chen Li , Deepu Rajan , Basura Fernando

This paper presents a novel method, termed Bridge to Answer, to infer correct answers for questions about a given video by leveraging adequate graph interactions of heterogeneous crossmodal graphs. To realize this, we learn question…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Jungin Park , Jiyoung Lee , Kwanghoon Sohn

Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently requires modeling temporal dynamics and evolving visual…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Zhaochong An , Zirui Li , Mingqiao Ye , Feng Qiao , Jiaang Li , Zongwei Wu , Vishal Thengane , Chengzu Li , Lei Li , Luc Van Gool , Guolei Sun , Serge Belongie

Understanding a scene by decoding the visual relationships depicted in an image has been a long studied problem. While the recent advances in deep learning and the usage of deep neural networks have achieved near human accuracy on many…

计算机视觉与模式识别 · 计算机科学 2020-05-19 Aniket Agarwal , Ayush Mangal , Vipul

Vision-and-Language Navigation (VLN) requires an agent to navigate in a real-world environment following natural language instructions. From both the textual and visual perspectives, we find that the relationships among the scene, its…

计算机视觉与模式识别 · 计算机科学 2020-12-29 Yicong Hong , Cristian Rodriguez-Opazo , Yuankai Qi , Qi Wu , Stephen Gould

Text-based Visual Question Answering~(TextVQA) aims to produce correct answers for given questions about the images with multiple scene texts. In most cases, the texts naturally attach to the surface of the objects. Therefore, spatial…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Hao Li , Jinfa Huang , Peng Jin , Guoli Song , Qi Wu , Jie Chen

Understanding relations between objects is crucial for understanding the semantics of a visual scene. It is also an essential step in order to bridge visual and language models. However, current state-of-the-art computer vision models still…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Palaash Agrawal , Haidi Azaman , Cheston Tan

Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ermanno Bartoli , Dennis Rotondi , Buwei He , Patric Jensfelt , Kai O. Arras , Iolanda Leite

Identifying relations between objects is central to understanding the scene. While several works have been proposed for relation modeling in the image domain, there have been many constraints in the video domain due to challenging dynamics…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Sangmin Woo , Junhyug Noh , Kangil Kim

We addressed the challenging task of video question answering, which requires machines to answer questions about videos in a natural language form. Previous state-of-the-art methods attempt to apply spatio-temporal attention mechanism on…

计算机视觉与模式识别 · 计算机科学 2020-08-21 Deng Huang , Peihao Chen , Runhao Zeng , Qing Du , Mingkui Tan , Chuang Gan

Recent approaches on visual scene understanding attempt to build a scene graph -- a computational representation of objects and their pairwise relationships. Such rich semantic representation is very appealing, yet difficult to obtain from…

计算机视觉与模式识别 · 计算机科学 2018-11-08 Paul Gay , Stuart James , Alessio Del Bue

Consider the scenario where a human cleans a table and a robot observing the scene is instructed with the task "Remove the cloth using which I wiped the table". Instruction following with temporal reasoning requires the robot to identify…

机器人学 · 计算机科学 2024-10-11 Riya Arora , Niveditha Narendranath , Aman Tambi , Sandeep S. Zachariah , Souvik Chakraborty , Rohan Paul

Temporal Action Detection (TAD), the task of localizing and classifying actions in untrimmed video, remains challenging due to action overlaps and variable action durations. Recent findings suggest that TAD performance is dependent on the…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Aglind Reka , Diana Laura Borza , Dominick Reilly , Michal Balazia , Francois Bremond

The growing interest in embodied agents increases the demand for spatiotemporal video understanding, yet existing benchmarks largely emphasize extractive reasoning, where answers can be explicitly presented within spatiotemporal events. It…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Seunghwan Bang , Hwanjun Song

Visual dialog is a challenging task that requires the comprehension of the semantic dependencies among implicit visual and textual contexts. This task can refer to the relation inference in a graphical model with sparse contexts and unknown…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Dan Guo , Hui Wang , Hanwang Zhang , Zheng-Jun Zha , Meng Wang

Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper addresses two key challenges: (1) conventional…

计算机视觉与模式识别 · 计算机科学 2025-05-15 Dayong Liang , Changmeng Zheng , Zhiyuan Wen , Yi Cai , Xiao-Yong Wei , Qing Li

Identifying different objects (man and cup) is an important problem on its own, but identifying the relationship between them (holding) is critical for many real world use cases. This paper describes an approach to reduce a visual…

计算机视觉与模式识别 · 计算机科学 2018-09-27 Toshiyuki Fukuzawa

We present Language-binding Object Graph Network, the first neural reasoning method with dynamic relational structures across both visual and textual domains with applications in visual question answering. Relaxing the common assumption…

计算机视觉与模式识别 · 计算机科学 2021-02-19 Thao Minh Le , Vuong Le , Svetha Venkatesh , Truyen Tran