English
Related papers

Related papers: SGTA: Scene-Graph Based Multi-Modal Traffic Agent …

200 papers

The scene graph generation (SGG) task aims to detect visual relationship triplets, i.e., subject, predicate, object, in an image, providing a structural vision layout for scene understanding. However, current models are stuck in common…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Yuyu Guo , Lianli Gao , Xuanhan Wang , Yuxuan Hu , Xing Xu , Xu Lu , Heng Tao Shen , Jingkuan Song

Answering questions that require reading texts in an image is challenging for current models. One key difficulty of this task is that rare, polysemous, and ambiguous words frequently appear in images, e.g., names of places, products, and…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Difei Gao , Ke Li , Ruiping Wang , Shiguang Shan , Xilin Chen

3D scene graphs provide a structured representation of object entities and their relationships, enabling high-level interpretation and reasoning for robots while remaining intuitively understandable to humans. Existing approaches for 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Zirui Wang , Ruiping Liu , Yufan Chen , Junwei Zheng , Weijia Fan , Kunyu Peng , Di Wen , Jiale Wei , Jiaming Zhang , Rainer Stiefelhagen

For the classification of traffic scenes, a description model is necessary that can describe the scene in a uniform way, independent of its domain. A model to describe a traffic scene in a semantic way is described in this paper. The…

Machine Learning · Computer Science 2022-06-30 Maximilian Zipfl , J. Marius Zöllner

This paper introduces the schemes of Team LingJing's experiments in NLPCC-2022-Shared-Task-4 Multi-modal Dialogue Understanding and Generation (MDUG). The MDUG task can be divided into two phases: multi-modal context understanding and…

Computation and Language · Computer Science 2022-07-06 Bin Li , Yixuan Weng , Ziyu Ma , Bin Sun , Shutao Li

The complex driving environment brings great challenges to the visual perception of autonomous vehicles. It's essential to extract clear and explainable information from the complex road and traffic scenarios and offer clues to decision and…

Computer Vision and Pattern Recognition · Computer Science 2022-06-03 Yiyue Zhao , Xinyu Yun , Chen Chai , Zhiyu Liu , Wenxuan Fan , Xiao Luo

Traffic anomaly detection (TAD) in driving videos is critical for ensuring the safety of autonomous driving and advanced driver assistance systems. Previous single-stage TAD methods primarily rely on frame prediction, making them vulnerable…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Rongqin Liang , Yuanman Li , Jiantao Zhou , Xia Li

Predicting the future trajectories of multiple interacting agents in a scene has become an increasingly important problem for many different applications ranging from control of autonomous vehicles and social robots to security and…

Computer Vision and Pattern Recognition · Computer Science 2019-07-18 Vineet Kosaraju , Amir Sadeghian , Roberto Martín-Martín , Ian Reid , S. Hamid Rezatofighi , Silvio Savarese

Global and local relational reasoning enable scene understanding models to perform human-like scene analysis and understanding. Scene understanding enables better semantic segmentation and object-to-object interaction detection. In the…

Image and Video Processing · Electrical Eng. & Systems 2022-01-31 Lalithkumar Seenivasan , Sai Mitheran , Mobarakol Islam , Hongliang Ren

Objects in a scene are not always related. The execution efficiency of the one-stage scene graph generation approaches are quite high, which infer the effective relation between entity pairs using sparse proposal sets and a few queries.…

Computer Vision and Pattern Recognition · Computer Science 2022-12-20 Yuxiang Zhang , Zhenbo Liu , Shuai Wang

Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control signals at inference time to guide the generative model…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Mirlan Karimov , Teodora Spasojevic , Markus Braun , Julian Wiederer , Vasileios Belagiannis , Marc Pollefeys

Action recognition is a fundamental task in video understanding. Existing methods typically extract unified features to process all actions in one video, which makes it challenging to model the interactions between different objects in…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Tianci Wu , Guangming Zhu , Jiang Lu , Siyuan Wang , Ning Wang , Nuoye Xiong , Zhang Liang

Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approaches often ignore the fact that videos are essentially…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Anoop Cherian , Chiori Hori , Tim K. Marks , Jonathan Le Roux

Traffic prediction has gradually attracted the attention of researchers because of the increase in traffic big data. Therefore, how to mine the complex spatio-temporal correlations in traffic data to predict traffic conditions more…

Machine Learning · Computer Science 2021-12-07 Yuchen Fang , Yanjun Qin , Haiyong Luo , Fang Zhao , Chenxing Wang

In high-conflict mixed-traffic scenarios involving human-driven and autonomous vehicles, most existing autonomous driving systems default to overly conservative behaviors, lack proactive interaction, and consequently suffer from limited…

Robotics · Computer Science 2026-04-28 Xinwei Dong , Jiyang Li , Jiabin Xie , Yang Yi , Tianshang Jia , Shiyu Fang , Ye Tian , Peng Hang

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Traffic Accident Detection (TAD) and Traffic Accident Understanding (TAU). However, existing studies mainly focus on describing and interpreting accident videos,…

Computation and Language · Computer Science 2026-04-24 Zijin Zhou , Songan Zhang

There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Yunnan Wang , Ziqiang Li , Zequn Zhang , Wenyao Zhang , Baao Xie , Xihui Liu , Wenjun Zeng , Xin Jin

Multimodal emotion recognition (MER) is crucial for enabling emotionally intelligent systems that perceive and respond to human emotions. However, existing methods suffer from limited cross-modal interaction and imbalanced contributions…

Multimedia · Computer Science 2025-07-30 Zeyu Deng , Yanhui Lu , Jiashu Liao , Shuang Wu , Chongfeng Wei

Most existing traffic sign-related works are dedicated to detecting and recognizing part of traffic signs individually, which fails to analyze the global semantic logic among signs and may convey inaccurate traffic instruction. Following…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Chuang Yang , Kai Zhuang , Mulin Chen , Haozhao Ma , Xu Han , Tao Han , Changxing Guo , Han Han , Bingxuan Zhao , Qi Wang

In this paper, we address the problem of inferring the layout of complex road scenes from video sequences. To this end, we formulate it as a top-view road attributes prediction problem and our goal is to predict these attributes for each…

Computer Vision and Pattern Recognition · Computer Science 2020-07-03 Buyu Liu , Bingbing Zhuang , Samuel Schulter , Pan Ji , Manmohan Chandraker