English
Related papers

Related papers: Action Emergence from Streaming Intent

200 papers

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Haomiao Xiong , Zongxin Yang , Jiazuo Yu , Yunzhi Zhuge , Lu Zhang , Jiawen Zhu , Huchuan Lu

Recent advances in diffusion transformers have empowered video generation models to generate high-quality video clips from texts or images. However, world models with the ability to predict long-horizon futures from past observations and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yixuan Zhu , Jiaqi Feng , Wenzhao Zheng , Yuan Gao , Xin Tao , Pengfei Wan , Jie Zhou , Jiwen Lu

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiamin Wang , Yichen Yao , Xiang Feng , Hang Wu , Yaming Wang , Qingqiu Huang , Yuexin Ma , Xinge Zhu

Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Models (Video-LLMs)…

Using generative models to synthesize new data has become a de-facto standard in autonomous driving to address the data scarcity issue. Though existing approaches are able to boost perception models, we discover that these approaches fail…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Enhui Ma , Lijun Zhou , Tao Tang , Zhan Zhang , Dong Han , Junpeng Jiang , Kun Zhan , Peng Jia , Xianpeng Lang , Haiyang Sun , Di Lin , Kaicheng Yu

Understanding the short-term motion of vulnerable road users (VRUs) like pedestrians and cyclists is critical for safe autonomous driving, especially in urban scenarios with ambiguous or high-risk behaviors. While vision-language models…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Mihir Godbole , Xiangbo Gao , Zhengzhong Tu

Lane segment topology reasoning constructs a comprehensive road network by capturing the topological relationships between lane segments and their semantic types. This enables end-to-end autonomous driving systems to perform road-dependent…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Yiming Yang , Yueru Luo , Bingkun He , Hongbin Lin , Suzhong Fu , Chao Zheng , Zhipeng Cao , Erlong Li , Chao Yan , Shuguang Cui , Zhen Li

In this work, we aim to achieve efficient end-to-end learning of driving policies in dynamic multi-agent environments. Predicting and anticipating future events at the object level are critical for making informed driving decisions. We…

Robotics · Computer Science 2021-01-18 Jinkun Cao , Xin Wang , Trevor Darrell , Fisher Yu

Despite significant progress in Visual-Language-Action (VLA), in highly complex and dynamic environments that involve real-time unpredictable interactions (such as 3D open worlds and large-scale PvP games), existing approaches remain…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Zheyuan Zhou , Liang Du , Zixun Sun , Xiaoyu Zhou , Ruimin Ye , Qihao Chen , Yinda Chen , Lemiao Qiu

Large Vision Language Models (LVLMs) exhibit strong Chain-of-Thought (CoT) capabilities, yet most existing paradigms assume full-video availability before inference, a batch-style process misaligned with real-world video streams where…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jialiang Zhang , Junlong Tong , Junyan Lin , Hao Wu , Yirong Sun , Yunpu Ma , Xiaoyu Shen

Recent progress in video large language models (Video-LLMs) has enabled strong offline reasoning over long and complex videos. However, real-world deployments increasingly require streaming perception and proactive interaction, where video…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Junho Kim , Hosu Lee , James M. Rehg , Minsu Kim , Yong Man Ro

The growing complexity of networks and the variety of future scenarios with diverse and often stringent performance requirements call for a higher level of automation. Intent-based management emerges as a solution to attain high level of…

Networking and Internet Architecture · Computer Science 2024-07-26 Erciyes Karakaya , Ozgur Ercetin , Huseyin Ozkan , Mehmet Karaca , Elham Dehghan Biyar , Alexandros Palaios

Future spacecraft operations require autonomy that can interpret high-level mission intent while preserving safety. However, existing trajectory optimization still relies heavily on expert-crafted formulations and does not support…

Systems and Control · Electrical Eng. & Systems 2026-05-29 Yuji Takubo , Simone D'Amico

Since the emergence of autonomous driving technology, it has advanced rapidly over the past decade. It is becoming increasingly likely that autonomous vehicles (AVs) would soon coexist with human-driven vehicles (HVs) on the roads.…

Robotics · Computer Science 2025-04-29 Jing Wang , Yan Jin , Hamid Taghavifar , Fei Ding , Chongfeng Wei

Real-time perception, or streaming perception, is a crucial aspect of autonomous driving that has yet to be thoroughly explored in existing research. To address this gap, we present DAMO-StreamNet, an optimized framework that combines…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Jun-Yan He , Zhi-Qi Cheng , Chenyang Li , Wangmeng Xiang , Binghui Chen , Bin Luo , Yifeng Geng , Xuansong Xie

Pedestrian intention prediction needs to be accurate for autonomous vehicles to navigate safely in urban environments. We present a lightweight, socially informed architecture for pedestrian intention prediction. It fuses four behavioral…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Sima Ashayer , Hoang H. Nguyen , Yu Liang , Mina Sartipi

Trajectory prediction is an essential component in autonomous driving, particularly for collision avoidance systems. Considering the inherent uncertainty of the task, numerous studies have utilized generative models to produce multiple…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Chen Liu , Shibo He , Haoyu Liu , Jiming Chen

A multi-modal framework to generate user intention distributions when operating a mobile vehicle is proposed in this work. The model learns from past observed trajectories and leverages traversability information derived from the visual…

Robotics · Computer Science 2022-03-17 Kavindie Katuwandeniya , Stefan H. Kiss , Lei Shi , Jaime Valls Miro

Predicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However, current best-performing single-modality methods or multi-modality fusion perception methods are only able to predict…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Yining Shi , Kun Jiang , Ke Wang , Jiusi Li , Yunlong Wang , Mengmeng Yang , Diange Yang

Current LLM agents operate under an implicit but universal assumption: execution is a transaction -- the user submits a request, the agent works in isolation, and only upon completion does the dialogue resume. This forces users into a…

Machine Learning · Computer Science 2026-04-28 Zhiyuan Zhai , Ming Li , Xin Wang