English
Related papers

Related papers: LangDriveCTRL: Natural Language Controllable Drivi…

200 papers

Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recent generative…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Kaicong Huang , Talha Azfar , Weisong Shi , Ruimin Ke

We present WorldCanvas, a framework for promptable world events that enables rich, user-directed simulation by combining text, trajectories, and reference images. Unlike text-only approaches and existing trajectory-controlled image-to-video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Hanlin Wang , Hao Ouyang , Qiuyu Wang , Yue Yu , Yihao Meng , Wen Wang , Ka Leong Cheng , Shuailei Ma , Qingyan Bai , Yixuan Li , Cheng Chen , Yanhong Zeng , Xing Zhu , Yujun Shen , Qifeng Chen

Mobile robot navigation systems are increasingly relied upon in dynamic and complex environments, yet they often struggle with map inaccuracies and the resulting inefficient path planning. This paper presents MRHaD, a Mixed Reality-based…

Robotics · Computer Science 2025-07-29 Takumi Taki , Masato Kobayashi , Eduardo Iglesius , Naoya Chiba , Shizuka Shirai , Yuki Uranishi

Human drivers can recognise fast abnormal driving situations to avoid accidents. Similar to humans, automated vehicles are supposed to perform anomaly detection. In this work, we propose the spatio-temporal graph auto-encoder for learning…

Robotics · Computer Science 2021-10-29 Julian Wiederer , Arij Bouazizi , Marco Troina , Ulrich Kressel , Vasileios Belagiannis

This paper presents a system for procedurally generating agent-based narratives using large language models (LLMs). Users could drag and drop multiple agents and objects into a scene, with each entity automatically assigned semantic…

Graphics · Computer Science 2025-12-24 Vinayak Regmi , Christos Mousas

We present ChronoDreamer, an action-conditioned world model for contact-rich robotic manipulation. Given a history of egocentric RGB frames, contact maps, actions, and joint states, ChronoDreamer predicts future video frames, contact…

Artificial Intelligence · Computer Science 2025-12-23 Zhenhao Zhou , Dan Negrut

Decision-making for urban autonomous driving is challenging due to the stochastic nature of interactive traffic participants and the complexity of road structures. Although reinforcement learning (RL)-based decision-making scheme is…

Machine Learning · Computer Science 2023-08-28 Haochen Liu , Zhiyu Huang , Xiaoyu Mo , Chen Lv

Recent progress in 4D representations, such as Dynamic NeRF and 4D Gaussian Splatting (4DGS), has enabled dynamic 4D scene reconstruction. However, text-driven 4D scene editing remains under-explored due to the challenge of ensuring both…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Dong In Lee , Hyungjun Doh , Seunggeun Chi , Runlin Duan , Sangpil Kim , Karthik Ramani

Recent diffusion models have achieved remarkable success in image relighting, and this success has quickly been extended to video relighting. However, existing methods offer limited explicit control over illumination in the relighted…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yizuo Peng , Xuelin Chen , Kai Zhang , Xiaodong Cun

Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Joonghyuk Shin , Daehyeon Choi , Jaesik Park

We consider the problem of generating realistic traffic scenes automatically. Existing methods typically insert actors into the scene according to a set of hand-crafted heuristics and are limited in their ability to model the true…

Computer Vision and Pattern Recognition · Computer Science 2021-01-19 Shuhan Tan , Kelvin Wong , Shenlong Wang , Sivabalan Manivasagam , Mengye Ren , Raquel Urtasun

Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current methods aggregate high-definition (HD) maps and 3D bounding boxes…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Zhechao Wang , Yiming Zeng , Lufan Ma , Zeqing Fu , Chen Bai , Ziyao Lin , Cheng Lu

Vision language navigation is the task that requires an agent to navigate through a 3D environment based on natural language instructions. One key challenge in this task is to ground instructions with the current visual information that the…

Computation and Language · Computer Science 2021-04-21 Jialu Li , Hao Tan , Mohit Bansal

Service robots are increasingly deployed in diverse and dynamic environments, where both physical layouts and social contexts change over time and across locations. In these unstructured settings, conventional navigation systems that rely…

Robotics · Computer Science 2025-07-16 Yanbo Wang , Zipeng Fang , Lei Zhao , Weidong Chen

Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Implicit textual prompts lack precision, while explicit trajectory conditioning imposes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Songlin Yang , Zhe Wang , Xuyi Yang , Songchun Zhang , Xianghao Kong , Taiyi Wu , Xiaotong Zhao , Ran Zhang , Alan Zhao , Anyi Rao

In robotic, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or videos may offer overly detailed specifications. To tackle these…

Widespread adoption of self-driving cars will depend not only on their safety but largely on their ability to interact with human users. Just like human drivers, self-driving cars will be expected to understand and safely follow…

Robotics · Computer Science 2019-10-18 Junha Roh , Chris Paxton , Andrzej Pronobis , Ali Farhadi , Dieter Fox

Existing video-language models can generate factual descriptions of road events but lack control over how these events are expressed: their tone, urgency, or style. This limits deployment in communication-critical settings where the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Chirag Parikh , Siddhi Pravin Lipare , Ravi Kiran Sarvadevabhatla

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

Over the past few years there is a growing interest in the learning-based self driving system. To ensure safety, such systems are first developed and validated in simulators before being deployed in the real world. However, most of the…

Robotics · Computer Science 2021-03-15 Quanyi Li , Zhenghao Peng , Qihang Zhang , Chunxiao Liu , Bolei Zhou