English
Related papers

Related papers: ScenarioControl: Vision-Language Controllable Vect…

200 papers

Video generation has achieved impressive quality, but it still suffers from artifacts such as temporal inconsistency and violation of physical laws. Leveraging 3D scenes can fundamentally resolve these issues by providing precise control…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Zhaofang Qian , Abolfazl Sharifi , Tucker Carroll , Ser-Nam Lim

Recent advances in deep learning have enabled the development of autonomous systems that use deep neural networks for perception. Formal verification of these systems is challenging due to the size and complexity of the perception DNNs as…

Machine Learning · Computer Science 2025-04-30 Christopher Watson , Rajeev Alur , Divya Gopinath , Ravi Mangal , Corina S. Pasareanu

In this work, we present CineMaster, a novel framework for 3D-aware and controllable text-to-video generation. Our goal is to empower users with comparable controllability as professional film directors: precise placement of objects within…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Qinghe Wang , Yawen Luo , Xiaoyu Shi , Xu Jia , Huchuan Lu , Tianfan Xue , Xintao Wang , Pengfei Wan , Di Zhang , Kun Gai

Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development of diffusion models recently. Existing methods mainly rely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Zirui Pan , Xin Wang , Yipeng Zhang , Hong Chen , Kwan Man Cheng , Yaofei Wu , Wenwu Zhu

Recent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion. Previous…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Omri Avrahami , Thomas Hayes , Oran Gafni , Sonal Gupta , Yaniv Taigman , Devi Parikh , Dani Lischinski , Ohad Fried , Xi Yin

Understanding a visual scene goes beyond recognizing individual objects in isolation. Relationships between objects also constitute rich semantic information about the scene. In this work, we explicitly model the objects and their…

Computer Vision and Pattern Recognition · Computer Science 2017-04-13 Danfei Xu , Yuke Zhu , Christopher B. Choy , Li Fei-Fei

Dynamic scene understanding is the ability of a computer system to interpret and make sense of the visual information present in a video of a real-world scene. In this thesis, we present a series of frameworks for dynamic scene…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Salman Khan

Despite the recent progress of generative adversarial networks (GANs) at synthesizing photo-realistic images, producing complex urban scenes remains a challenging problem. Previous works break down scene generation into two consecutive…

Computer Vision and Pattern Recognition · Computer Science 2021-06-04 Guillaume Le Moing , Tuan-Hung Vu , Himalaya Jain , Patrick Pérez , Matthieu Cord

We present GALA3D, generative 3D GAussians with LAyout-guided control, for effective compositional text-to-3D generation. We first utilize large language models (LLMs) to generate the initial layout and introduce a layout-guided 3D Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Xiaoyu Zhou , Xingjian Ran , Yajiao Xiong , Jinlin He , Zhiwei Lin , Yongtao Wang , Deqing Sun , Ming-Hsuan Yang

Cinematic camera control relies on a tight feedback loop between director and cinematographer, where camera motion and framing are continuously reviewed and refined. Recent generative camera systems can produce diverse, text-conditioned…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Mengtian Li , Yuwei Lu , Feifei Li , Chenqi Gan , Zhifeng Xie , Xi Wang

Story visualization aims to generate a series of images that match the story described in texts, and it requires the generated images to satisfy high quality, alignment with the text description, and consistency in character identities.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Wen Wang , Canyu Zhao , Hao Chen , Zhekai Chen , Kecheng Zheng , Chunhua Shen

Synthesising safe controllers from visual data typically requires extensive supervised labelling of safety-critical data, which is often impractical in real-world settings. Recent advances in world models enable reliable prediction in…

Robotics · Computer Science 2025-07-21 Mehul Anand , Shishir Kolathaya

Ensuring the safety and robustness of autonomous driving systems necessitates a comprehensive evaluation in safety-critical scenarios. However, these safety-critical scenarios are rare and difficult to collect from real-world driving data,…

Artificial Intelligence · Computer Science 2025-08-19 Mingxing Peng , Yuting Xie , Xusen Guo , Ruoyu Yao , Hai Yang , Jun Ma

Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process. However, most existing pipelines only utilize the language modality for scene descriptions or reasoning and lack the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Sicheng Zuo , Yuxuan Li , Wenzhao Zheng , Zheng Zhu , Jie Zhou , Jiwen Lu

Scene understanding and reasoning has been a fundamental problem in 3D computer vision, requiring models to identify objects, their properties, and spatial or comparative relationships among the objects. Existing approaches enable this by…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Vivek Madhavaram , Vartika Sengar , Arkadipta De , Charu Sharma

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Bohan Li , Shuojue Yang , Baorui Peng , Xianda Guo , Erli Zhang , Youqi Tao , Junfeng Duan , Daguang Xu , Qi Dou , Xin Jin , Wenjun Zeng , Hao Zhao , Yueming Jin

This presentation introduces a self-supervised learning approach to the synthesis of new video clips from old ones, with several new key elements for improved spatial resolution and realism: It conditions the synthesis process on contextual…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Guillaume Le Moing , Jean Ponce , Cordelia Schmid

Many safety-critical applications, especially in autonomous driving, require reliable object detectors. They can be very effectively assisted by a method to search for and identify potential failures and systematic errors before these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Valentyn Boreiko , Matthias Hein , Jan Hendrik Metzen

Simulation-based testing has emerged as an essential tool for verifying and validating autonomous vehicles (AVs). However, contemporary methodologies, such as deterministic and imitation learning-based driver models, struggle to capture the…

Robotics · Computer Science 2025-11-04 Cheng Wang , Lingxin Kong , Massimiliano Tamborski , Stefano V. Albrecht

Machine learning based autonomous driving systems often face challenges with safety-critical scenarios that are rare in real-world data, hindering their large-scale deployment. While increasing real-world training data coverage could…

Machine Learning · Computer Science 2024-09-13 Yuan Yin , Pegah Khayatan , Éloi Zablocki , Alexandre Boulch , Matthieu Cord
‹ Prev 1 8 9 10 Next ›