中文
相关论文

相关论文: Genie: Generative Interactive Environments

200 篇论文

Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively…

计算机视觉与模式识别 · 计算机科学 2023-10-02 Anthony Hu , Lloyd Russell , Hudson Yeo , Zak Murez , George Fedoseev , Alex Kendall , Jamie Shotton , Gianluca Corrado

In this paper, we introduce ControlVAE, a novel model-based framework for learning generative motion control policies based on variational autoencoders (VAE). Our framework can learn a rich and flexible latent representation of skills and a…

图形学 · 计算机科学 2022-10-13 Heyuan Yao , Zhenhua Song , Baoquan Chen , Libin Liu

In this work we present an adversarial training algorithm that exploits correlations in video to learn --without supervision-- an image generator model with a disentangled latent space. The proposed methodology requires only a few…

计算机视觉与模式识别 · 计算机科学 2019-10-25 Facundo Tuesca , Lucas C. Uzal

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Zhenya Yang , Zhe Liu , Yuxiang Lu , Liping Hou , Chenxuan Miao , Siyi Peng , Bailan Feng , Xiang Bai , Hengshuang Zhao

In this paper, we propose a generative model, Temporal Generative Adversarial Nets (TGAN), which can learn a semantic representation of unlabeled videos, and is capable of generating videos. Unlike existing Generative Adversarial Nets…

机器学习 · 计算机科学 2017-08-21 Masaki Saito , Eiichi Matsumoto , Shunta Saito

The field of video generation has expanded significantly in recent years, with controllable and compositional video generation garnering considerable interest. Most methods rely on leveraging annotations such as text, objects' bounding…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Aram Davtyan , Sepehr Sameni , Björn Ommer , Paolo Favaro

The development of robust and generalizable robot learning models is critically contingent upon the availability of large-scale, diverse training data and reliable evaluation benchmarks. Collecting data in the physical world poses…

Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable environments for perception, reasoning, and action. Yet current research still lacks…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Jianjie Fang , Yingshan Lei , Qin Wan , Ziyou Wang , Yuchao Huang , Yongyan Xu , Baining Zhao , Weichen Zhang , Chen Gao , Xinlei Chen , Yong Li

In this paper, we introduce a novel dIffusion language modEl pre-training framework for text generation, which we call GENIE. GENIE is a large-scale pretrained diffusion language model that consists of an encoder and a diffusion-based…

计算与语言 · 计算机科学 2023-02-20 Zhenghao Lin , Yeyun Gong , Yelong Shen , Tong Wu , Zhihao Fan , Chen Lin , Nan Duan , Weizhu Chen

Generative adversarial networks are the state of the art approach towards learned synthetic image generation. Although early successes were mostly unsupervised, bit by bit, this trend has been superseded by approaches based on labelled…

计算机视觉与模式识别 · 计算机科学 2020-12-17 Ricard Durall , Kalun Ho , Franz-Josef Pfreundt , Janis Keuper

We address the following action-effect prediction task. Given an image depicting an initial state of the world and an action expressed in text, predict an image depicting the state of the world following the action. The prediction should…

计算机视觉与模式识别 · 计算机科学 2022-08-03 Fangjun Li , David C. Hogg , Anthony G. Cohn

Generative models have emerged as an essential building block for many image synthesis and editing tasks. Recent advances in this field have also enabled high-quality 3D or video content to be generated that exhibits either multi-view or…

计算机视觉与模式识别 · 计算机科学 2023-08-10 Sherwin Bahmani , Jeong Joon Park , Despoina Paschalidou , Hao Tang , Gordon Wetzstein , Leonidas Guibas , Luc Van Gool , Radu Timofte

In this work, we introduce a two-step framework for generative modeling of temporal data. Specifically, the generative adversarial networks (GANs) setting is employed to generate synthetic scenes of moving objects. To do so, we propose a…

计算机视觉与模式识别 · 计算机科学 2019-02-01 Isabela Albuquerque , João Monteiro , Tiago H. Falk

Generative AI is transforming image synthesis, enabling the creation of high-quality, diverse, and photorealistic visuals across industries like design, media, healthcare, and autonomous systems. Advances in techniques such as…

计算机视觉与模式识别 · 计算机科学 2025-01-31 Fouad Bousetouane

Generative world models hold significant potential for simulating interactions with visuomotor policies in varied environments. Frontier video models can enable generation of realistic observations and environment interactions in a scalable…

We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describing the targeted transformation, our generated images…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Tomáš Souček , Dima Damen , Michael Wray , Ivan Laptev , Josef Sivic

Generative artificial intelligence revolutionized society. Current models are trained by minimizing the distance between the produced data and the training set. Consequently, development is plateauing as they are intrinsically data-hungry…

机器学习 · 计算机科学 2025-06-09 Mattia Miotto , Lorenzo Monacelli

Humans naturally build mental models of object interactions and dynamics, allowing them to imagine how their surroundings will change if they take a certain action. While generative models today have shown impressive results on…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Sruthi Sudhakar , Ruoshi Liu , Basile Van Hoorick , Carl Vondrick , Richard Zemel

While recent advances in image editing have enabled impressive visual synthesis capabilities, current methods remain constrained by explicit textual instructions and limited editing operations, lacking deep comprehension of implicit user…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Dong Zhang , Lingfeng He , Rui Yan , Fei Shen , Jinhui Tang

Producing realistic character animations is one of the essential tasks in human-AI interactions. Considered as a sequence of poses of a humanoid, the task can be considered as a sequence generation problem with spatiotemporal smoothness and…

计算机视觉与模式识别 · 计算机科学 2020-05-26 Maryam Sadat Mirzaei , Kourosh Meshgi , Etienne Frigo , Toyoaki Nishida