English
Related papers

Related papers: Mitty: Diffusion-based Human-to-Robot Video Genera…

200 papers

Motion in-betweening, a fundamental task in character animation, consists of generating motion sequences that plausibly interpolate user-provided keyframe constraints. It has long been recognized as a labor-intensive and challenging…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Setareh Cohan , Guy Tevet , Daniele Reda , Xue Bin Peng , Michiel van de Panne

Despite significant advancements in customizing text-to-image and video generation models, generating images and videos that effectively integrate multiple personalized concepts remains a challenging task. To address this, we present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Gihyun Kwon , Jong Chul Ye

Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yujiang Pu , Zhanbo Huang , Vishnu Boddeti , Yu Kong

This paper investigates a solution for enabling in-context capabilities of video diffusion transformers, with minimal tuning required for activation. Specifically, we propose a simple pipeline to leverage in-context generation:…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Zhengcong Fei , Di Qiu , Debang Li , Changqian Yu , Mingyuan Fan

Diffusion generative modeling has become a promising approach for learning robotic manipulation tasks from stochastic human demonstrations. In this paper, we present Diffusion-EDFs, a novel SE(3)-equivariant diffusion-based approach for…

While diffusion models have shown great potential in portrait generation, generating expressive, coherent, and controllable cinematic portrait videos remains a significant challenge. Existing intermediate signals for portrait generation,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Junyi Wang , Yudong Guo , Boyang Guo , Shengming Yang , Juyong Zhang

Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a challenging, resource-intensive process requiring deliberate…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Alejandro Pardo , Fabio Pizzati , Tong Zhang , Alexander Pondaven , Philip Torr , Juan Camilo Perez , Bernard Ghanem

Learning rewards from expert videos offers an affordable and effective solution to specify the intended behaviors for reinforcement learning (RL) tasks. In this work, we propose Diffusion Reward, a novel framework that learns rewards from…

Machine Learning · Computer Science 2024-08-12 Tao Huang , Guangqi Jiang , Yanjie Ze , Huazhe Xu

Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous demonstrations, and end-to-end pixel-to-action pipelines may degenerate under background and viewpoint…

Machine Learning · Computer Science 2025-12-23 Yao Feng , Hengkai Tan , Xinyi Mao , Chendong Xiang , Guodong Liu , Shuhe Huang , Hang Su , Jun Zhu

We present Stable Video Diffusion - a latent video diffusion model for high-resolution, state-of-the-art text-to-video and image-to-video generation. Recently, latent diffusion models trained for 2D image synthesis have been turned into…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Andreas Blattmann , Tim Dockhorn , Sumith Kulal , Daniel Mendelevitch , Maciej Kilian , Dominik Lorenz , Yam Levi , Zion English , Vikram Voleti , Adam Letts , Varun Jampani , Robin Rombach

With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Jianwen Jiang , Chao Liang , Jiaqi Yang , Gaojie Lin , Tianyun Zhong , Yanbo Zheng

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

Event cameras, mimicking the human retina, capture brightness changes with unparalleled temporal resolution and dynamic range. Integrating events into intensities poses a highly ill-posed challenge, marred by initial condition ambiguities.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Jinxiu Liang , Bohan Yu , Yixin Yang , Yiming Han , Boxin Shi

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

We present a method to create diffusion-based video models from pretrained Text-to-Image (T2I) models. Recently, AnimateDiff proposed freezing the T2I model while only training temporal layers. We advance this method by proposing a unique…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Mingi Kwon , Seoung Wug Oh , Yang Zhou , Difan Liu , Joon-Young Lee , Haoran Cai , Baqiao Liu , Feng Liu , Youngjung Uh

Computer-assisted interventions can improve intra-operative guidance, particularly through deep learning methods that harness the spatiotemporal information in surgical videos. However, the severe data imbalance often found in surgical…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Danush Kumar Venkatesh , Isabel Funke , Micha Pfeiffer , Fiona Kolbinger , Hanna Maria Schmeiser , Juergen Weitz , Marius Distler , Stefanie Speidel

Visual generation and understanding are two deeply interconnected aspects of human intelligence, yet they have been traditionally treated as separate tasks in machine learning. In this paper, we propose Jodi, a diffusion framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yifeng Xu , Zhenliang He , Meina Kan , Shiguang Shan , Xilin Chen

Generating multi-view human images from a single view is a complex and significant challenge. Although recent advancements in multi-view object generation have shown impressive results with diffusion models, novel view synthesis for humans…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Yu Feng , Shunsi Zhang , Jian Shu , Hanfeng Zhao , Guoliang Pang , Chi Zhang , Hao Wang

Constructing robots to accomplish long-horizon tasks is a long-standing challenge within artificial intelligence. Approaches using generative methods, particularly Diffusion Models, have gained attention due to their ability to model…

Robotics · Computer Science 2026-04-30 Sigmund Hennum Høeg , Aksel Vaaler , Chaoqi Liu , Olav Egeland , Yilun Du

Robots hold great promise for performing repetitive or hazardous tasks, but achieving human-like dexterity, especially in contact-rich and dynamic environments, remains challenging. Rigid robots, which rely on position or velocity control,…

Robotics · Computer Science 2024-10-28 Malek Aburub , Cristian C. Beltran-Hernandez , Tatsuya Kamijo , Masashi Hamaya