中文
相关论文

相关论文: Dream2Flow: Bridging Video Generation and Open-Wor…

200 篇论文

Recent advances in video generation enable a new paradigm for 3D scene creation: generating camera-controlled videos that simulate scene walkthroughs, then lifting them to 3D via feed-forward reconstruction techniques. This generative…

Image-to-3D generation faces inherent semantic ambiguity under occlusion, where partial observation alone is often insufficient to determine object category. In this work, we formalize text-driven amodal 3D generation, where text prompts…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Jiayin Zhu , Guoji Fu , Xiaolu Liu , Qiyuan He , Yicong Li , Angela Yao

Video motion transfer aims to generate a target video that inherits motion patterns from a source video while rendering new scenes. Existing training-free approaches focus on constructing motion guidance based on the intermediate outputs of…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Zhen Wang , Youcan Xu , Jun Xiao , Long Chen

Humans can easily segment moving objects without knowing what they are. That objectness could emerge from continuous visual observations motivates us to model grouping and movement concurrently from unlabeled videos. Our premise is that a…

计算机视觉与模式识别 · 计算机科学 2021-11-12 Runtao Liu , Zhirong Wu , Stella X. Yu , Stephen Lin

Generating articulated objects, such as laptops and microwaves, is a crucial yet challenging task with extensive applications in Embodied AI and AR/VR. Current image-to-3D methods primarily focus on surface geometry and texture, neglecting…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Ruijie Lu , Yu Liu , Jiaxiang Tang , Junfeng Ni , Yuxiang Wang , Diwen Wan , Gang Zeng , Yixin Chen , Siyuan Huang

Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual and motion priors. However, they still face key challenges: a…

机器人学 · 计算机科学 2026-02-04 Yixiang Chen , Peiyan Li , Jiabing Yang , Keji He , Xiangnan Wu , Yuan Xu , Kai Wang , Jing Liu , Nianfeng Liu , Yan Huang , Liang Wang

Optical flow is the motion of a pixel between at least two consecutive video frames and can be estimated through an end-to-end trainable convolutional neural network. To this end, large training datasets are required to improve the accuracy…

计算机视觉与模式识别 · 计算机科学 2021-04-19 Roman Seidel , André Apitzsch , Gangolf Hirtz

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that…

机器人学 · 计算机科学 2025-05-16 Jun Guo , Xiaojian Ma , Yikai Wang , Min Yang , Huaping Liu , Qing Li

Recent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in videos generated by any…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Changgu Chen , Junwei Shu , Gaoqi He , Changbo Wang , Yang Li

We tackle the problem of text-driven 3D generation from a geometry alignment perspective. Given a set of text prompts, we aim to generate a collection of objects with semantically corresponding parts aligned across them. Recent methods…

Motion-controllable image animation is a fundamental task with a wide range of potential applications. Recent works have made progress in controlling camera or object motion via various motion representations, while they still struggle to…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yingjie Chen , Yifang Men , Yuan Yao , Miaomiao Cui , Liefeng Bo

We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Yunuo Chen , Junli Cao , Vidit Goel , Sergei Korolev , Chenfanfu Jiang , Jian Ren , Sergey Tulyakov , Anil Kag

Real-time in-between motion generation is universally required in games and highly desirable in existing animation pipelines. Its core challenge lies in the need to satisfy three critical conditions simultaneously: quality, controllability…

图形学 · 计算机科学 2022-05-06 Xiangjun Tang , He Wang , Bo Hu , Xu Gong , Ruifan Yi , Qilong Kou , Xiaogang Jin

Understanding and manipulating articulated objects, such as doors and drawers, is crucial for robots operating in human environments. We wish to develop a system that can learn to articulate novel objects with no prior interaction, after…

机器人学 · 计算机科学 2024-05-03 Harry Zhang , Ben Eisner , David Held

Teaching robots dexterous manipulation skills often requires collecting hundreds of demonstrations using wearables or teleoperation, a process that is challenging to scale. Videos of human-object interactions are easier to collect and…

机器人学 · 计算机科学 2025-08-19 Tyler Ga Wei Lum , Olivia Y. Lee , C. Karen Liu , Jeannette Bohg

Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise patterns inherent to real sensors. In this work, inspired by…

机器人学 · 计算机科学 2025-12-09 Xiujian Liang , Jiacheng Liu , Mingyang Sun , Qichen He , Cewu Lu , Jianhua Sun

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal…

Advances in video generation have significantly improved the realism and quality of created scenes. This has fueled interest in developing intuitive tools that let users leverage video generation as world simulators. Text-to-video (T2V)…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Zuhao Liu , Aleksandar Yanev , Ahmad Mahmood , Ivan Nikolov , Saman Motamed , Wei-Shi Zheng , Xi Wang , Lei Sun , Luc Van Gool , Danda Pani Paudel

Unsupervised video object segmentation (VOS), also known as video salient object detection, aims to detect the most prominent object in a video at the pixel level. Recently, two-stream approaches that leverage both RGB images and optical…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Suhwan Cho , Minhyeok Lee , Jungho Lee , Donghyeong Kim , Seunghoon Lee , Sungmin Woo , Sangyoun Lee

Creating 4D fields of Gaussian Splatting from images or videos is a challenging task due to its under-constrained nature. While the optimization can draw photometric reference from the input videos or be regulated by generative models,…

计算机视觉与模式识别 · 计算机科学 2024-05-15 Quankai Gao , Qiangeng Xu , Zhe Cao , Ben Mildenhall , Wenchao Ma , Le Chen , Danhang Tang , Ulrich Neumann