中文
相关论文

相关论文: AnimateScene: Camera-controllable Animation in Any…

200 篇论文

Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Xinhang Gao , Junlin Guan , Shuhan Luo , Wenzhuo Li , Guanghuan Tan , Jiacheng Wang

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Shuolin Xu , Siming Zheng , Ziyi Wang , HC Yu , Jinwei Chen , Huaqi Zhang , Daquan Zhou , Tong-Yee Lee , Bo Li , Peng-Tao Jiang

Human animation aims to generate temporally coherent and visually consistent videos over long sequences, yet modeling long-range dependencies while preserving frame quality remains challenging. Inspired by the human ability to leverage past…

Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate progress in kinematic motion generation, they often fail to…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Kaiyang Ji , Ye Shi , Zichen Jin , Kangyi Chen , Lan Xu , Yuexin Ma , Jingyi Yu , Jingya Wang

In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yuhong Zhang , Guanlin Wu , Ling-Hao Chen , Zhuokai Zhao , Jing Lin , Xiaoke Jiang , Jiamin Wu , Zhuoheng Li , Hao Frank Yang , Haoqian Wang , Lei Zhang

We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Tanuj Sur , Shashank Tripathi , Nikos Athanasiou , Ha Linh Nguyen , Kai Xu , Michael J. Black , Angela Yao

We present a method that simultaneously addresses the tasks of dynamic scene novel-view synthesis and six degree-of-freedom (6-DOF) tracking of all dense scene elements. We follow an analysis-by-synthesis framework, inspired by recent work…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Jonathon Luiten , Georgios Kopanas , Bastian Leibe , Deva Ramanan

Generating high-quality, textured 3D scenes from a single image remains a fundamental challenge in vision and graphics. Recent image-to-3D generators recover reasonable geometry from single views, but their object-centric training limits…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Kaizhi Zheng , Yue Fan , Jing Gu , Zishuo Xu , Xuehai He , Xin Eric Wang

Dynamic environments that include unstructured moving objects pose a hard problem for Simultaneous Localization and Mapping (SLAM) performance. The motion of rigid objects can be typically tracked by exploiting their texture and geometric…

计算机视觉与模式识别 · 计算机科学 2021-08-03 Huayan Zhang , Tianwei Zhang , Tin Lun Lam , Sethu Vijayakumar

Comprehending natural language instructions is a charming property for 3D indoor scene synthesis systems. Existing methods directly model object joint distributions and express object relations implicitly within a scene, thereby hindering…

计算机视觉与模式识别 · 计算机科学 2024-02-08 Chenguo Lin , Yadong Mu

Interactive 3D scenes are increasingly vital for embodied intelligence, yet existing datasets remain limited due to the labor-intensive process of annotating part segmentation, kinematic types, and motion trajectories. We present REACT3D, a…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Zhao Huang , Boyang Sun , Alexandros Delitzas , Jiaqi Chen , Marc Pollefeys

This paper tackles the problem of physics-aware human motion synthesis in a dynamic scene. Unlike existing works which mainly tend to generate physically unrealistic motions due to limited contact modeling, typically restricted to hands, in…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Chaoyue Xing , Wei Mao , Miaomiao Liu

We present a novel diffusion-based approach for coherent 3D scene reconstruction from a single RGB image. Our method utilizes an image-conditioned 3D scene diffusion model to simultaneously denoise the 3D poses and geometries of all objects…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Manuel Dahnert , Angela Dai , Norman Müller , Matthias Nießner

Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, which construct 3D environments without active agents, and (2)…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Yitong Wang , Fangyun Wei , Hongyang Zhang , Bo Dai , Yan Lu

Synthesizing 3D human avatars interacting realistically with a scene is an important problem with applications in AR/VR, video games and robotics. Towards this goal, we address the task of generating a virtual human -- hands and full body…

机器人学 · 计算机科学 2023-03-30 Purva Tendulkar , Dídac Surís , Carl Vondrick

Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Jisu Nam , Yicong Hong , Chun-Hao Paul Huang , Feng Liu , JoungBin Lee , Jiyoung Kim , Siyoon Jin , Yunsung Lee , Jaeyoon Jung , Suhwan Choi , Seungryong Kim , Yang Zhou

Recent advancements in 4D scene reconstruction using neural radiance fields (NeRF) have demonstrated the ability to represent dynamic scenes from multi-view videos. However, they fail to reconstruct the dynamic scenes and struggle to fit…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Seoha Kim , Jeongmin Bae , Youngsik Yun , Hahyun Lee , Gun Bang , Youngjung Uh

Generating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Lingteng Qiu , Shenhao Zhu , Qi Zuo , Xiaodong Gu , Yuan Dong , Junfei Zhang , Chao Xu , Zhe Li , Weihao Yuan , Liefeng Bo , Guanying Chen , Zilong Dong

Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable. While several recent works investigate how to disentangle…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Michael Niemeyer , Andreas Geiger

Most current action recognition methods heavily rely on appearance information by taking an RGB sequence of entire image regions as input. While being effective in exploiting contextual information around humans, e.g., human appearance and…

计算机视觉与模式识别 · 计算机科学 2021-04-16 Gyeongsik Moon , Heeseung Kwon , Kyoung Mu Lee , Minsu Cho