中文
相关论文

相关论文: VerseCrafter: Dynamic Realistic Video World Model …

200 篇论文

Novel view synthesis of dynamic scenes is becoming important in various applications, including augmented and virtual reality. We propose a novel 4D Gaussian Splatting (4DGS) algorithm for dynamic scenes from casually recorded monocular…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Mijeong Kim , Jongwoo Lim , Bohyung Han

We present InfiniCube, a scalable method for generating unbounded dynamic 3D driving scenes with high fidelity and controllability. Previous methods for scene generation either suffer from limited scales or lack geometric and appearance…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Yifan Lu , Xuanchi Ren , Jiawei Yang , Tianchang Shen , Zhangjie Wu , Jun Gao , Yue Wang , Siheng Chen , Mike Chen , Sanja Fidler , Jiahui Huang

Recent progress in text-to-video generation has achieved remarkable realism, yet fine-grained control over camera motion and orientation remains elusive, especially with extreme trajectories (e.g., a 180-degree turnaround, or looking…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Frédéric Fortier-Chouinard , Yannick Hold-Geoffroy , Valentin Deschaintre , Matheus Gadelha , Jean-François Lalonde

Representing and rendering dynamic scenes has been an important but challenging task. Especially, to accurately model complex motions, high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Guanjun Wu , Taoran Yi , Jiemin Fang , Lingxi Xie , Xiaopeng Zhang , Wei Wei , Wenyu Liu , Qi Tian , Xinggang Wang

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Gwanghyun Kim , Xueting Li , Ye Yuan , Koki Nagano , Tianye Li , Jan Kautz , Se Young Chun , Umar Iqbal

Recent Text-to-Video (T2V) models have demonstrated powerful capability in visual simulation of real-world geometry and physical laws, indicating its potential as implicit world models. Inspired by this, we explore the feasibility of…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Yu Li , Menghan Xia , Gongye Liu , Jianhong Bai , Xintao Wang , Conglang Zhang , Yuxuan Lin , Ruihang Chu , Pengfei Wan , Yujiu Yang

Virtual 3D meetings offer the potential to enhance copresence, increase engagement and thus improve effectiveness of remote meetings compared to standard 2D video calls. However, representing people in 3D meetings remains a challenge;…

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic environments and enable open-vocabulary querying in complex…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Xianfeng Wu , Yajing Bai , Minghan Li , Xianzu Wu , Xueqi Zhao , Zhongyuan Lai , Wenyu Liu , Xinggang Wang

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

Realistic reconstruction of dynamic 4D scenes from monocular videos is essential for understanding the physical world. Despite recent progress in neural rendering, existing methods often struggle to recover accurate 3D geometry and…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Haoran Zhou , Gim Hee Lee

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

Digitizing humans and synthesizing photorealistic avatars with explicit 3D pose and camera controls are central to VR, telepresence, and entertainment. Existing skinning-based workflows require laborious manual rigging or template-based…

Wrist-view observations are crucial for VLA models as they capture fine-grained hand-object interactions that directly enhance manipulation performance. Yet large-scale datasets rarely include such recordings, resulting in a substantial gap…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Zezhong Qian , Xiaowei Chi , Yuming Li , Shizun Wang , Zhiyuan Qin , Xiaozhu Ju , Sirui Han , Shanghang Zhang

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform general dynamic scenes. In…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yuheng Jiang , Zhehao Shen , Chengcheng Guo , Yu Hong , Zhuo Su , Yingliang Zhang , Marc Habermann , Lan Xu

Image view synthesis has seen great success in reconstructing photorealistic visuals, thanks to deep learning and various novel representations. The next key step in immersive virtual experiences is view synthesis of dynamic scenes.…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Kai-En Lin , Guowei Yang , Lei Xiao , Feng Liu , Ravi Ramamoorthi

With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Weixing Xie , Xiao Dong , Yong Yang , Qiqin Lin , Jingze Chen , Junfeng Yao , Xiaohu Guo

A general concept of 3D volumetric visualization systems is described based on 3D discrete voxel scenes (worlds) representation. Definitions of 3D discrete voxel scene (world) basic elements and main steps of the image synthesis algorithm…

图形学 · 计算机科学 2017-02-07 Anas M. Al-Oraiqat , E. A. Bashkov , S. A. Zori , Aladdein M. Amro

Generating realistic and controllable weather effects in videos is valuable for many applications. Physics-based weather simulation requires precise reconstructions that are hard to scale to in-the-wild videos, while current video editing…

图形学 · 计算机科学 2025-07-22 Chih-Hao Lin , Zian Wang , Ruofan Liang , Yuxuan Zhang , Sanja Fidler , Shenlong Wang , Zan Gojcic

Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome this limitation with a highly scalable self-supervised framework capable of leveraging…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Avinash Paliwal , Adithya Iyer , Shivin Yadav , Muhammad Ali Afridi , Midhun Harikumar

Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appearance across…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Yuanze Lin , Yi-Wen Chen , Yi-Hsuan Tsai , Ronald Clark , Ming-Hsuan Yang
‹ 上一页 1 8 9 10 下一页 ›