中文
相关论文

相关论文: ActCam: Zero-Shot Joint Camera and 3D Motion Contr…

200 篇论文

Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions,…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Qiucheng Wu , Handong Zhao , Zhixin Shu , Jing Shi , Yang Zhang , Shiyu Chang

Recent advancements in generative models have revolutionized the field of artificial intelligence, enabling the creation of highly-realistic and detailed images. In this study, we propose a novel Mask Conditional Text-to-Image Generative…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Rami Skaik , Leonardo Rossi , Tomaso Fontanini , Andrea Prati

We propose ReCamDriving, a purely vision-based, camera-controlled novel-trajectory video generation framework. While repair-based methods fail to restore complex artifacts and LiDAR-based approaches rely on sparse and incomplete cues,…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Yaokun Li , Shuaixian Wang , Mantang Guo , Jiehui Huang , Taojun Ding , Mu Hu , Kaixuan Wang , Shaojie Shen , Guang Tan

Zero-shot Text-to-Video synthesis generates videos based on prompts without any videos. Without motion information from videos, motion priors implied in prompts are vital guidance. For example, the prompt "airplane landing on the runway"…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Sitong Su , Litao Guo , Lianli Gao , Hengtao Shen , Jingkuan Song

Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Ruineng Li , Daitao Xing , Huiming Sun , Yuanzhou Ha , Jinglin Shen , Chiuman Ho

Interactive 3D scenes are increasingly vital for embodied intelligence, yet existing datasets remain limited due to the labor-intensive process of annotating part segmentation, kinematic types, and motion trajectories. We present REACT3D, a…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Zhao Huang , Boyang Sun , Alexandros Delitzas , Jiaqi Chen , Marc Pollefeys

Image-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accurately modeling…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Chenhui Zhu , Yilu Wu , Shuai Wang , Gangshan Wu , Limin Wang

Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error,…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Koichi Namekata , Sherwin Bahmani , Ziyi Wu , Yash Kant , Igor Gilitschenski , David B. Lindell

Image caption generation is one of the most challenging problems at the intersection of vision and language domains. In this work, we propose a realistic captioning task where the input scenes may incorporate visual objects with no…

计算机视觉与模式识别 · 计算机科学 2022-07-04 Berkan Demirel , Ramazan Gokberk Cinbis

This paper proposes a novel application system for the generation of three-dimensional (3D) character animation driven by markerless human body motion capturing. The entire pipeline of the system consists of five stages: 1) the capturing of…

计算机视觉与模式识别 · 计算机科学 2022-12-13 Jinbao Wang , Ke Lu , Jian Xue

Aerial cinematography is revolutionizing industries that require live and dynamic camera viewpoints such as entertainment, sports, and security. However, safely piloting a drone while filming a moving target in the presence of obstacles is…

Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Beiyuan Zhang , Yue Ma , Chunlei Fu , Xinyang Song , Zhenan Sun , Ziqiang Li

We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion…

计算机视觉与模式识别 · 计算机科学 2022-11-24 Liang Xu , Ziyang Song , Dongliang Wang , Jing Su , Zhicheng Fang , Chenjing Ding , Weihao Gan , Yichao Yan , Xin Jin , Xiaokang Yang , Wenjun Zeng , Wei Wu

Recent text-to-video generation approaches rely on computationally heavy training and require large-scale video datasets. In this paper, we introduce a new task of zero-shot text-to-video generation and propose a low-cost approach (without…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Levon Khachatryan , Andranik Movsisyan , Vahram Tadevosyan , Roberto Henschel , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

Scalable sensor simulation is an important yet challenging open problem for safety-critical domains such as self-driving. Current works in image simulation either fail to be photorealistic or do not model the 3D environment and the dynamic…

计算机视觉与模式识别 · 计算机科学 2021-05-18 Yun Chen , Frieda Rong , Shivam Duggal , Shenlong Wang , Xinchen Yan , Sivabalan Manivasagam , Shangjie Xue , Ersin Yumer , Raquel Urtasun

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal…

Controllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cinematic technique known as Frame In and Frame Out.…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Boyang Wang , Xuweiyi Chen , Matheus Gadelha , Zezhou Cheng

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Video generation models have advanced rapidly and are beginning to show a strong understanding of physical dynamics. In this paper, we investigate how far an advanced video generation model such as Veo-3 can support generalizable robotic…

机器人学 · 计算机科学 2026-04-07 Zhongru Zhang , Chenghan Yang , Qingzhou Lu , Yanjiang Guo , Jianke Zhang , Yucheng Hu , Jianyu Chen

Self-captured full-body videos are popular, but most deployments require mounted cameras, carefully-framed shots, and repeated practice. We propose a more convenient solution that enables full-body video capture using handheld mobile…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Bowei Chen , Brian Curless , Ira Kemelmacher-Shlizerman , Steven M. Seitz