English
Related papers

Related papers: MTVCraft: Tokenizing 4D Motion for Arbitrary Chara…

200 papers

Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware and markers limits scalability and real-world deployment. Advancing reliable markerless…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Yeeun Park , Miqdad Naduthodi , Suryansh Kumar

Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Hila Chefer , Shiran Zada , Roni Paiss , Ariel Ephrat , Omer Tov , Michael Rubinstein , Lior Wolf , Tali Dekel , Tomer Michaeli , Inbar Mosseri

Recent portrait animation methods have made significant strides in generating realistic lip synchronization. However, they often lack explicit control over head movements and facial expressions, and cannot produce videos from multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yukang Lin , Hokit Fung , Jianjin Xu , Zeping Ren , Adela S. M. Lau , Guosheng Yin , Xiu Li

Video 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal information from sequential 2D poses, which cannot model the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Zhongwei Qiu , Qiansheng Yang , Jian Wang , Dongmei Fu

Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide accurate 3D geometry, their performance drops markedly when…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yu Hu , Chong Cheng , Sicheng Yu , Xiaoyang Guo , Hao Wang

Creating compelling 3D character animations typically requires either expert use of professional software or expensive motion capture systems operated by skilled actors. We present DancingBox, a lightweight, vision-based system that makes…

Graphics · Computer Science 2026-03-19 Haocheng Yuan , Adrien Bousseau , Hao Pan , Lei Zhong , Changjian Li

Real-time animation of virtual characters has traditionally been accomplished by playing short sequences of animations structured in the form of a graph. These methods are time-consuming to set up and scale poorly with the number of motions…

Graphics · Computer Science 2023-10-10 Jose Luis Ponton

Current advances in human head modeling allow the generation of plausible-looking 3D head models via neural representations, such as NeRFs and SDFs. Nevertheless, constructing complete high-fidelity head models with explicitly controlled…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Artem Sevastopolsky , Philip-William Grassal , Simon Giebenhain , ShahRukh Athar , Luisa Verdoliva , Matthias Niessner

Pose-guided human image animation aims to synthesize realistic videos of a reference character driven by a sequence of poses. While diffusion-based methods have achieved remarkable success, most existing approaches are limited to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yingcheng Hu , Haowen Gong , Chuanguang Yang , Zhulin An , Yongjun Xu , Songhua Liu

Recent advances in deep learning have significantly pushed the state-of-the-art in photorealistic video animation given a single image. In this paper, we extrapolate those advances to the 3D domain, by studying 3D image-to-video translation…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Rolandos Alexandros Potamias , Jiali Zheng , Stylianos Ploumpis , Giorgos Bouritsas , Evangelos Ververas , Stefanos Zafeiriou

Metaverse platforms are rapidly evolving to provide immersive spaces for user interaction and content creation. However, the generation of dynamic and interactive 3D objects remains challenging due to the need for advanced 3D modeling and…

Human-Computer Interaction · Computer Science 2025-05-01 Ryutaro Kurai , Takefumi Hiraki , Yuichi Hiroi , Yutaro Hirao , Monica Perusquía-Hernández , Hideaki Uchiyama , Kiyoshi Kiyokawa

Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Ying Zang , Xuanyi Liu , Yidong Han , Deyi Ji , Chaotao Ding , Yuanqi Hu , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. In content creation workflows, precise and simultaneous control over camera motion, object motion, and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Sixiao Zheng , Zimian Peng , Yanpeng Zhou , Yi Zhu , Hang Xu , Xiangru Huang , Yanwei Fu

Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ashkan Taghipour , Morteza Ghahremani , Mohammed Bennamoun , Farid Boussaid , Aref Miri Rekavandi , Zinuo Li , Qiuhong Ke , Hamid Laga

Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is still a black box, where all attributes (e.g., appearance,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Jiwen Yu , Xiaodong Cun , Chenyang Qi , Yong Zhang , Xintao Wang , Ying Shan , Jian Zhang

Recent advances in video diffusion models have significantly improved character animation techniques. However, current approaches rely on basic structural conditions such as DWPose or SMPL-X to animate character images, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Muyao Niu , Mingdeng Cao , Yifan Zhan , Qingtian Zhu , Mingze Ma , Jiancheng Zhao , Yanhong Zeng , Zhihang Zhong , Xiao Sun , Yinqiang Zheng

We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In MoMask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Chuan Guo , Yuxuan Mu , Muhammad Gohar Javed , Sen Wang , Li Cheng

We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts. VoiceCraft employs a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Puyuan Peng , Po-Yao Huang , Shang-Wen Li , Abdelrahman Mohamed , David Harwath

Autonomous motion capture (mocap) systems for outdoor scenarios involving flying or mobile cameras rely on i) a robotic front-end to track and follow a human subject in real-time while he/she performs physical activities, and ii) an…

Recent techniques for text-to-4D generation synthesize dynamic 3D scenes using supervision from pre-trained text-to-video models. However, existing representations for motion, such as deformation models or time-dependent neural…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sherwin Bahmani , Xian Liu , Wang Yifan , Ivan Skorokhodov , Victor Rong , Ziwei Liu , Xihui Liu , Jeong Joon Park , Sergey Tulyakov , Gordon Wetzstein , Andrea Tagliasacchi , David B. Lindell