English
Related papers

Related papers: DIMO: Diverse 3D Motion Generation for Arbitrary O…

200 papers

Generating 3D human motion based on textual descriptions has been a research focus in recent years. It requires the generated motion to be diverse, natural, and conform to the textual description. Due to the complex spatio-temporal nature…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Chongyang Zhong , Lei Hu , Zihao Zhang , Shihong Xia

Generating diverse and natural human motion is one of the long-standing goals for creating intelligent characters in the animated world. In this paper, we propose a self-supervised method for generating long-range, diverse and plausible…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Jingwei Xu , Huazhe Xu , Bingbing Ni , Xiaokang Yang , Xiaolong Wang , Trevor Darrell

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets and paired text descriptions. However, how to effectively…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Xiaoyan Cong , Zekun Li , Zhiyang Dou , Hongyu Li , Omid Taheri , Chuan Guo , Abhay Mittal , Sizhe An , Taku Komura , Wojciech Matusik , Michael J. Black , Srinath Sridhar

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent…

Graphics · Computer Science 2026-05-20 Yi-Yang Zhang , Tengjiao Sun , Pengcheng Fang , Deng-Bao Wang , Xiaohao Cai , Min-Ling Zhang , Hansung Kim

Simple as it seems, moving an object to another location within an image is, in fact, a challenging image-editing task that requires re-harmonizing the lighting, adjusting the pose based on perspective, accurately filling occluded regions,…

Graphics · Computer Science 2025-03-12 Xin Yu , Tianyu Wang , Soo Ye Kim , Paul Guerrero , Xi Chen , Qing Liu , Zhe Lin , Xiaojuan Qi

Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowledge of the motion…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Ekkasit Pinyoanuntapong , Muhammad Usama Saleem , Pu Wang , Minwoo Lee , Srijan Das , Chen Chen

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a single image, seamlessly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Minjun Kang , Inkyu Shin , Taeyeop Lee , In So Kweon , Kuk-Jin Yoon

We propose Neural 3D Articulation Prior (NAP), the first 3D deep generative model to synthesize 3D articulated object models. Despite the extensive research on generating 3D objects, compositions, or scenes, there remains a lack of focus on…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Jiahui Lei , Congyue Deng , Bokui Shen , Leonidas Guibas , Kostas Daniilidis

Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have shown promising…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Juil Koo , Paul Guerrero , Chun-Hao Paul Huang , Duygu Ceylan , Minhyuk Sung

Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional objects from single images--particularly under occlusions--remains challenging. Existing methods often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Hui Shan , Keyang Luo , Ming Li , Sizhe Zheng , Yanwei Fu , Zhen Chen , Xiangru Huang

In this work, we introduce a novel approach for creating controllable dynamics in 3D-generated Gaussians using casually captured reference videos. Our method transfers the motion of objects from reference videos to a variety of generated 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Zhoujie Fu , Jiacheng Wei , Wenhao Shen , Chaoyue Song , Xiaofeng Yang , Fayao Liu , Xulei Yang , Guosheng Lin

Predicting future motion is crucial in video understanding and controllable video generation. Dense point trajectories are a compact, expressive motion representation, but modeling their future evolution from observed video remains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Zewei Zhang , Jia Jun Cheng Xian , Kaiwen Liu , Ming Liang , Hang Chu , Jun Chen , Renjie Liao

Attributes such as style, fine-grained text, and trajectory are specific conditions for describing motion. However, existing methods often lack precise user control over motion attributes and suffer from limited generalizability to unseen…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Mingjie Wei , Xuemei Xie , Guangming Shi

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas

We propose ArtiLatent, a generative framework that synthesizes human-made 3D objects with fine-grained geometry, accurate articulation, and realistic appearance. Our approach jointly models part geometry and articulation dynamics by…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Honghua Chen , Yushi Lan , Yongwei Chen , Xingang Pan

Deep generative models are often used for human motion prediction as they are able to model multi-modal data distributions and characterize diverse human behavior. While much care has been taken into designing and learning deep generative…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Ye Yuan , Kris Kitani

Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in…

Robotics · Computer Science 2026-03-06 Yichen Cai , Jianfeng Gao , Christoph Pohl , Tamim Asfour

Driven by powerful image diffusion models, recent research has achieved the automatic creation of 3D objects from textual or visual guidance. By performing score distillation sampling (SDS) iteratively across different views, these methods…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Zeyu Li , Ruitong Gan , Chuanchen Luo , Yuxi Wang , Jiaheng Liu , Ziwei Zhu Man Zhang , Qing Li , Xucheng Yin , Zhaoxiang Zhang , Junran Peng

Object-level mapping builds a 3D map of objects in a scene with detailed shapes and poses from multi-view sensor observations. Conventional methods struggle to build complete shapes and estimate accurate poses due to partial occlusions and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Ziwei Liao , Binbin Xu , Steven L. Waslander

Robotic manipulation systems operating in complex environments rely on perception systems that provide information about the geometry (pose and 3D shape) of the objects in the scene along with other semantic information such as object…

Robotics · Computer Science 2023-05-17 Shubham Agrawal , Nikhil Chavan-Dafle , Isaac Kasahara , Selim Engin , Jinwook Huh , Volkan Isler