English
Related papers

Related papers: HY-World 2.0: A Multi-Modal World Model for Recons…

200 papers

Digitizing the physical world into accurate simulation-ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D reconstruction and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Hongchi Xia , Chih-Hao Lin , Hao-Yu Hsu , Quentin Leboutet , Katelyn Gao , Michael Paulitsch , Benjamin Ummenhofer , Shenlong Wang

We present GauStudio, a novel modular framework for modeling 3D Gaussian Splatting (3DGS) to provide standardized, plug-and-play components for users to easily customize and implement a 3DGS pipeline. Supported by our framework, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Chongjie Ye , Yinyu Nie , Jiahao Chang , Yuantao Chen , Yihao Zhi , Xiaoguang Han

Photorealistic 3D full-body human reconstruction from a single image is a critical yet challenging task for applications in films and video games due to inherent ambiguities and severe self-occlusions. While recent approaches leverage SMPL…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Wenyue Chen , Peng Li , Wangguandong Zheng , Chengfeng Zhao , Mengfei Li , Yaolong Zhu , Zhiyang Dou , Ronggang Wang , Yuan Liu

Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Zhiheng Liu , Xueqing Deng , Shoufa Chen , Angtian Wang , Qiushan Guo , Mingfei Han , Zeyue Xue , Mengzhao Chen , Ping Luo , Linjie Yang

Building multimodal language models is fundamentally challenging: it requires aligning vision and language modalities, curating high-quality instruction data, and avoiding the degradation of existing text-only capabilities once vision is…

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module publicly available. The achievement of HunyuanImage 3.0 relies…

Generating realistic hand-object interactions (HOI) videos is a significant challenge due to the difficulty of modeling physical constraints (e.g., contact and occlusion between hands and manipulated objects). Current methods utilize HOI…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Haodong Yan , Hang Yu , Zhide Zhong , Weilin Yuan , Xin Gong , Zehang Luo , Chengxi Heyu , Junfeng Li , Wenxuan Song , Shunbo Zhou , Haoang Li

We propose a new view synthesis method via synthesizing a 3D neural field from both single or few-view input images. To address the ill-posed nature of the image-to-3D generation problem, we devise a two-stage method that involves a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Tung Do , Thuan Hoang Nguyen , Anh Tuan Tran , Rang Nguyen , Binh-Son Hua

In this paper, we present a method to reconstruct the world and multiple dynamic humans in 3D from a monocular video input. As a key idea, we represent both the world and multiple humans via the recently emerging 3D Gaussian Splatting…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Inhee Lee , Byungjun Kim , Hanbyul Joo

Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReels-V3, a conditional video generation model, built upon a…

World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Zeqi Xiao , Yushi Lan , Yifan Zhou , Wenqi Ouyang , Shuai Yang , Yanhong Zeng , Xingang Pan

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Weijie Wang , Xiaoxuan He , Youping Gu , Yifan Yang , Zeyu Zhang , Yefei He , Yanbo Ding , Xirui Hu , Donny Y. Chen , Zhiyuan He , Yuqing Yang , Bohan Zhuang

Reconstructing dynamic scenes with multiple interacting humans and objects from sparse-view inputs is a critical yet challenging task, essential for creating high-fidelity digital twins for robotics and VR/AR. This problem, which we term…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Weiquan Wang , Jun Xiao , Feifei Shao , Yi Yang , Yueting Zhuang , Long Chen

We describe Generative Blocks World to interact with the scene of a generated image by manipulating simple geometric abstractions. Our method represents scenes as assemblies of convex 3D primitives, and the same scene can be represented by…

Graphics · Computer Science 2026-03-23 Vaibhav Vavilala , Seemandhar Jain , Rahul Vasanth , D. A. Forsyth , Anand Bhattad

Perpetual view generation aims to synthesize a long-term video corresponding to an arbitrary camera trajectory solely from a single input image. Recent methods commonly utilize a pre-trained text-to-image diffusion model to synthesize new…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Bo Pan , Yang Chen , Yingwei Pan , Ting Yao , Wei Chen , Tao Mei

Generating high-quality 3D content from text, single images, or sparse view images remains a challenging task with broad applications. Existing methods typically employ multi-view diffusion models to synthesize multi-view images, followed…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junlin Han , Jianyuan Wang , Andrea Vedaldi , Philip Torr , Filippos Kokkinos

We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) and the demands of embodied agents, our models are developed…

3D semantic scene graphs (3DSSG) provide compact structured representations of environments by explicitly modeling objects, attributes, and relationships. While 3DSSGs have shown promise in robotics and embodied AI, many existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Marian Renz , Felix Igelbrink , Martin Atzmueller

Recent progress in 3D object generation has been fueled by the strong priors offered by diffusion models. However, existing models are tailored to specific tasks, accommodating only one modality at a time and necessitating retraining to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Yijun Fan , Yiwei Ma , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Jiahao Lu , Weitao Xiong , Jiacheng Deng , Peng Li , Tianyu Huang , Zhiyang Dou , Cheng Lin , Sai-Kit Yeung , Yuan Liu