中文
相关论文

相关论文: MoCam: Unified Novel View Synthesis via Structured…

200 篇论文

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Korrawe Karunratanakul , Konpat Preechakul , Supasorn Suwajanakorn , Siyu Tang

Reconstructing 3D scenes and synthesizing novel views from sparse input views is a highly challenging task. Recent advances in video diffusion models have demonstrated strong temporal reasoning capabilities, making them a promising tool for…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Yuqi Zhang , Guanying Chen , Jiaxing Chen , Chuanyu Fu , Chuan Huang , Shuguang Cui

Recent advances in diffusion models have significantly improved 3D generation, enabling the use of assets generated from an image for embodied AI simulations. However, the one-to-many nature of the image-to-3D problem limits their use due…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Onat Şahin , Mohammad Altillawi , George Eskandar , Carlos Carbone , Ziyuan Liu

Human motion prediction is important for many virtual and augmented reality (VR/AR) applications such as collision avoidance and realistic avatar generation. Existing methods have synthesised body motion only from observed past motion,…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Haodong Yan , Zhiming Hu , Syn Schmitt , Andreas Bulling

Recent advancements in generative models have significantly improved novel view synthesis (NVS) from multi-view data. However, existing methods depend on external multi-view alignment processes, such as explicit pose estimation or…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Lingen Li , Zhaoyang Zhang , Yaowei Li , Jiale Xu , Wenbo Hu , Xiaoyu Li , Weihao Cheng , Jinwei Gu , Tianfan Xue , Ying Shan

Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motion synthesis tasks, and typically use a model of a small size…

计算机视觉与模式识别 · 计算机科学 2022-12-07 Jianxin Ma , Shuai Bai , Chang Zhou

Generating realistic 3D objects from single-view images requires natural appearance, 3D consistency, and the ability to capture multiple plausible interpretations of unseen regions. Existing approaches often rely on fine-tuning pretrained…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Pufan Li , Bi'an Du , Wei Hu

Image synthesis under multi-modal priors is a useful and challenging task that has received increasing attention in recent years. A major challenge in using generative models to accomplish this task is the lack of paired data containing all…

计算机视觉与模式识别 · 计算机科学 2022-06-13 Nithin Gopalakrishnan Nair , Wele Gedara Chaminda Bandara , Vishal M Patel

Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce ViewMask-1-to-3, formulating multi-view synthesis as a…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Ruishu Zhu , Zhihao Huang , Jiacheng Sun , Ping Luo , Hongyuan Zhang , Xuelong Li

We propose a novel framework for diffusion-based novel view synthesis in which we leverage external representations as conditions, harnessing their geometric and semantic correspondence properties for enhanced geometric consistency in…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Min-Seop Kwak , Minkyung Kwon , Jinhyeok Choi , Jiho Park , Seungryong Kim

We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Bharath Raj Nagoor Kani , Hsin-Ying Lee , Sergey Tulyakov , Shubham Tulsiani

In this paper, we present a novel diffusion model called that generates multiview-consistent images from a single-view image. Using pretrained large-scale 2D diffusion models, recent work Zero123 demonstrates the ability to generate…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Yuan Liu , Cheng Lin , Zijiao Zeng , Xiaoxiao Long , Lingjie Liu , Taku Komura , Wenping Wang

This paper introduces a novel approach to synthesize texture to dress up a given 3D object, given a text prompt. Based on the pretrained text-to-image (T2I) diffusion model, existing methods usually employ a project-and-inpaint approach, in…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Yuxin Liu , Minshan Xie , Hanyuan Liu , Tien-Tsin Wong

This paper describes the Qualcomm AI Research solution to the RealADSim-NVS challenge, hosted at the RealADSim Workshop at ICCV 2025. The challenge concerns novel view synthesis in street scenes, and participants are required to generate,…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Mohamed Omran , Farhad Zanjani , Davide Abati , Jens Petersen , Amirhossein Habibian

Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Yaqi Zhao , Wang Lin , Zijian Zhang , Miles Yang , Jingyuan Chen , Wentao Zhang , Zhao Zhong , Liefeng Bo

Reference-driven image completion, which restores missing regions in a target view using additional images, is particularly challenging when the target view differs significantly from the references. Existing generative methods rely solely…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Beibei Lin , Tingting Chen , Robby T. Tan

Pixel-space diffusion has recently re-emerged as a strong alternative to latent diffusion, enabling high-quality generation without pretrained autoencoders. However, standard pixel-space diffusion models receive relatively weak semantic…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Han Lin , Xichen Pan , Zun Wang , Yue Zhang , Chu Wang , Jaemin Cho , Mohit Bansal

Conventional geometry-based SLAM systems lack dense 3D reconstruction capabilities since their data association usually relies on feature correspondences. Additionally, learning-based SLAM systems often fall short in terms of real-time…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Zhongche Qu , Zhi Zhang , Cong Liu , Jianhua Yin

The task of novel view synthesis aims to generate unseen perspectives of an object or scene from a limited set of input images. Nevertheless, synthesizing novel views from a single image still remains a significant challenge in the realm of…

计算机视觉与模式识别 · 计算机科学 2023-10-05 Yifan Jiang , Hao Tang , Jen-Hao Rick Chang , Liangchen Song , Zhangyang Wang , Liangliang Cao

Generating street-view images from satellite imagery is a challenging task, particularly in maintaining accurate pose alignment and incorporating diverse environmental conditions. While diffusion models have shown promise in generative…

图像与视频处理 · 电气工程与系统科学 2025-06-04 Xianghui Ze , Zhenbo Song , Qiwei Wang , Jianfeng Lu , Yujiao Shi