中文
相关论文

相关论文: Pippo: High-Resolution Multi-View Humans from a Si…

200 篇论文

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Junyi Chen , Tong He , Zhoujie Fu , Pengfei Wan , Kun Gai , Weicai Ye

Multi-view diffusion models have recently emerged as a powerful paradigm for novel view synthesis, yet the underlying mechanism that enables their view-consistency remains unclear. In this work, we first verify that the attention maps of…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Minkyung Kwon , Jinhyeok Choi , Jiho Park , Seonghu Jeon , Jinhyuk Jang , Junyoung Seo , Minseop Kwak , Jin-Hwa Kim , Seungryong Kim

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables conditional generation from inputs such as text, depth, pose,…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Duong H. Le , Tuan Pham , Sangho Lee , Christopher Clark , Aniruddha Kembhavi , Stephan Mandt , Ranjay Krishna , Jiasen Lu

This paper presents a new large multiview dataset called HUMBI for human body expressions with natural clothing. The goal of HUMBI is to facilitate modeling view-specific appearance and geometry of gaze, face, hand, body, and garment from…

计算机视觉与模式识别 · 计算机科学 2020-05-26 Zhixuan Yu , Jae Shin Yoon , In Kyu Lee , Prashanth Venkatesh , Jaesik Park , Jihun Yu , Hyun Soo Park

Recent advancements in open-world 3D object generation have been remarkable, with image-to-3D methods offering superior fine-grained control over their text-to-3D counterparts. However, most existing models fall short in simultaneously…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Minghua Liu , Ruoxi Shi , Linghao Chen , Zhuoyang Zhang , Chao Xu , Xinyue Wei , Hansheng Chen , Chong Zeng , Jiayuan Gu , Hao Su

High resolution panoramic video content is paramount for immersive experiences in Virtual Reality, but is non-trivial to collect as it requires specialized equipment and intricate camera setups. In this work, we introduce VideoPanda, a…

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Zehuan Huang , Yuan-Chen Guo , Xingqiao An , Yunhan Yang , Yangguang Li , Zi-Xin Zou , Ding Liang , Xihui Liu , Yan-Pei Cao , Lu Sheng

There is a rapidly growing interest in controlling consistency across multiple generated images using diffusion models. Among various methods, recent works have found that simply manipulating attention modules by concatenating features from…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Jiaojiao Fan , Haotian Xue , Qinsheng Zhang , Yongxin Chen

The one-shot talking-head generation learns to synthesize a talking-head video with one source portrait image under the driving of same or different identity video. Usually these methods require plane-based pixel transformations via Jacobin…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Luchuan Song , Pinxin Liu , Guojun Yin , Chenliang Xu

Despite recent progress in diffusion models, generating realistic head portraits from novel viewpoints remains a significant challenge. Most current approaches are constrained to limited angular ranges, predominantly focusing on frontal or…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Stathis Galanakis , Alexandros Lattas , Stylianos Moschoglou , Bernhard Kainz , Stefanos Zafeiriou

Recent advancements in diffusion techniques have propelled image and video generation to unprecedented levels of quality, significantly accelerating the deployment and application of generative AI. However, 3D shape generation technology…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Yangguang Li , Zi-Xin Zou , Zexiang Liu , Dehu Wang , Yuan Liang , Zhipeng Yu , Xingchao Liu , Yuan-Chen Guo , Ding Liang , Wanli Ouyang , Yan-Pei Cao

Beyond the superiority of the text-to-image diffusion model in generating high-quality images, recent studies have attempted to uncover its potential for adapting the learned semantic knowledge to visual perception tasks. In this work,…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Kangyang Xie , Binbin Yang , Hao Chen , Meng Wang , Cheng Zou , Hui Xue , Ming Yang , Chunhua Shen

CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Antonio D'Orazio , Maria Rosaria Briglia , Donato Crisostomi , Dario Loi , Emanuele Rodolà , Iacopo Masi

We propose \textbf{DMV3D}, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model incorporates a triplane NeRF representation and can denoise…

计算机视觉与模式识别 · 计算机科学 2023-11-16 Yinghao Xu , Hao Tan , Fujun Luan , Sai Bi , Peng Wang , Jiahao Li , Zifan Shi , Kalyan Sunkavalli , Gordon Wetzstein , Zexiang Xu , Kai Zhang

In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., Multi-Image Composition (MICo), remains a challenging problem, partly hindered by the lack of high-quality training data. To…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Xinyu Wei , Kangrui Cen , Hongyang Wei , Zhen Guo , Kai Cui , Bairui Li , Zeqing Wang , Jinrui Zhang , Lei Zhang

We study how to synthesize novel views of human body from a single image. Though recent deep learning based methods work well for rigid objects, they often fail on objects with large articulation, like human bodies. The core step of…

计算机视觉与模式识别 · 计算机科学 2018-04-13 Hao Zhu , Hao Su , Peng Wang , Xun Cao , Ruigang Yang

Recent advancements in image generation have made significant progress, yet existing models present limitations in perceiving and generating an arbitrary number of interrelated images within a broad context. This limitation becomes…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Ying Shen , Yizhe Zhang , Shuangfei Zhai , Lifu Huang , Joshua M. Susskind , Jiatao Gu

Video models have recently been applied with success to problems in content generation, novel view synthesis, and, more broadly, world simulation. Many applications in generation and transfer rely on conditioning these models, typically…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Edoardo A. Dominici , Thomas Deixelberger , Konstantinos Vardis , Markus Steinberger

We address the challenge of creating 3D assets for household articulated objects from a single image. Prior work on articulated object creation either requires multi-view multi-state input, or only allows coarse control over the generation…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Jiayi Liu , Denys Iliash , Angel X. Chang , Manolis Savva , Ali Mahdavi-Amiri

We present GASPACHO, a method for generating photorealistic, controllable renderings of human-object interactions from multi-view RGB video. Unlike prior work that reconstructs only the human and treats objects as background, GASPACHO…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Aymen Mir , Arthur Moreau , Helisa Dhamo , Zhensong Zhang , Gerard Pons-Moll , Eduardo Pérez-Pellitero