中文
相关论文

相关论文: UniUGG: Unified 3D Understanding and Generation vi…

200 篇论文

Recent advances in 3D generation have improved the fidelity and geometric details of synthesized 3D assets. However, due to the inherent ambiguity of single-view observations and the lack of robust global structural priors caused by limited…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Wenyue Chen , Wenjue Chen , Peng Li , Qinghe Wang , Xu Jia , Heliang Zheng , Rongfei Jia , Yuan Liu , Ronggang Wang

Generating high-quality meshes with complex structures and realistic surfaces is the primary goal of 3D generative models. Existing methods typically employ sequence data or deformable tetrahedral grids for mesh generation. However,…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Ruowei Wang , Jiaqi Li , Dan Zeng , Xueqi Ma , Zixiang Xu , Jianwei Zhang , Qijun Zhao

3D shape generation aims to produce innovative 3D content adhering to specific conditions and constraints. Existing methods often decompose 3D shapes into a sequence of localized components, treating each element in isolation without…

计算机视觉与模式识别 · 计算机科学 2024-07-15 Ruikai Cui , Weizhe Liu , Weixuan Sun , Senbo Wang , Taizhang Shang , Yang Li , Xibin Song , Han Yan , Zhennan Wu , Shenzhou Chen , Hongdong Li , Pan Ji

Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently optimize understanding via sparse text signals and…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Songsong Yu , Yuxin Chen , Ying Shan , Yanwei Li

We propose L3DG, the first approach for generative 3D modeling of 3D Gaussians through a latent 3D Gaussian diffusion formulation. This enables effective generative 3D modeling, scaling to generation of entire room-scale scenes which can be…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Barbara Roessle , Norman Müller , Lorenzo Porzi , Samuel Rota Bulò , Peter Kontschieder , Angela Dai , Matthias Nießner

Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related tasks. While current unified frameworks based on large language…

人工智能 · 计算机科学 2026-01-21 Guocun Wang , Kenkun Liu , Jing Lin , Guorui Song , Jian Li , Xiaoguang Han

Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the UniEmo, a unified framework that seamlessly integrates…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yijie Zhu , Lingsen Zhang , Zitong Yu , Rui Shao , Tao Tan , Liqiang Nie

Three-dimensional scene generation holds significant potential in gaming, film, and virtual reality. However, most existing methods adopt a single-step generation process, making it difficult to balance scene complexity with minimal user…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Jiacheng Hong , Kunzhen Wu , Mingrui Yu , Yichao Gu , Shengze Xue , Shuangjiu Xiao , Deli Dong

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Xiang Deng , Feng Gao , Yong Zhang , Youxin Pang , Xu Xiaoming , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Chengyue Wu , Xiaokang Chen , Zhiyu Wu , Yiyang Ma , Xingchao Liu , Zizheng Pan , Wen Liu , Zhenda Xie , Xingkai Yu , Chong Ruan , Ping Luo

Generative self-supervised learning on graphs, particularly graph masked autoencoders, has emerged as a popular learning paradigm and demonstrated its efficacy in handling non-Euclidean data. However, several remaining issues limit the…

机器学习 · 计算机科学 2024-02-14 Yijun Tian , Chuxu Zhang , Ziyi Kou , Zheyuan Liu , Xiangliang Zhang , Nitesh V. Chawla

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

机器人学 · 计算机科学 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

Lighting has a strong influence on visual appearance, yet understanding and representing lighting in images remains notoriously difficult. Various lighting representations exist, such as environment maps, irradiance, spherical harmonics, or…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Zitian Zhang , Iliyan Georgiev , Michael Fischer , Yannick Hold-Geoffroy , Jean-François Lalonde , Valentin Deschaintre

Unified remote sensing multimodal models exhibit a pronounced spatial reversal curse: Although they can accurately recognize and describe object locations in images, they often fail to faithfully execute the same spatial relations during…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Weiyu Zhang , Yuan Hu , Yong Li , Yu Liu

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Haoyu Zhang , Meng Liu , Zaijing Li , Haokun Wen , Weili Guan , Yaowei Wang , Liqiang Nie

Dexterous grasping aims to produce diverse grasping postures with a high grasping success rate. Regression-based methods that directly predict grasping parameters given the object may achieve a high success rate but often lack diversity.…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Jiaxin Lu , Hao Kang , Haoxiang Li , Bo Liu , Yiding Yang , Qixing Huang , Gang Hua

D shape generation is a fundamental operation in computer graphics. While significant progress has been made, especially with recent deep generative models, it remains a challenge to synthesize high-quality shapes with rich geometric…

图形学 · 计算机科学 2022-05-31 Jie Yang , Kaichun Mo , Yu-Kun Lai , Leonidas J. Guibas , Lin Gao

Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames.…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Trong-Thuan Nguyen , Pha Nguyen , Jackson Cothren , Alper Yilmaz , Khoa Luu

Remote sensing vision tasks require extensive labeled data across multiple, interconnected domains. However, current generative data augmentation frameworks are task-isolated, i.e., each vision task requires training an independent…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Datao Tang , Hao Wang , Yudeng Xin , Hui Qiao , Dongsheng Jiang , Yin Li , Zhiheng Yu , Xiangyong Cao

Limited by the computational efficiency and accuracy, generating complex 3D scenes remains a challenging problem for existing generation networks. In this work, we propose DepthGAN, a novel method of generating depth maps with only semantic…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Yidi Li , Yiqun Wang , Zhengda Lu , Jun Xiao