Omni-View:基于多视图图像的统一 3D 模型生成驱动理解
计算机视觉与模式识别
2026-02-02 v2
摘要
本文提出 Omni-View,扩展了基于多视图图像的统一多模态理解和生成到 3D 场景的能力,探索着“生成促进理解”的原理。Omni-View 由理解模型、纹理模块和几何模块组成,联合建模场景理解、新视图合成和几何估计,使 3D 场景理解和生成任务之间实现协同互动。通过设计,纹理模块负责外观合成,利用其时空建模能力;几何模块提供明确的几何约束,因而丰富了模型对 3D 场景的整体理解。经过两阶段训练,Omni-View 在 VSI-Bench benchmark 上取得了 55.4 的领先水平,超越了现有的专用 3D 理解模型,同时在新视图合成和 3D 场景生成中均表现出色。代码和预训练模型已开源于 https://github.com/AIDC-AI/Omni-View。
引用
@article{arxiv.2511.07222,
title = {Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images},
author = {JiaKui Hu and Shanshan Zhao and Qing-Guo Chen and Xuerui Qiu and Jialun Liu and Zhao Xu and Weihua Luo and Kaifu Zhang and Yanye Lu},
journal= {arXiv preprint arXiv:2511.07222},
year = {2026}
}
备注
Accepted by ICLR 2026