English
Related papers

Related papers: SpatialFusion: Endowing Unified Image Generation w…

200 papers

Spatial understanding is essential for Multimodal Large Language Models (MLLMs) to support perception, reasoning, and planning in embodied environments. Despite recent progress, existing studies reveal that MLLMs still struggle with spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Wanyue Zhang , Yibin Huang , Yangbin Xu , JingJing Huang , Helu Zhi , Shuo Ren , Wang Xu , Jiajun Zhang

2D portrait animation has experienced significant advancements in recent years. Much research has utilized the prior knowledge embedded in large generative diffusion models to enhance high-quality image manipulation. However, most methods…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Xinya Ji , Gaspard Zoss , Prashanth Chandran , Lingchen Yang , Xun Cao , Barbara Solenthaler , Derek Bradley

Most models of generative AI for images assume that images are inherently low-dimensional objects embedded within a high-dimensional space. Additionally, it is often implicitly assumed that thematic image datasets form smooth or piecewise…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Leah Bar , Liron Mor Yosef , Shai Zucker , Neta Shoham , Inbar Seroussi , Nir Sochen

Recent advances in imitation learning have shown significant promise for robotic control and embodied intelligence. However, achieving robust generalization across diverse mounted camera observations remains a critical challenge. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Travis Davies , Jiahuan Yan , Xiang Chen , Yu Tian , Yueting Zhuang , Yiqi Huang , Luhui Hu

Current 6D object pose estimation methods usually require a 3D model for each object. These methods also require additional training in order to incorporate new objects. As a result, they are difficult to scale to a large number of objects…

Computer Vision and Pattern Recognition · Computer Science 2020-06-15 Keunhong Park , Arsalan Mousavian , Yu Xiang , Dieter Fox

Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data. Current methods primarily utilize task-specific models, while recent foundation models for…

Machine Learning · Computer Science 2025-10-16 Dominik J. Mühlematter , Lin Che , Ye Hong , Martin Raubal , Nina Wiedemann

This paper introduces TBAC-UniImage, a novel unified model for multimodal understanding and generation. We achieve this by deeply integrating a pre-trained Diffusion Model, acting as a generative ladder, with a Multimodal Large Language…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Junzhe Xu , Yuyang Yin , Xi Chen

Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). Current methods predominantly rely on verbal descriptive…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Meng Cao , Haokun Lin , Haoyuan Li , Haoran Tang , Rongtao Xu , Dong An , Xue Liu , Ian Reid , Xiaodan Liang

Acquiring accurate three-dimensional depth information conventionally requires expensive multibeam LiDAR devices. Recently, researchers have developed a less expensive option by predicting depth information from two-dimensional color…

Computer Vision and Pattern Recognition · Computer Science 2019-12-03 Peng Yin , Jianing Qian , Yibo Cao , David Held , Howie Choset

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Xianjin Wu , Dingkang Liang , Tianrui Feng , Kui Xia , Yumeng Zhang , Xiaofan Li , Xiao Tan , Xiang Bai

Recently, deep learning-based image enhancement algorithms achieved state-of-the-art (SOTA) performance on several publicly available datasets. However, most existing methods fail to meet practical requirements either for visual perception…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Tao Wang , Yong Li , Jingyang Peng , Yipeng Ma , Xian Wang , Fenglong Song , Youliang Yan

The remarkable success of Large Language Models (LLMs) has extended to the multimodal domain, achieving outstanding performance in image understanding and generation. Recent efforts to develop unified Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Hao Li , Changyao Tian , Jie Shao , Xizhou Zhu , Zhaokai Wang , Jinguo Zhu , Wenhan Dou , Xiaogang Wang , Hongsheng Li , Lewei Lu , Jifeng Dai

While recent advances in generative latent spaces have driven substantial progress in single-image generation, the optimal latent space for novel view synthesis (NVS) remains largely unexplored. In particular, NVS requires geometrically…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Wooseok Jang , Seonghu Jeon , Jisang Han , Jinhyeok Choi , Minkyung Kwon , Seungryong Kim , Saining Xie , Sainan Liu

Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motion synthesis tasks, and typically use a model of a small size…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Jianxin Ma , Shuai Bai , Chang Zhou

Generating high-quality 3D objects from textual descriptions remains a challenging problem due to computational cost, the scarcity of 3D data, and complex 3D representations. We introduce Geometry Image Diffusion (GIMDiffusion), a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Slava Elizarov , Ciara Rowles , Simon Donné

Generative depth estimation methods leverage the rich visual priors stored in pre-trained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, parameter updates during training lead to catastrophic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Hongkai Lin , Dingkang Liang , Mingyang Du , Xin Zhou , Xiang Bai

Recent advancements in autonomous driving, augmented reality, robotics, and embodied intelligence have necessitated 3D perception algorithms. However, current 3D perception methods, especially specialized small models, exhibit poor…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Fan Yang , Sicheng Zhao , Yanhao Zhang , Hui Chen , Haonan Lu , Jungong Han , Guiguang Ding

Image classification is a fundamental computer vision task and an important baseline for deep metric learning. In decades efforts have been made on enhancing image classification accuracy by using deep learning models while less attention…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Yunfeng Zhao , Huiyu Zhou , Fei Wu , Xifeng Wu

In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imagination to ensure efficient and accurate operation. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Xinyan Cai , Shiguang Wu , Dafeng Chi , Yuzheng Zhuang , Xingyue Quan , Jianye Hao , Qiang Guan

In this work, we present a novel framework built to simplify 3D asset generation for amateur users. To enable interactive generation, our method supports a variety of input modalities that can be easily provided by a human, including…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Yen-Chi Cheng , Hsin-Ying Lee , Sergey Tulyakov , Alexander Schwing , Liangyan Gui