English
Related papers

Related papers: Generation Models Know Space: Unleashing Implicit …

200 papers

Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Duo Zheng , Shijia Huang , Yanyang Li , Liwei Wang

Recent advances in 3D generation have improved the fidelity and geometric details of synthesized 3D assets. However, due to the inherent ambiguity of single-view observations and the lack of robust global structural priors caused by limited…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Wenyue Chen , Wenjue Chen , Peng Li , Qinghe Wang , Xu Jia , Heliang Zheng , Rongfei Jia , Yuan Liu , Ronggang Wang

Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chun-Hsiao Yeh , Shengyi Qian , Manchen Wang , Yi Ma , Joseph Tighe , Fanyi Xiao

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision-Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge. Existing 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Haijier Chen , Bo Xu , Shoujian Zhang , Haoze Liu , Jiaxuan Lin , Jingrong Wang

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Duo Zheng , Shijia Huang , Liwei Wang

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models…

Robotics · Computer Science 2025-11-25 Tao Lin , Gen Li , Yilei Zhong , Yanwen Zou , Yuxin Du , Jiting Liu , Encheng Gu , Bo Zhao

Leveraging 3D information within Multimodal Large Language Models (MLLMs) has recently shown significant advantages for indoor scene understanding. However, existing methods, including those using explicit ground-truth 3D positional…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Chushan Zhang , Ruihan Lu , Jinguang Tong , Yikai Wang , Hongdong Li

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in 2D visual tasks but still exhibit limited physical spatial awareness when processing real-world visual streams. Recently, feed-forward geometric foundation…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Chongyu Wang , Ting Huang , Chunyu Sun , Xinyu Ning , Di Wang , Hao Tang

Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing approaches typically mitigate this limitation either…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Jiahua Chen , Qihong Tang , Weinong Wang , Qi Fan

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image data without explicit 3D geometric supervision, resulting in…

Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relationship in 3D space over time, largely due to the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Shengchao Zhou , Yuxin Chen , Yuying Ge , Wei Huang , Jiehong Lin , Ying Shan , Xiaojuan Qi

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiaxin Zhang , Junjun Jiang , Haijie Li , Youyu Chen , Kui Jiang , Dave Zhenyu Chen

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Haoyu Zhang , Meng Liu , Zaijing Li , Haokun Wen , Weili Guan , Yaowei Wang , Liqiang Nie

Trained on internet-scale video data, generative world models are increasingly recognized as powerful world simulators that can generate consistent and plausible dynamics over structure, motion, and physics. This raises a natural question:…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Kevin Zhang , Kuangzhi Ge , Xiaowei Chi , Renrui Zhang , Shaojun Shi , Zhen Dong , Sirui Han , Shanghang Zhang

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Haoyu Zhen , Xiaowen Qiu , Peihao Chen , Jincheng Yang , Xin Yan , Yilun Du , Yining Hong , Chuang Gan

Generating complete 3D objects under partial occlusions (i.e., amodal scenarios) is a practically important yet challenging problem, as large portions of object geometry are unobserved in real-world scenarios. Existing approaches either…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Junwei Zhou , Yu-Wing Tai

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Zhixue Fang , Xu He , Songlin Tang , Haoxian Zhang , Qingfeng Li , Xiaoqiang Liu , Pengfei Wan , Kun Gai

The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Jinlong Li , Cristiano Saltori , Fabio Poiesi , Nicu Sebe

Estimating accurate camera poses, 3D scene geometry, and object motion from in-the-wild videos is a long-standing challenge for classical structure from motion pipelines due to the presence of dynamic objects. Recent learning-based methods…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Zhuoyuan Wu , Xurui Yang , Jiahui Huang , Yue Wang , Jun Gao
‹ Prev 1 2 3 10 Next ›