中文
相关论文

相关论文: Reinforcing 3D Understanding in Point-VLMs via Geo…

200 篇论文

Recent advancements in reinforcement learning (RL) have enhanced the reasoning abilities of large language models (LLMs), yet the impact on multimodal LLMs (MLLMs) is limited. Particularly in vision-intensive tasks like geometric reasoning,…

计算与语言 · 计算机科学 2025-09-23 Guizhen Chen , Weiwen Xu , Hao Zhang , Hou Pong Chan , Deli Zhao , Anh Tuan Luu , Yu Rong

Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial drift, and depth violations in generated videos. To address this limitation, we propose a…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Tengjiao Yin , Jinglei Shi , Heng Guo , Xi Wang

While Multimodal Large Language Models (MLLMs) demonstrate proficiency in 2D scenes, extending their perceptual intelligence to 3D point cloud understanding remains a significant challenge. Current approaches focus primarily on aligning 3D…

Achieving robust perception-reasoning synergy is a central goal for advanced Vision-Language Models (VLMs). Recent advancements have pursued this goal via architectural designs or agentic workflows. However, these approaches are often…

人工智能 · 计算机科学 2026-05-15 Haozhe Wang , Qixin Xu , Changpeng Wang , Taofeng Xue , Chong Peng , Wenhu Chen , Fangzhen Lin

Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Shihua Zhang , Qiuhong Shen , Shizun Wang , Tianbo Pan , Xinchao Wang

Multimodal Large Language Models (MLLMs) struggle with complex geometric reasoning, largely because "black box" outcome-based supervision fails to distinguish between lucky guesses and rigorous deduction. To address this, we introduce a…

机器学习 · 计算机科学 2026-01-09 Jianlong Chen , Daocheng Fu , Shengze Xu , Jiawei Chen , Yuan Feng , Yue Yang , Junchi Yan , Hongyuan Zha , Renqiu Xia

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models,…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Jiaxin Zhang , Junjun Jiang , Haijie Li , Youyu Chen , Kui Jiang , Dave Zhenyu Chen

Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, maintaining multi-view consistency in edited results remains challenging, and the extreme scarcity of 3D-consistent editing paired…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Jiyuan Wang , Chunyu Lin , Lei Sun , Zhi Cao , Yuyang Yin , Lang Nie , Zhenlong Yuan , Xiangxiang Chu , Yunchao Wei , Kang Liao , Guosheng Lin

The rapid development of Large Multimodal Models (LMMs) has led to remarkable progress in 2D visual understanding; however, extending these capabilities to 3D scene understanding remains a significant challenge. Existing approaches…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Hongpei Zheng , Lintao Xiang , Qijun Yang , Qian Lin , Hujun Yin

Reward engineering, the manual specification of reward functions to induce desired agent behavior, remains a fundamental challenge in multi-agent reinforcement learning. This difficulty is amplified by credit assignment ambiguity,…

人工智能 · 计算机科学 2026-01-14 Haoran Su , Yandong Sun , Congjia Yu

While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space remains limited. To address this gap, we introduce geometrically referenced 3D scene…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jiangye Yuan , Gowri Kumar , Baoyuan Wang

The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Tuo Xiang , Xuemiao Xu , Bangzhen Liu , Jinyi Li , Yong Li , Shengfeng He

Auxiliary lines are essential for solving complex geometric problems but remain challenging for large vision-language models (LVLMs). Recent attempts construct auxiliary lines via code-driven rendering, a strategy that relies on accurate…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Shasha Guo , Liang Pang , Xi Wang , Yanling Wang , Huawei Shen , Jing Zhang

Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approaches rely on end-to-end video understanding or large-scale…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Fernando Ropero , Erkin Turkoz , Daniel Matos , Junqing Du , Antonio Ruiz , Yanfeng Zhang , Lu Liu , Mingwei Sun , Yongliang Wang

Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still face two fundamental limitations. First, they struggle with coarse semantic alignment, often…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Weimin Bai , Yubo Li , Weijian Luo , Zeqiang Lai , Yequan Wang , Wenzheng Chen , He Sun

Geometric reasoning inherently requires "thinking with constructions" -- the dynamic manipulation of visual aids to bridge the gap between problem conditions and solutions. However, existing Multimodal Large Language Models (MLLMs) are…

人工智能 · 计算机科学 2026-03-20 Haokun Zhao , Wanshi Xu , Haidong Yuan , Songjun Cao , Long Ma , Yanghua Xiao

Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. While raw geospatial imagery is abundant, the amount of task-direct supervision falls far…

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in their learned…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Haoyu Wu , Diankun Wu , Tianyu He , Junliang Guo , Yang Ye , Yueqi Duan , Jiang Bian

Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Zizun Li , Haoyu Guo , Runzhe Teng , Chunhua Shen , Tong He

360 panoramic images are increasingly used in virtual reality, autonomous driving, and robotics for holistic scene understanding. However, current Vision-Language Models (VLMs) struggle with 3D spatial reasoning on Equirectangular…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Zekai Lin , Xu Zheng
‹ 上一页 1 2 3 10 下一页 ›