中文
相关论文

相关论文: GeoWeaver: Grounding Visual Tokens with Geometric …

200 篇论文

We propose a model to learn visually grounded word embeddings (vis-w2v) to capture visual notions of semantic relatedness. While word embeddings trained using text have been extremely successful, they cannot uncover notions of semantic…

计算机视觉与模式识别 · 计算机科学 2016-06-30 Satwik Kottur , Ramakrishna Vedantam , José M. F. Moura , Devi Parikh

Spatial reasoning is a core aspect of human intelligence that allows perception, inference and planning in 3D environments. However, current vision-language models (VLMs) struggle to maintain geometric coherence and cross-view consistency…

人工智能 · 计算机科学 2025-12-03 Qiyao Xue , Weichen Liu , Shiqi Wang , Haoming Wang , Yuyang Wu , Wei Gao

Previous works leveraging video models for image-to-3D scene generation tend to suffer from geometric distortions and blurry content. In this paper, we renovate the pipeline of image-to-3D scene generation by unlocking the potential of…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Yuhao Wan , Lijuan Liu , Jingzhi Zhou , Zihan Zhou , Xuying Zhang , Dongbo Zhang , Shaohui Jiao , Qibin Hou , Ming-Ming Cheng

Visual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated…

计算机视觉与模式识别 · 计算机科学 2022-06-09 Li Yang , Yan Xu , Chunfeng Yuan , Wei Liu , Bing Li , Weiming Hu

The advancement of Large Vision-Language Models (LVLMs) requires precise local region-based reasoning that faithfully grounds the model's logic in actual visual evidence. However, existing datasets face limitations in scalability due to…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Byeonggeuk Lim , Kyeonghyun Kim , JungMin Yun , YoungBin Kim

Geometric problem solving, as a typical multimodal reasoning problem, has attracted much attention and made great progress recently, however most of works focus on plane geometry while usually fail in solid geometry due to 3D spatial…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Ruoran Xu , Haoyu Cheng , Bin Dong , Qiufeng Wang

Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and grounding. Existing methods, such as proxy encoding and geometry…

计算机视觉与模式识别 · 计算机科学 2025-01-24 Tianren Ma , Lingxi Xie , Yunjie Tian , Boyu Yang , Qixiang Ye

Visual perception and language understanding are - fundamental components of human intelligence, enabling them to understand and reason about objects and their interactions. It is crucial for machines to have this capacity to reason using…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Thao Minh Le

World models serve as essential building blocks toward Artificial General Intelligence (AGI), enabling intelligent agents to predict future states and plan actions by simulating complex physical interactions. However, existing interactive…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Junyi Chen , Haoyi Zhu , Xianglong He , Yifan Wang , Jianjun Zhou , Wenzheng Chang , Yang Zhou , Zizun Li , Zhoujie Fu , Jiangmiao Pang , Tong He

Generative AI has enabled novice designers to quickly create professional-looking visual representations for product concepts. However, novices have limited domain knowledge that could constrain their ability to write prompts that…

人机交互 · 计算机科学 2026-03-30 Sirui Tao , Ivan Liang , Cindy Peng , Zhiqing Wang , Srishti Palani , Steven P. Dow

Multimodal reasoning remains a fundamental challenge in artificial intelligence. Despite substantial advances in text-based reasoning, even state-of-the-art models such as GPT-o3 struggle to maintain strong performance in multimodal…

计算与语言 · 计算机科学 2025-09-09 Hao Liang , Ruitao Wu , Bohan Zeng , Junbo Niu , Wentao Zhang , Bin Dong

Multimodal Large Language Models (MLLMs) have demonstrated impressive progress in single-image grounding and general multi-image understanding. Recently, some methods begin to address multi-image grounding. However, they are constrained by…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Shurong Zheng , Yousong Zhu , Hongyin Zhao , Fan Yang , Yufei Zhan , Ming Tang , Jinqiao Wang

Obtaining large-scale, high-quality reasoning data is crucial for improving the geometric reasoning capabilities of multi-modal large language models (MLLMs). However, existing data generation methods, whether based on predefined tem plates…

计算与语言 · 计算机科学 2025-10-06 Weiming Wu , Jin Ye , Zi-kang Wang , Zhi Zhou , Yu-Feng Li , Lan-Zhe Guo

Automatic math problem solving has recently attracted increasing attention as a long-standing AI benchmark. In this paper, we focus on solving geometric problems, which requires a comprehensive understanding of textual descriptions, visual…

人工智能 · 计算机科学 2022-01-12 Jiaqi Chen , Jianheng Tang , Jinghui Qin , Xiaodan Liang , Lingbo Liu , Eric P. Xing , Liang Lin

Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains significantly limited. Current methods often rely on passive…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Meng Cao , Xingyu Li , Xue Liu , Ian Reid , Xiaodan Liang

Vision-Language Models (VLMs) have shown remarkable capabilities in spatial reasoning, yet they remain fundamentally limited to qualitative precision and lack the computational precision required for real-world robotics. Current approaches…

机器人学 · 计算机科学 2026-03-05 Yi Han , Enshen Zhou , Shanyu Rong , Jingkun An , Pengwei Wang , Zhongyuan Wang , Cheng Chi , Lu Sheng , Shanghang Zhang

The rapid advancement of generative models has intensified the challenge of detecting and interpreting visual forgeries, necessitating robust frameworks for image forgery detection while providing reasoning as well as localization. While…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Ipsita Praharaj , Yukta Butala , Badrikanath Praharaj , Yash Butala

Geospatial Knowledge Graphs (GeoKGs) model geoentities (e.g., places and natural features) and spatial relationships in an interconnected manner, providing strong knowledge support for geographic applications, including data retrieval,…

人工智能 · 计算机科学 2024-10-25 Lei Hu , Wenwen Li , Yunqiang Zhu

Vision-Language Models (VLMs) in remote sensing often fail at complex analytical tasks, a limitation stemming from their end-to-end training paradigm that bypasses crucial reasoning steps and leads to unverifiable outputs. To address this…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Jiaqi Liu , Lang Sun , Ronghao Fu , Bo Yang

A proper evaluation of stories generated for a sequence of images -- the task commonly referred to as visual storytelling -- must consider multiple aspects, such as coherence, grammatical correctness, and visual grounding. In this work, we…

人工智能 · 计算机科学 2023-10-30 Aditya K Surikuchi , Sandro Pezzelle , Raquel Fernández