中文
相关论文

相关论文: Gaslight, Gatekeep, V1-V3: Early Visual Cortex Ali…

200 篇论文

Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. Recent geometry-grounded vision models, such as VGGT~\cite{wang2025vggt},…

机器人学 · 计算机科学 2025-09-22 An Dinh Vuong , Minh Nhat Vu , Ian Reid

As large models gain traction, vision-language-action (VLA) systems are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection, progress has mainly focused on controlling simple gripper…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Jiawei He , Danshi Li , Xinqiang Yu , Zekun Qi , Wenyao Zhang , Jiayi Chen , Zhaoxiang Zhang , Zhizheng Zhang , Li Yi , He Wang

Vision-Language Models (VLMs) excel at photorealistic generation, yet often struggle to represent abstract meaning such as idiomatic interpretations of noun compounds. To study whether high visual fidelity interferes with idiomatic…

计算与语言 · 计算机科学 2026-04-21 Wei He

Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potential applications in enhanced radiology reporting. However, these models require large…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Zachary Huemann , Samuel Church , Joshua D. Warner , Daniel Tran , Xin Tie , Alan B McMillan , Junjie Hu , Steve Y. Cho , Meghan Lubner , Tyler J. Bradshaw

Large Vision-Language Models (LVLMs) have achieved impressive progress across various applications but remain vulnerable to malicious queries that exploit the visual modality. Existing alignment approaches typically fail to resist malicious…

密码学与安全 · 计算机科学 2025-11-18 Yitong Zhang , Jia Li , Liyi Cai , Ge Li

Learning visuomotor policies from scarce expert demonstrations remains a core challenge in robotic manipulation. A primary hurdle lies in distilling high-dimensional RGB representations into control-relevant geometry without overfitting.…

机器人学 · 计算机科学 2026-05-18 Davide Buoso , Andrea Protopapa , Stefano Di Carlo , Francesca Pistilli , Giuseppe Averta

Explanation mechanisms are increasingly used to support transparency and trust in vision-language models (VLMs), particularly in settings where model decisions require human oversight. However, the robustness of these explanations remains…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Narges Babadi , Hadis Karimipour

Vision Language Models (VLMs) demonstrate promising chart comprehension capabilities. Yet, prior explorations of their visualization literacy have been limited to assessing their response correctness and fail to explore their internal…

人机交互 · 计算机科学 2025-04-09 Lianghan Dong , Anamaria Crisan

When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro and GPT-5 only respond with text, which can be difficult for…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Brandon Collins , Logan Bolton , Hung Huy Nguyen , Mohammad Reza Taesiri , Trung Bui , Anh Totti Nguyen

We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate…

计算与语言 · 计算机科学 2026-05-04 Gregory N. Frank

A small but growing body of work has shown that machine learning models which better align with human vision have also exhibited higher robustness to adversarial examples, raising the question: can human-like perception make models more…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Blaine Hoak , Kunyang Li , Patrick McDaniel

A key solution to visual question answering (VQA) exists in how to fuse visual and language features extracted from an input image and question. We show that an attention mechanism that enables dense, bi-directional interactions between the…

计算机视觉与模式识别 · 计算机科学 2018-12-04 Duy-Kien Nguyen , Takayuki Okatani

Visual perception, the brain's construction of a stable world from sensory data, faces several long-standing, fundamental challenges. While often studied separately, these problems have resisted a single, unifying computational framework.…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Xin Li

Large Language Models (LLMs) are increasingly used in educational settings as interactive tools for collaboration. However, their tendency toward sycophancy, aligning with user beliefs even when incorrect, raises concerns for learning and…

人机交互 · 计算机科学 2026-05-22 Cansu Koyuturk , Sabrina Guidotti , Dimitri Ognibene

Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on binary classification across two clinical…

人工智能 · 计算机科学 2026-03-31 Doan Nam Long Vu , Simone Balloccu

Vision-Language Models (VLMs) deliver impressive performance in understanding visual content with language instructions. However, redundancy in vision tokens results in the degenerated inference efficiency of VLMs, which hinders real-time…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Qinyu Chen , Jiawen Qi

Vision-language models enable open-vocabulary object grounding through natural language queries, under the implicit assumption that semantically equivalent descriptions yield consistent outputs. We examine this assumption using a controlled…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Dawar Jyoti Deka , Amit Sethi , Syed Mohammad Ali

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image data without explicit 3D geometric supervision, resulting in…

Vision-language models (VLMs) have demonstrated strong performance in image geolocation, a capability further sharpened by frontier multimodal large reasoning models (MLRMs). This poses a significant privacy risk, as these widely accessible…

密码学与安全 · 计算机科学 2026-02-19 Ruixin Yang , Ethan Mendes , Arthur Wang , James Hays , Sauvik Das , Wei Xu , Alan Ritter