中文
相关论文

相关论文: Gaslight, Gatekeep, V1-V3: Early Visual Cortex Ali…

200 篇论文

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,''…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Shaoxiong Zhan , Yanlin Lai , Zheng Liu , Hai Lin , Shen Li , Xiaodong Cai , Zijian Lin , Wen Huang , Hai-Tao Zheng

The brain processes visual inputs having structure over a large range of spatial scales. The precise mechanisms or algorithms used by the brain to achieve this feat are largely unknown and an open problem in visual neuroscience. In…

神经元与认知 · 定量生物学 2018-07-04 Keith Hayton , Dimitrios Moirogiannis , Marcelo Magnasco

Occlusion perception, a critical foundation for human-level spatial understanding, embodies the challenge of integrating visual recognition and reasoning. Though multimodal large language models (MLLMs) have demonstrated remarkable…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Zhaochen Liu , Kaiwen Gao , Shuyi Liang , Bin Xiao , Limeng Qiao , Lin Ma , Tingting Jiang

This paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering)…

计算机视觉与模式识别 · 计算机科学 2019-12-05 Luowei Zhou , Hamid Palangi , Lei Zhang , Houdong Hu , Jason J. Corso , Jianfeng Gao

Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations into action plans. While prior work has studied abstention in LLMs, existing benchmarks are…

机器人学 · 计算机科学 2026-05-21 Doguhan Yeke , Elif Su Temirel , Ananth Shreekumar , Brandon Lee , Dongyan Xu , Z Berkay Celik

How large language models (LLMs) align with the neural representation and computation of human language is a central question in cognitive science. Using representational geometry as a mechanistic lens, we addressed this by tracking…

神经元与认知 · 定量生物学 2026-02-10 Yixuan Liu , Zhiyuan Ma , Likai Tang , Runmin Gan , Xinche Zhang , Jinhao Li , Chao Xie , Sen Song

Detecting levels of psychological defence mechanisms in supportive conversations is inherently ambiguous. In the PsyDefDetect shared task at BioNLP 2026 the eight positive defence categories share surface language and differ only in…

计算与语言 · 计算机科学 2026-05-11 Philipp Steigerwald , Eric Rudolph , Jens Albrecht

AEC drawings encode geometry and semantics through symbols, layout conventions, and dense annotation, yet it remains unclear whether modern multimodal and vision-language models can reliably interpret this graphical language. We present…

人工智能 · 计算机科学 2026-01-09 Aleksei Kondratenko , Mussie Birhane , Houssame E. Hsain , Guido Maciocci

Vision-Language Pre-training (VLP) has advanced the performance of many vision-language tasks, such as image-text retrieval, visual entailment, and visual reasoning. The pre-training mostly utilizes lexical databases and image queries in…

计算与语言 · 计算机科学 2023-06-30 Yasmine Karoui , Rémi Lebret , Negar Foroutan , Karl Aberer

Adversarial attacks are known to succeed on classifiers, but it has been an open question whether more complex vision systems are vulnerable. In this paper, we study adversarial examples for vision and language models, which incorporate…

人工智能 · 计算机科学 2018-04-09 Xiaojun Xu , Xinyun Chen , Chang Liu , Anna Rohrbach , Trevor Darrell , Dawn Song

Understanding the physical world - governed by laws of motion, spatial relations, and causality - poses a fundamental challenge for multimodal large language models (MLLMs). While recent advances such as OpenAI o3 and GPT-4o demonstrate…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Zhuobai Dong , Junchao Yi , Ziyuan Zheng , Haochen Han , Xiangxi Zheng , Alex Jinpeng Wang , Fangming Liu , Linjie Li

Telling an LLM to "be enthusiastic" raises its sycophancy rate from 30\% to 50\% on a lightly-aligned model, but has zero effect on a strongly-aligned one. We define this gap as the alignment floor,…

人机交互 · 计算机科学 2026-05-29 Xing Zhang , Guanghui Wang , Yanwei Cui , Wei Qiu , Ziyuan Li , Bing Zhu , Peiyang He

This research investigates both explicit and implicit social biases exhibited by Vision-Language Models (VLMs). The key distinction between these bias types lies in the level of awareness: explicit bias refers to conscious, intentional…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Jen-tse Huang , Jiantong Qin , Jianping Zhang , Youliang Yuan , Wenxuan Wang , Jieyu Zhao

Although value-aligned language models (LMs) appear unbiased in explicit bias evaluations, they often exhibit stereotypes in implicit word association tasks, raising concerns about their fair usage. We investigate the mechanisms behind this…

计算与语言 · 计算机科学 2025-06-10 Lihao Sun , Chengzhi Mao , Valentin Hofmann , Xuechunzi Bai

Vision-language models (VLMs) exhibit a striking paradox: they can generate executable code that reconstructs a 3D scene from geometric primitives with correct object counts, classes, and approximate positions, yet the same models fail at…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Junze Liu , Kun Qian , Florian Dubost , Kai Zhong , Arvind Srinivasan , Nan Chen , Anping Wang , Sam Zhang , Alejandro Mottini , Qingjun Cui , Tian Wang

We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Through a large-scale computational study (n=3,213 scene-context pairs from COCO-2017) using Qwen-VL 30B and LLaVA-1.5-13B subject…

计算与语言 · 计算机科学 2026-03-06 Murad Farzulla

Vision-language models (VLMs) have demonstrated strong cross-modal capabilities, yet most work remains limited to 2D data and assumes binary supervision (i.e., positive vs. negative pairs), overlooking the continuous and structured…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Ailar Mahdizadeh , Puria Azadi Moghadam , Xiangteng He , Shahriar Mirabbasi , Panos Nasiopoulos , Leonid Sigal

Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vision-language models.…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Maijunxian Wang , Yijiang Li , Bingyang Wang , Tianwei Zhao , Ran Ji , Qingying Gao , Emmy Liu , Hokin Deng , Dezhi Luo

To interpret our surroundings, the brain uses a visual categorization process. Current theories and models suggest that this process comprises a hierarchy of different computations that transforms complex, high-dimensional inputs into…

神经元与认知 · 定量生物学 2024-06-21 Y. Duan , J. Zhan , J. Gross , R. A. A. Ince , P. G. Schyns

Autonomous driving systems often infer pedestrian yielding behavior from geometric and kinematic cues alone, limiting their ability to reason about visual scene context and age-dependent behavioral variability. This limitation can produce…

系统与控制 · 电气工程与系统科学 2026-04-28 Qingwen Pu , Kun Xie , Yuxiang Liu