English
Related papers

Related papers: Visuospatial short-term memory and dorsal visual g…

200 papers

Vision Language Models (VLMs), exemplified by GPT-4V, adeptly integrate text and vision modalities. This integration enhances Large Language Models' ability to mimic human perception, allowing them to process image inputs. Despite VLMs'…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Messi H. J. Lee , Jacob M. Montgomery , Calvin K. Lai

Memorability of an image is a characteristic determined by the human observers' ability to remember images they have seen. Yet recent work on image memorability defines it as an intrinsic property that can be obtained independent of the…

Computer Vision and Pattern Recognition · Computer Science 2019-03-07 Erdem Akagunduz , Adrian G. Bors , Karla K. Evans

We propose an efficient multi-view stereo (MVS) network for infering depth value from multiple RGB images. Recent studies have shown that mapping the geometric relationship in real space to neural network is an essential topic of the MVS…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Zihang Wan

Vision language models (VLMs) are designed to extract relevant visuospatial information from images. Some research suggests that VLMs can exhibit humanlike scene understanding, while other investigations reveal difficulties in their ability…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Sangeet Khemlani , Tyler Tran , Nathaniel Gyory , Anthony M. Harrison , Wallace E. Lawson , Ravenna Thielstrom , Hunter Thompson , Taaren Singh , J. Gregory Trafton

Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning process capable of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Wenbo Hu , Jingli Lin , Yilin Long , Yunlong Ran , Lihan Jiang , Yifan Wang , Chenming Zhu , Runsen Xu , Tai Wang , Jiangmiao Pang

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Zuntao Liu , Yi Du , Taimeng Fu , Shaoshu Su , Cherie Ho , Chen Wang

Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reasoning tasks that require multimodal or multilingual…

Machine Learning · Computer Science 2025-07-09 Wenyi Wu , Zixuan Song , Kun Zhou , Yifei Shao , Zhiting Hu , Biwei Huang

Human memory exhibits significant vulnerability in cognitive tasks and daily life. Comparisons between visual working memory and new perceptual input (e.g., during cognitive tasks) can lead to unintended memory distortions. Previous studies…

Neurons and Cognition · Quantitative Biology 2025-07-31 Yuang Cao , Jiachen Zou , Chen Wei , Quanying Liu

Accurate automatic classification of major tissue classes and the cerebrospinal fluid in pediatric MR scans of early childhood brains remains a challenge. A poor and highly variable grey matter and white matter contrast on T1-weighted MR…

Image and Video Processing · Electrical Eng. & Systems 2020-05-08 Nataliya Portman , Paule-J Toussaint , Alan C. Evans McConnell

Vision-and-Language Navigation (VLN) in large-scale urban environments requires embodied agents to ground linguistic instructions in complex scenes and recall relevant experiences over extended time horizons. Prior modular pipelines offer…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Lixuan He , Haoyu Dong , Zhenxing Chen , Yangcheng Yu , Jie Feng , Yong Li

Vision Language Models (VLMs) extend remarkable capabilities of text-only large language models and vision-only models, and are able to learn from and process multi-modal vision-text input. While modern VLMs perform well on a number of…

Computation and Language · Computer Science 2025-07-22 Hannah Sterz , Jonas Pfeiffer , Ivan Vulić

To measure the volume of specific image structures, a typical approach is to first segment those structures using a neural network trained on voxel-wise (strong) labels and subsequently compute the volume from the segmentation. A more…

Image and Video Processing · Electrical Eng. & Systems 2020-04-14 Oliver Werner , Kimberlin M. H. van Wijnen , Wiro J. Niessen , Marius de Groot , Meike W. Vernooij , Florian Dubost , Marleen de Bruijne

We study the Voronoi volume function (VVF) -- the distribution of cell volumes (or inverse local number density) in the Voronoi tessellation of any set of cosmological tracers (galaxies/haloes). We show that the shape of the VVF of biased…

Cosmology and Nongalactic Astrophysics · Physics 2020-05-27 Aseem Paranjape , Shadab Alam

In this work, we propose valley-coupled spin-hall memories (VSH-MRAMs) based on monolayer WSe2. The key features of the proposed memories are (a) the ability to switch magnets with perpendicular magnetic anisotropy (PMA) via VSH effect and…

Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relationship in 3D space over time, largely due to the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Shengchao Zhou , Yuxin Chen , Yuying Ge , Wei Huang , Jiehong Lin , Ying Shan , Xiaojuan Qi

Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) architecture that has been designed to address the vanishing and exploding gradient problems of conventional RNNs. Unlike feedforward neural networks, RNNs have cyclic…

Neural and Evolutionary Computing · Computer Science 2014-02-06 Haşim Sak , Andrew Senior , Françoise Beaufays

The visual crowding makes it difficult to identify the patterns in peripheral vision, but the neural mechanism for this phenomenon is still unclear because of different opinions. In order to study the separation effect of V1 under different…

Neurons and Cognition · Quantitative Biology 2019-05-27 Xieyi Liu , Junjun Zhang , Ling Li

The computational expense of redundant vision tokens in Large Vision-Language Models (LVLMs) has led many existing methods to compress them via a vision projector. However, this compression may lose visual information that is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Ze Feng , Jiang-jiang Liu , Sen Yang , Lingyu Xiao , Zhibin Quan , Zhenhua Feng , Wankou Yang , Jingdong Wang

Introduction: Quantitative Susceptibility Mapping (QSM) is generally acquired with full brain coverage, even though many QSM brain-iron studies focus on the deep grey matter (DGM) region only. Reducing the spatial coverage to the DGM…

Quantitative Methods · Quantitative Biology 2021-06-02 Xuanyu Zhu , Yang Gao , Feng Liu , Stuart Crozier , Hongfu Sun

Valley-spin hall (VSH) effect in monolayer WSe2 has been shown to exhibit highly beneficial features for nonvolatile memory (NVM) design. Key advantages of VSH-based magnetic random-access memory (VSH-MRAM) over spin orbit torque (SOT)-MRAM…

Systems and Control · Electrical Eng. & Systems 2022-09-20 Karam Cho , Sumeet Kumar Gupta