English
Related papers

Related papers: Visual Commonsense in Pretrained Unimodal and Mult…

200 papers

Multimodal Large Language Models (MLLMs) have made notable advances in visual understanding, yet their abilities to recognize objects modified by specific attributes remain an open question. To address this, we explore MLLMs' reasoning…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Jiaxuan Li , Junwen Mo , MinhDuc Vo , Akihiro Sugimoto , Hideki Nakayama

Causality knowledge is vital to building robust AI systems. Deep learning models often perform poorly on tasks that require causal reasoning, which is often derived using some form of commonsense knowledge not immediately available in the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-23 Aman Chadha , Vinija Jain

Large scale visual understanding is challenging, as it requires a model to handle the widely-spread and imbalanced distribution of <subject, relation, object> triples. In real-world scenarios with large numbers of objects and relations,…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Ji Zhang , Yannis Kalantidis , Marcus Rohrbach , Manohar Paluri , Ahmed Elgammal , Mohamed Elhoseiny

Bottom-up and top-down visual cues are two types of information that helps the visual saliency models. These salient cues can be from spatial distributions of the features (space-based saliency) or contextual / task-dependent features…

Computer Vision and Pattern Recognition · Computer Science 2018-07-05 Nevrez Imamoglu , Wataru Shimoda , Chi Zhang , Yuming Fang , Asako Kanezaki , Keiji Yanai , Yoshifumi Nishida

Reporting bias arises when people assume that some knowledge is universally understood and hence, do not necessitate explicit elaboration. In this paper, we focus on the wide existence of reporting bias in visual-language datasets, embodied…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Qiyu Wu , Mengjie Zhao , Yutong He , Lang Huang , Junya Ono , Hiromi Wakaki , Yuki Mitsufuji

In our work, we explore the synergistic capabilities of pre-trained vision-and-language models (VLMs) and large language models (LLMs) on visual commonsense reasoning (VCR) problems. We find that VLMs and LLMs-based decision pipelines are…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Kaiwen Zhou , Kwonjoon Lee , Teruhisa Misu , Xin Eric Wang

Multimodal learning has mainly focused on learning large models on, and fusing feature representations from, different modalities for better performances on downstream tasks. In this work, we take a detour from this trend and study the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-08 Yifeng Shi , Marc Niethammer

Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In this work, we demonstrate that VLMs are selectively blind. They modulate the amount of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Wan-Cyuan Fan , Jiayun Luo , Declan Kutscher , Leonid Sigal , Ritwik Gupta

Infants learn not only object categories but also fine-grained visual attributes such as color, size, and texture from limited experience. Prior infant-scale vision--language models have mainly been evaluated on object recognition, leaving…

Machine Learning · Computer Science 2026-05-14 Patrick Batsell , Satoshi Tsutsui , Bihan Wen

Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying on visual evidence. While prior work attempts to mitigate this issue through decoding…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Seulbi Lee , Sangheum Hwang

Humans understand language based on the rich background knowledge about how the physical world works, which in turn allows us to reason about the physical world through language. In addition to the properties of objects (e.g., boats require…

Computation and Language · Computer Science 2019-08-09 Maxwell Forbes , Ari Holtzman , Yejin Choi

This study evaluates the effectiveness of Vision Language Models (VLMs) in representing and utilizing multimodal content for fact-checking. To be more specific, we investigate whether incorporating multimodal content improves performance…

Computation and Language · Computer Science 2024-12-09 Recep Firat Cekinel , Pinar Karagoz , Cagri Coltekin

Language models (LMs) trained on raw texts have no direct access to the physical world. Gordon and Van Durme (2013) point out that LMs can thus suffer from reporting bias: texts rarely report on common facts, instead focusing on the unusual…

Computation and Language · Computer Science 2022-09-27 Fangyu Liu , Julian Martin Eisenschlos , Jeremy R. Cole , Nigel Collier

We study the perception of color illusions by vision-language models. Color illusion, where a person's visual system perceives color differently from actual color, is well-studied in human vision. However, it remains underexplored whether…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Lingjun Mao , Zineng Tang , Alane Suhr

While vision-language models (VLMs) have achieved remarkable performance improvements recently, there is growing evidence that these models also posses harmful biases with respect to social attributes such as gender and race. Prior studies…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Phillip Howard , Avinash Madasu , Tiep Le , Gustavo Lujan Moreno , Vasudev Lal

Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reasoning remains unclear. Existing benchmarks report strong…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yuandong Wang , Yao Cui , Yuxin Zhao , Zhen Yang , Yangfu Zhu , Zhenzhou Shao

Perceptual constancy is the ability to maintain stable perceptions of objects despite changes in sensory input, such as variations in distance, angle, or lighting. This ability is crucial for visual understanding in a dynamic world. Here,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Haoran Sun , Bingyang Wang , Suyang Yu , Yijiang Li , Qingying Gao , Haiyun Lyu , Lianyu Huang , Zelong Hong , Jiahui Ge , Qianli Ma , Hang He , Yifan Zhou , Lingzi Guo , Lantao Mei , Maijunxian Wang , Dezhi Luo , Hokin Deng

Reasoning is an important ability that we learn from a very early age. Yet, reasoning is extremely hard for algorithms. Despite impressive recent progress that has been reported on tasks that necessitate reasoning, such as visual question…

Computer Vision and Pattern Recognition · Computer Science 2020-01-10 Jingxiang Lin , Unnat Jain , Alexander G. Schwing

The design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to…

Computation and Language · Computer Science 2021-10-25 Fangyu Liu , Emanuele Bugliarello , Edoardo Maria Ponti , Siva Reddy , Nigel Collier , Desmond Elliott

Current work on multimodal machine translation (MMT) has suggested that the visual modality is either unnecessary or only marginally beneficial. We posit that this is a consequence of the very simple, short and repetitive sentences used in…

Computation and Language · Computer Science 2019-06-04 Ozan Caglayan , Pranava Madhyastha , Lucia Specia , Loïc Barrault
‹ Prev 1 4 5 6 7 8 10 Next ›