English
Related papers

Related papers: Find Someone Who: Visual Commonsense Understanding…

200 papers

Image-text retrieval of natural scenes has been a popular research topic. Since image and text are heterogeneous cross-modal data, one of the key challenges is how to learn comprehensive yet unified representations to express the…

Computer Vision and Pattern Recognition · Computer Science 2019-10-14 Sijin Wang , Ruiping Wang , Ziwei Yao , Shiguang Shan , Xilin Chen

We propose the ambiguity problem for the foreground object segmentation task and motivate the importance of estimating and accounting for this ambiguity when designing vision systems. Specifically, we distinguish between images which lead…

Computer Vision and Pattern Recognition · Computer Science 2017-05-02 Danna Gurari , Kun He , Bo Xiong , Jianming Zhang , Mehrnoosh Sameki , Suyog Dutt Jain , Stan Sclaroff , Margrit Betke , Kristen Grauman

Context reasoning is critical in a wide variety of applications where current inputs need to be interpreted in the light of previous experience and knowledge. Both spatial and temporal contextual information play a critical role in the…

Computer Vision and Pattern Recognition · Computer Science 2019-09-26 Mengmi Zhang , Claire Tseng , Karla Montejo , Joseph Kwon , Gabriel Kreiman

Humans have the capacity to question what we see and to recognize when our vision is unreliable (e.g., when we realize that we are experiencing a visual illusion). Inspired by this capacity, we present MetaCOG: a hierarchical probabilistic…

Artificial Intelligence · Computer Science 2024-07-10 Marlene D. Berke , Zhangir Azerbayev , Mario Belledonne , Zenna Tavares , Julian Jara-Ettinger

We present a visually-grounded language understanding model based on a study of how people verbally describe objects in scenes. The emphasis of the model is on the combination of individual word meanings to produce meanings for complex…

Artificial Intelligence · Computer Science 2011-07-04 P. Gorniak , D. Roy

Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed for real-world applications such as multimodal code…

Computation and Language · Computer Science 2025-06-02 Iñigo Alonso , Gorka Azkune , Ander Salaberria , Jeremy Barnes , Oier Lopez de Lacalle

Always, some individuals in images are more important/attractive than others in some events such as presentation, basketball game or speech. However, it is challenging to find important people among all individuals in images directly based…

Computer Vision and Pattern Recognition · Computer Science 2017-11-07 Wei-Hong Li , Benchao Li , Wei-Shi Zheng

Multi-image spatial reasoning remains challenging for current multimodal large language models (MLLMs). While single-view perception is inherently 2D, reasoning over multiple views requires building a coherent scene understanding across…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Xuejun Zhang , Aditi Tiwari , Zhenhailong Wang , Heng Ji

Current multi-modal models exhibit a notable misalignment with the human visual system when identifying objects that are visually assimilated into the background. Our observations reveal that these multi-modal models cannot distinguish…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ruolin Shen , Xiaozhong Ji , Kai WU , Jiangning Zhang , Yijun He , HaiHua Yang , Xiaobin Hu , Xiaoyu Sun

We introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cognitive sciences which requires zero-shot visual inferences…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Tyler Bonnen , Stephanie Fu , Yutong Bai , Thomas O'Connell , Yoni Friedman , Nancy Kanwisher , Joshua B. Tenenbaum , Alexei A. Efros

"How does the person in the bounding box feel?" Achieving human-level recognition of the apparent emotion of a person in real world situations remains an unsolved task in computer vision. Facial expressions are not enough: body pose,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Yasaman Etesam , Özge Nilay Yalçın , Chuxuan Zhang , Angelica Lim

This paper proposes a novel approach to learn commonsense from images, instead of limited raw texts or costly constructed knowledge bases, for the commonsense reasoning problem in NLP. Our motivation comes from the fact that an image is…

Computation and Language · Computer Science 2020-10-13 Wanqing Cui , Yanyan Lan , Liang Pang , Jiafeng Guo , Xueqi Cheng

Many vision and language models suffer from poor visual grounding - often falling back on easy-to-learn language priors rather than basing their decisions on visual concepts in the image. In this work, we propose a generic approach called…

Computer Vision and Pattern Recognition · Computer Science 2019-10-29 Ramprasaath R. Selvaraju , Stefan Lee , Yilin Shen , Hongxia Jin , Shalini Ghosh , Larry Heck , Dhruv Batra , Devi Parikh

Dense video understanding requires answering several questions such as who is doing what to whom, with what, how, why, and where. Recently, Video Situation Recognition (VidSitu) is framed as a task for structured prediction of multiple…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Zeeshan Khan , C. V. Jawahar , Makarand Tapaswi

Human communication often relies on visual cues to resolve ambiguity. While humans can intuitively integrate these cues, AI systems often find it challenging to engage in sophisticated multimodal reasoning. We introduce VAGUE, a benchmark…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Heejeong Nam , Jinwoo Ahn , Keummin Ka , Jiwan Chung , Youngjae Yu

Answering complex questions about images is an ambitious goal for machine intelligence, which requires a joint understanding of images, text, and commonsense knowledge, as well as a strong reasoning ability. Recently, multimodal…

Computer Vision and Pattern Recognition · Computer Science 2021-12-17 Zhecan Wang , Haoxuan You , Liunian Harold Li , Alireza Zareian , Suji Park , Yiqing Liang , Kai-Wei Chang , Shih-Fu Chang

The ability of human beings to precisely recog- nize others intents is a significant mental activity in reasoning about actions, such as, what other people are doing and what they will do next. Recent research has revealed that human…

Human-Computer Interaction · Computer Science 2018-03-13 Xiang Zhang

Visual commonsense reasoning (VCR) is a challenging multi-modal task, which requires high-level cognition and commonsense reasoning ability about the real world. In recent years, large-scale pre-training approaches have been developed and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Cheng Yang , Rui Xu , Ye Guo , Peixiang Huang , Yiru Chen , Wenkui Ding , Zhongyuan Wang , Hong Zhou

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Lixiong Qin , Shilong Ou , Miaoxuan Zhang , Jiangning Wei , Yuhang Zhang , Xiaoshuai Song , Yuchen Liu , Mei Wang , Weiran Xu

Modern mobile devices are able to provide context-aware and personalized services to the users, by leveraging on their sensing capabilities to infer the activity and situation in which a person is currently involved. Current solutions for…

Machine Learning · Computer Science 2023-07-10 Mattia Giovanni Campana , Franca Delmastro