English
Related papers

Related papers: Find Someone Who: Visual Commonsense Understanding…

200 papers

Visual understanding goes well beyond object recognition. With one glance at an image, we can effortlessly imagine the world beyond the pixels: for instance, we can infer people's actions, goals, and mental states. While this task is easy…

Computer Vision and Pattern Recognition · Computer Science 2019-03-27 Rowan Zellers , Yonatan Bisk , Ali Farhadi , Yejin Choi

Visual commonsense plays a vital role in understanding and reasoning about the visual world. While commonsense knowledge bases like ConceptNet provide structured collections of general facts, they lack visually grounded representations.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Xiangqing Shen , Fanfan Wang , Siwei Wu , Rui Xia

Even from a single frame of a still image, people can reason about the dynamic story of the image before, after, and beyond the frame. For example, given an image of a man struggling to stay afloat in water, we can reason that the man fell…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Jae Sung Park , Chandra Bhagavatula , Roozbeh Mottaghi , Ali Farhadi , Yejin Choi

We present a method for inferring diverse 3D models of human-object interactions from images. Reasoning about how humans interact with objects in complex scenes from a single 2D image is a challenging task given ambiguities arising from the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Xi Wang , Gen Li , Yen-Ling Kuo , Muhammed Kocabas , Emre Aksan , Otmar Hilliges

Large-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for suffering from the inherent…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Yuan Yao , Tianyu Yu , Ao Zhang , Mengdi Li , Ruobing Xie , Cornelius Weber , Zhiyuan Liu , Hai-Tao Zheng , Stefan Wermter , Tat-Seng Chua , Maosong Sun

We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Claire Yuqing Cui , Apoorv Khandelwal , Yoav Artzi , Noah Snavely , Hadar Averbuch-Elor

The aspiration for artificial general intelligence, fueled by the rapid progress of multimodal models, demands human-comparable performance across diverse environments. We propose HumanPCR, an evaluation suite for probing MLLMs' capacity…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Keliang Li , Hongze Shen , Hao Shi , Ruibing Hou , Hong Chang , Jie Huang , Chenghao Jia , Wen Wang , Yiling Wu , Dongmei Jiang , Shiguang Shan , Xilin Chen

Contextual commonsense inference is the task of generating various types of explanations around the events in a dyadic dialogue, including cause, motivation, emotional reaction, and others. Producing a coherent and non-trivial explanation…

Computation and Language · Computer Science 2022-11-04 Siqi Shen , Deepanway Ghosal , Navonil Majumder , Henry Lim , Rada Mihalcea , Soujanya Poria

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

Computation and Language · Computer Science 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the resulting semantic…

Computer Vision and Pattern Recognition · Computer Science 2019-11-26 Yongfei Liu , Bo Wan , Xiaodan Zhu , Xuming He

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Yi Yuan , Jingdong Chen , Le Wang

Person detection is a key problem for many computer vision tasks. While face detection has reached maturity, detecting people under a full variation of camera view-points, human poses, lighting conditions and occlusions is still a difficult…

Computer Vision and Pattern Recognition · Computer Science 2015-11-26 Tuan-Hung Vu , Anton Osokin , Ivan Laptev

Alternatively inferring on the visual facts and commonsense is fundamental for an advanced VQA system. This ability requires models to go beyond the literal understanding of commonsense. The system should not just treat objects as the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Difei Gao , Ruiping Wang , Shiguang Shan , Xilin Chen

Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile…

Computer Vision and Pattern Recognition · Computer Science 2023-03-13 Shixiang Tang , Cheng Chen , Qingsong Xie , Meilin Chen , Yizhou Wang , Yuanzheng Ci , Lei Bai , Feng Zhu , Haiyang Yang , Li Yi , Rui Zhao , Wanli Ouyang

Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Linhui Xiao , Xiaoshan Yang , Xiangyuan Lan , Yaowei Wang , Changsheng Xu

In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Lingru Zhou , Yiqi Gao , Manqing Zhang , Peng Wu , Peng Wang , Yanning Zhang

Visual Dialog requires an agent to engage in a conversation with humans grounded in an image. Many studies on Visual Dialog focus on the understanding of the dialog history or the content of an image, while a considerable amount of…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Shunyu Zhang , Xiaoze Jiang , Zequn Yang , Tao Wan , Zengchang Qin

Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention as a benchmark for Artificial Intelligence (AI). Although there…

Computation and Language · Computer Science 2025-09-16 Fenghua Cheng , Jinxiang Wang , Sen Wang , Zi Huang , Xue Li

When answering questions about an image, it not only needs knowing what -- understanding the fine-grained contents (e.g., objects, relationships) in the image, but also telling why -- reasoning over grounding visual cues to derive the…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Jianwei Yang , Jiayuan Mao , Jiajun Wu , Devi Parikh , David D. Cox , Joshua B. Tenenbaum , Chuang Gan

Visual grounding (VG) aims to establish fine-grained alignment between vision and language. Ideally, it can be a testbed for vision-and-language models to evaluate their understanding of the images and texts and their reasoning abilities…

Computer Vision and Pattern Recognition · Computer Science 2023-07-24 Zhihong Chen , Ruifei Zhang , Yibing Song , Xiang Wan , Guanbin Li
‹ Prev 1 2 3 10 Next ›