中文
相关论文

相关论文: Spot The Ball: A Benchmark for Visual Social Infer…

200 篇论文

Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in comprehending complex visual content. However, the mechanisms underlying how VLMs process visual information remain largely unexplored. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Omri Kaduri , Shai Bagon , Tali Dekel

Visual relationship detection is an intermediate image understanding task that detects two objects and classifies a predicate that explains the relationship between two objects in an image. The three components are linguistically and…

计算机视觉与模式识别 · 计算机科学 2019-04-17 Jaewon Jung , Jongyoul Park

Predicting outcomes in sports is important for teams, leagues, bettors, media, and fans. Given the growing amount of player tracking data, sports analytics models are increasingly utilizing spatially-derived features built upon player…

机器学习 · 计算机科学 2022-07-29 Peter Xenopoulos , Claudio Silva

Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a critical bottleneck. Strikingly, MLLMs can produce correct answers even while misinterpreting crucial visual elements, masking these…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Aditya Kanade , Tanuja Ganu

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language…

机器人学 · 计算机科学 2024-11-01 Guanyan Chen , Meiling Wang , Te Cui , Yao Mu , Haoyang Lu , Tianxing Zhou , Zicai Peng , Mengxiao Hu , Haizhou Li , Yuan Li , Yi Yang , Yufeng Yue

Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant…

计算与语言 · 计算机科学 2025-10-14 Shiqi Chen , Tongyao Zhu , Ruochen Zhou , Jinghan Zhang , Siyang Gao , Juan Carlos Niebles , Mor Geva , Junxian He , Jiajun Wu , Manling Li

Humans are social creatures who readily recognize various social interactions from simple display of moving shapes. While previous research has often focused on visual features, we examine what semantic representations that humans employ to…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Yiling Yun , Hongjing Lu

Predicting athletes' performance has relied mostly on statistical data. Besides the traditional data, various types of data, including video, have become available. However, it is challenging to use them for deep learning, especially when…

人机交互 · 计算机科学 2023-04-07 Jaeyoun You , Jinhan Choi , Ho-Jae Shin , Bongwon Suh

The lack of reasoning capabilities in Vision-Language Models (VLMs) has remained at the forefront of research discourse. We posit that this behavior stems from a reporting bias in their training data. That is, how people communicate about…

计算与语言 · 计算机科学 2026-02-27 Amita Kamath , Jack Hessel , Khyathi Chandu , Jena D. Hwang , Kai-Wei Chang , Ranjay Krishna

Discovering social relations in images can make machines better interpret the behavior of human beings. However, automatically recognizing social relations in images is a challenging task due to the significant gap between the domains of…

计算机视觉与模式识别 · 计算机科学 2019-01-11 Meng Zhang , Xinchen Liu , Wu Liu , Anfu Zhou , Huadong Ma , Tao Mei

Recent research on Vision Language Models (VLMs) suggests that they rely on inherent biases learned during training to respond to questions about visual properties of an image. These biases are exacerbated when VLMs are asked highly…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Saurav Sengupta , Nazanin Moradinasab , Jiebei Liu , Donald E. Brown

While Multimodal Large Language Models (MLLMs) excel at many vision tasks, it is unknown if they exhibit human-like perceptual behaviors. To evaluate this, we introduce HVSBench, the first large-scale benchmark with over 85,000 samples…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Jiaying Lin , Shuquan Ye , Dan Xu , Wanli Ouyang , Rynson W. H. Lau

Large vision language models (VLMs) have demonstrated significant potential for integration into daily life, making it crucial for them to incorporate human values when making decisions in real-world situations. This paper introduces VIVA,…

计算与语言 · 计算机科学 2024-10-11 Zhe Hu , Yixiao Ren , Jing Li , Yu Yin

Understanding social interaction, which encompasses perceiving numerous and subtle multimodal cues, inferring unobservable mental states and relations, and dynamically predicting others' behavior, is the foundation for achieving…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Fanqi Kong , Weiqin Zu , Xinyu Chen , Yaodong Yang , Song-Chun Zhu , Xue Feng

Iconicity, the resemblance between linguistic form and meaning, is pervasive in signed languages, offering a natural testbed for visual grounding. For vision-language models (VLMs), the challenge is to recover such essential mappings from…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Onur Keleş , Aslı Özyürek , Gerardo Ortega , Kadir Gökgöz , Esam Ghaleb

Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain limited to single-agent or text-only environments. In contrast, real-world scenarios often…

人工智能 · 计算机科学 2026-04-14 Zelai Xu , Zhexuan Xu , Xiangmin Yi , Huining Yuan , Mo Guang , Kaiwen Long , Xinlei Chen , Yi Wu , Chao Yu , Yu Wang

We present Saliency Benchmark (SalBench), a novel benchmark designed to assess the capability of Large Vision-Language Models (LVLM) in detecting visually salient features that are readily apparent to humans, such as a large circle amidst a…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yasser Dahou , Ngoc Dung Huynh , Phuc H. Le-Khac , Wamiq Reyaz Para , Ankit Singh , Sanath Narayan

Visual Language Models (VLMs) are often used for zero-shot detection of visual attributes in the image. We present a zero-shot evaluation of open-source VLMs for privacy-related attribute recognition. We identify the attributes for which…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Olena Hrynenko , Darya Baranouskaya , Alina Elena Baia , Andrea Cavallaro

Vision-Language Models (VLMs) are expected to be capable of reasoning with commonsense knowledge as human beings. One example is that humans can reason where and when an image is taken based on their knowledge. This makes us wonder if,…

计算机视觉与模式识别 · 计算机科学 2024-01-01 Gengyuan Zhang , Yurui Zhang , Kerui Zhang , Volker Tresp

Humans can effortlessly describe what they see, yet establishing a shared representational format between vision and language remains a significant challenge. Emerging evidence suggests that human brain representations in both vision and…

神经元与认知 · 定量生物学 2025-07-30 Katerina Marie Simkova , Adrien Doerig , Clayton Hickey , Ian Charest