中文
相关论文

相关论文: GameVerse: Can Vision-Language Models Learn from V…

200 篇论文

Large Language Models (LLMs) exhibit robust problem-solving capabilities for diverse tasks. However, most LLM-based agents are designed as specific task solvers with sophisticated prompt engineering, rather than agents capable of learning…

人工智能 · 计算机科学 2024-06-10 Wenqi Zhang , Ke Tang , Hai Wu , Mengna Wang , Yongliang Shen , Guiyang Hou , Zeqi Tan , Peng Li , Yueting Zhuang , Weiming Lu

Given the rapidly growing capabilities of vision-language models (VLMs), extending them to interactive decision-making tasks such as video games has emerged as a promising frontier. However, existing approaches either rely on large-scale…

Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, even though the visual change is obvious to humans. This…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Xiaoxiao Sun , Mingyang Li , Kun Yuan , Min Woo Sun , Mark Endo , Shengguang Wu , Changlin Li , Yuhui Zhang , Zeyu Wang , Serena Yeung-Levy

With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for the sector's sustained growth. Recent advancements in Vision-Language Models (VLMs) offer…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Mohammad Reza Taesiri , Abhijay Ghildyal , Saman Zadtootaghaj , Nabajeet Barman , Cor-Paul Bezemer

Recent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yujie Lu , Dongfu Jiang , Wenhu Chen , William Yang Wang , Yejin Choi , Bill Yuchen Lin

We can think of Visual Question Answering as a (multimodal) conversation between a human and an AI system. Here, we explore the sensitivity of Vision Language Models (VLMs) through the lens of cooperative principles of conversation proposed…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Monika Shah , Sudarshan Balaji , Somdeb Sarkhel , Sanorita Dey , Deepak Venugopal

We present a perception in reflection paradigm designed to transcend the limitations of current large vision-language models (LVLMs), which are expected yet often fail to achieve perfect perception initially. Specifically, we propose…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Yana Wei , Liang Zhao , Kangheng Lin , En Yu , Yuang Peng , Runpei Dong , Jianjian Sun , Haoran Wei , Zheng Ge , Xiangyu Zhang , Vishal M. Patel

Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper,…

人工智能 · 计算机科学 2025-01-03 Shudong Liu , Yiqiao Jin , Cheng Li , Derek F. Wong , Qingsong Wen , Lichao Sun , Haipeng Chen , Xing Xie , Jindong Wang

Text-based role-playing models can imitate character styles, yet they often fail to reflect a scene's atmosphere and evolving tension, both essential for immersive applications such as Virtual Reality (VR) games and interactive narratives.…

人工智能 · 计算机科学 2026-05-07 Miao Wang , Yuling Shi , Yijiang Li , Yeheng Chen , Xiaodong Gu , Bin Li , Bo Gao , Yaduan Ruan

Multimodal Vision Language Models (VLMs) have emerged as a transformative topic at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through both visual and textual…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Zongxia Li , Xiyang Wu , Hongyang Du , Fuxiao Liu , Huy Nghiem , Guangyao Shi

Social deduction games like Werewolf combine language, reasoning, and strategy, providing a testbed for studying natural language and social intelligence. However, most studies reduce the game to LLM-based self-play, yielding templated…

计算与语言 · 计算机科学 2025-10-14 Zirui Song , Yuan Huang , Junchang Liu , Haozhe Luo , Chenxi Wang , Lang Gao , Zixiang Xu , Mingfei Han , Xiaojun Chang , Xiuying Chen

Recent advancements in video-based large language models (Video LLMs) have witnessed the emergence of diverse capabilities to reason and interpret dynamic visual content. Among them, gameplay videos stand out as a distinctive data source,…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Meng Cao , Haoran Tang , Haoze Zhao , Hangyu Guo , Jiaheng Liu , Ge Zhang , Ruyang Liu , Qiang Sun , Ian Reid , Xiaodan Liang

Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their…

计算与语言 · 计算机科学 2025-12-01 Philip Schroeder , Ondrej Biza , Thomas Weng , Hongyin Luo , James Glass

Large Language Models (LLMs) have achieved remarkable success across diverse natural language tasks, yet the reward models employed for aligning LLMs often encounter challenges of reward hacking, where the approaches predominantly rely on…

计算与语言 · 计算机科学 2026-03-06 Biao Liu , Ning Xu , Junming Yang , Hao Xu , Xin Geng

Large vision-language models (LVLMs) have shown promising performance on a variety of vision-language tasks. However, they remain susceptible to hallucinations, generating outputs misaligned with visual content or instructions. While…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Jinrui Zhang , Teng Wang , Haigang Zhang , Ping Lu , Feng Zheng

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, struggling to simulate…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Fan Yang , Shurong Zheng , Hongyin Zhao , Yufei Zhan , Xin Li , Yousong Zhu , Chaoyang Zhao Ming Tang , Jinqiao Wang

Interaction between learner and feedback-giver has come into focus recently for post-training of Large Language Models (LLMs), through the use of reward models that judge the appropriateness of a model's response. In this paper, we…

Multimodal Large Language Models (MLLMs) have shown great potential in revolutionizing Graphical User Interface (GUI) automation. However, existing GUI models mostly rely on learning from nearly error-free offline trajectories, thus lacking…

人工智能 · 计算机科学 2025-06-10 Penghao Wu , Shengnan Ma , Bo Wang , Jiaheng Yu , Lewei Lu , Ziwei Liu

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Yiming Zhao , Yu Zeng , Yukun Qi , YaoYang Liu , Xikun Bao , Lin Chen , Zehui Chen , Qing Miao , Chenxi Liu , Jie Zhao , Feng Zhao

Language provides a natural interface to specify and evaluate performance on visual tasks. To realize this possibility, vision language models (VLMs) must successfully integrate visual and linguistic information. Our work compares VLMs to a…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Stephanie Fu , Tyler Bonnen , Devin Guillory , Trevor Darrell