中文
相关论文

相关论文: Baba Is AI: Break the Rules to Beat the Benchmark

200 篇论文

AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is scattered across heterogeneous sources that often contradict one another, new information can…

机器学习 · 计算机科学 2026-05-19 Haonian Ji , Kaiwen Xiong , Siwei Han , Peng Xia , Shi Qiu , Yiyang Zhou , Jiaqi Liu , Jinlong Li , Bingzhou Li , Zeyu Zheng , Cihang Xie , Huaxiu Yao

Large Language Models (LLMs) represent a landmark achievement in Artificial Intelligence (AI), demonstrating unprecedented proficiency in procedural tasks such as text generation, code completion, and conversational coherence. These…

人工智能 · 计算机科学 2025-05-07 Schaun Wheeler , Olivier Jeunen

The massive successes of large language models (LLMs) encourage the emerging exploration of LLM-augmented Autonomous Agents (LAAs). An LAA is able to generate actions with its core LLM and interact with environments, which facilitates the…

As the complexity and scope of games increase, game testing, also called playtesting, becomes an essential activity to ensure the quality of video games. Yet, the manual, ad-hoc nature of game testing leaves space for automation. In this…

Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination, opaque operation, and subjective preferences. To address…

Human beings use compositionality to generalise from past experiences to novel experiences. We assume a separation of our experiences into fundamental atomic components that can be recombined in novel ways to support our ability to engage…

计算与语言 · 计算机科学 2023-12-20 Kevin Denamganaï , Sondess Missaoui , James Alfred Walker

Video Games are boring when they are too easy, and frustrating when they are too hard. In terms of providing game experience such as enjoyment to the player by match players with different levels of ability to player ability, We assume that…

人机交互 · 计算机科学 2021-10-22 Junjie Xu

Atari games have been a long-standing benchmark in the reinforcement learning (RL) community for the past decade. This benchmark was proposed to test general competency of RL algorithms. Previous work has achieved good average performance…

Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely on historical dialogue for responses, these agents must…

人工智能 · 计算机科学 2025-06-30 Peijie Yu , Yifan Yang , Jinjian Li , Zelong Zhang , Haorui Wang , Xiao Feng , Feng Zhang

The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments. Agents trained solely to achieve their objectives may adopt…

人工智能 · 计算机科学 2025-12-09 Dena Mujtaba , Brian Hu , Anthony Hoogs , Arslan Basharat

As machine learning (ML) is more tightly woven into society, it is imperative that we better characterize ML's strengths and limitations if we are to employ it responsibly. Existing benchmark environments for ML, such as board and video…

机器学习 · 计算机科学 2022-07-22 Eric Pulick , Shubham Bharti , Yiding Chen , Vladimir Menkov , Yonatan Mintz , Paul Kantor , Vicki M. Bier

Recent advancements in natural language and Large Language Models (LLMs) have enabled AI agents to simulate human-like interactions within virtual worlds. However, these interactions still face limitations in complexity and flexibility,…

计算与语言 · 计算机科学 2023-07-25 Yuanzhi Liang , Linchao Zhu , Yi Yang

This paper introduces the Word Synchronization Challenge, a novel benchmark to evaluate large language models (LLMs) in Human-Computer Interaction (HCI). This benchmark uses a dynamic game-like framework to test LLMs ability to mimic human…

人机交互 · 计算机科学 2026-01-15 Tanguy Cazalets , Joni Dambre

In role-playing games a Game Master (GM) is the player in charge of the game, who must design the challenges the players face and narrate the outcomes of their actions. In this work we discuss some challenges to model GMs from an…

计算与语言 · 计算机科学 2023-10-03 Santiago Góngora , Luis Chiruzzo , Gonzalo Méndez , Pablo Gervás

User-configured chatbots built on top of large language models are increasingly available through centralized marketplaces such as OpenAI's GPT Store. While these platforms enforce usage policies intended to prevent harmful or inappropriate…

计算与语言 · 计算机科学 2025-12-22 David Rodriguez , William Seymour , Jose M. Del Alamo , Jose Such

Many real-world multi-party negotiations unfold as sequences of binding, action-level commitments rather than a single final outcome, yet this regime remains under-studied in existing benchmarks. We introduce a benchmark and evaluation…

多智能体系统 · 计算机科学 2026-05-14 Leo Benac , Jonas Raedler , Zilin Ma , Finale Doshi-Velez

AI ethics is an emerging field with multiple, competing narratives about how to best solve the problem of building human values into machines. Two major approaches are focused on bias and compliance, respectively. But neither of these ideas…

人工智能 · 计算机科学 2023-02-24 Thomas Krendl Gilbert , Megan Welle Brozek , Andrew Brozek

Recent advances in large multimodal models have enabled new opportunities in embodied AI, particularly in robotic manipulation. These models have shown strong potential in generalization and reasoning, but achieving reliable and responsible…

机器人学 · 计算机科学 2025-12-05 Lei Zhang , Ju Dong , Kaixin Bai , Minheng Ni , Zoltan-Csaba Marton , Zhaopeng Chen , Jianwei Zhang

Building embodied autonomous agents capable of participating in social interactions with humans is one of the main challenges in AI. This problem motivated many research directions on embodied language use. Current approaches focus on…

机器学习 · 计算机科学 2021-04-28 Grgur Kovač , Rémy Portelas , Katja Hofmann , Pierre-Yves Oudeyer

Large Language Models (LLMs) have shown remarkable promise in communicating with humans. Their potential use as artificial partners with humans in sociological experiments involving conversation is an exciting prospect. But how viable is…

人工智能 · 计算机科学 2025-02-04 James Flamino , Mohammed Shahid Modi , Boleslaw K. Szymanski , Brendan Cross , Colton Mikolajczyk