中文
相关论文

相关论文: Baba Is AI: Break the Rules to Beat the Benchmark

200 篇论文

Large language models (LLMs) are known to perform well on language tasks, but struggle with reasoning tasks. This paper explores the ability of LLMs to play the 2D puzzle game Baba is You, in which players manipulate rules by rearranging…

人工智能 · 计算机科学 2025-06-25 Fien van Wetten , Aske Plaat , Max van Duijn

The Keke AI Competition introduces an artificial agent competition for the game Baba is You - a Sokoban-like puzzle game where players can create rules that influence the mechanics of the game. Altering a rule can cause temporary or…

人工智能 · 计算机科学 2022-09-13 M Charity , Julian Togelius

This paper describes a new version of the mixed-initiative collaborative level designing system: Baba is Y'all, as well as the results of a user study on the system. Baba is Y'all is a prototype for AI-assisted game design in collaboration…

人机交互 · 计算机科学 2022-10-11 M Charity , Isha Dave , Ahmed Khalifa , Julian Togelius

We establish the undecidability of 2019 puzzle game Baba is You through a reduction from the Post correspondence problem. In particular, we consider a restricted form of the Post correspondence problem introduced by Neary (arXiv:1312.6700)…

计算复杂性 · 计算机科学 2024-06-17 Jonathan Geller

We present a collaborative mixed-initiative system for building levels for the puzzle game "Baba is You". Unlike previous mixed-initiative systems, Baba is Y'all is designed for collaborative asynchronous creation by multiple users over the…

人机交互 · 计算机科学 2020-06-04 Megan Charity , Ahmed Khalifa , Julian Togelius

While games have been used extensively as milestones to evaluate game-playing AI, there exists no standardised framework for reporting the obtained observations. As a result, it remains difficult to draw general conclusions about the…

人工智能 · 计算机科学 2020-07-07 Vanessa Volz , Boris Naujoks

The advancement of data-driven artificial intelligence (AI), particularly machine learning, heavily depends on large-scale benchmarks. Despite remarkable progress across domains ranging from pattern recognition to intelligent…

人工智能 · 计算机科学 2026-02-03 Chao Li , Shangdong Yang , Chiheng Zhan , Zhenxing Ge , Yujing Hu , Bingkun Bao , Xingguo Chen , Yang Gao

Decision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' decision-making through the lens of Game Theory. However,…

We introduce WebGames, a comprehensive benchmark suite designed to evaluate general-purpose web-browsing AI agents through a collection of 50+ interactive challenges. These challenges are specifically crafted to be straightforward for…

Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform…

计算与语言 · 计算机科学 2023-06-13 Aarohi Srivastava , Abhinav Rastogi , Abhishek Rao , Abu Awal Md Shoeb , Abubakar Abid , Adam Fisch , Adam R. Brown , Adam Santoro , Aditya Gupta , Adrià Garriga-Alonso , Agnieszka Kluska , Aitor Lewkowycz , Akshat Agarwal , Alethea Power , Alex Ray , Alex Warstadt , Alexander W. Kocurek , Ali Safaya , Ali Tazarv , Alice Xiang , Alicia Parrish , Allen Nie , Aman Hussain , Amanda Askell , Amanda Dsouza , Ambrose Slone , Ameet Rahane , Anantharaman S. Iyer , Anders Andreassen , Andrea Madotto , Andrea Santilli , Andreas Stuhlmüller , Andrew Dai , Andrew La , Andrew Lampinen , Andy Zou , Angela Jiang , Angelica Chen , Anh Vuong , Animesh Gupta , Anna Gottardi , Antonio Norelli , Anu Venkatesh , Arash Gholamidavoodi , Arfa Tabassum , Arul Menezes , Arun Kirubarajan , Asher Mullokandov , Ashish Sabharwal , Austin Herrick , Avia Efrat , Aykut Erdem , Ayla Karakaş , B. Ryan Roberts , Bao Sheng Loe , Barret Zoph , Bartłomiej Bojanowski , Batuhan Özyurt , Behnam Hedayatnia , Behnam Neyshabur , Benjamin Inden , Benno Stein , Berk Ekmekci , Bill Yuchen Lin , Blake Howald , Bryan Orinion , Cameron Diao , Cameron Dour , Catherine Stinson , Cedrick Argueta , César Ferri Ramírez , Chandan Singh , Charles Rathkopf , Chenlin Meng , Chitta Baral , Chiyu Wu , Chris Callison-Burch , Chris Waites , Christian Voigt , Christopher D. Manning , Christopher Potts , Cindy Ramirez , Clara E. Rivera , Clemencia Siro , Colin Raffel , Courtney Ashcraft , Cristina Garbacea , Damien Sileo , Dan Garrette , Dan Hendrycks , Dan Kilman , Dan Roth , Daniel Freeman , Daniel Khashabi , Daniel Levy , Daniel Moseguí González , Danielle Perszyk , Danny Hernandez , Danqi Chen , Daphne Ippolito , Dar Gilboa , David Dohan , David Drakard , David Jurgens , Debajyoti Datta , Deep Ganguli , Denis Emelin , Denis Kleyko , Deniz Yuret , Derek Chen , Derek Tam , Dieuwke Hupkes , Diganta Misra , Dilyar Buzan , Dimitri Coelho Mollo , Diyi Yang , Dong-Ho Lee , Dylan Schrader , Ekaterina Shutova , Ekin Dogus Cubuk , Elad Segal , Eleanor Hagerman , Elizabeth Barnes , Elizabeth Donoway , Ellie Pavlick , Emanuele Rodola , Emma Lam , Eric Chu , Eric Tang , Erkut Erdem , Ernie Chang , Ethan A. Chi , Ethan Dyer , Ethan Jerzak , Ethan Kim , Eunice Engefu Manyasi , Evgenii Zheltonozhskii , Fanyue Xia , Fatemeh Siar , Fernando Martínez-Plumed , Francesca Happé , Francois Chollet , Frieda Rong , Gaurav Mishra , Genta Indra Winata , Gerard de Melo , Germán Kruszewski , Giambattista Parascandolo , Giorgio Mariani , Gloria Wang , Gonzalo Jaimovitch-López , Gregor Betz , Guy Gur-Ari , Hana Galijasevic , Hannah Kim , Hannah Rashkin , Hannaneh Hajishirzi , Harsh Mehta , Hayden Bogar , Henry Shevlin , Hinrich Schütze , Hiromu Yakura , Hongming Zhang , Hugh Mee Wong , Ian Ng , Isaac Noble , Jaap Jumelet , Jack Geissinger , Jackson Kernion , Jacob Hilton , Jaehoon Lee , Jaime Fernández Fisac , James B. Simon , James Koppel , James Zheng , James Zou , Jan Kocoń , Jana Thompson , Janelle Wingfield , Jared Kaplan , Jarema Radom , Jascha Sohl-Dickstein , Jason Phang , Jason Wei , Jason Yosinski , Jekaterina Novikova , Jelle Bosscher , Jennifer Marsh , Jeremy Kim , Jeroen Taal , Jesse Engel , Jesujoba Alabi , Jiacheng Xu , Jiaming Song , Jillian Tang , Joan Waweru , John Burden , John Miller , John U. Balis , Jonathan Batchelder , Jonathan Berant , Jörg Frohberg , Jos Rozen , Jose Hernandez-Orallo , Joseph Boudeman , Joseph Guerr , Joseph Jones , Joshua B. Tenenbaum , Joshua S. Rule , Joyce Chua , Kamil Kanclerz , Karen Livescu , Karl Krauth , Karthik Gopalakrishnan , Katerina Ignatyeva , Katja Markert , Kaustubh D. Dhole , Kevin Gimpel , Kevin Omondi , Kory Mathewson , Kristen Chiafullo , Ksenia Shkaruta , Kumar Shridhar , Kyle McDonell , Kyle Richardson , Laria Reynolds , Leo Gao , Li Zhang , Liam Dugan , Lianhui Qin , Lidia Contreras-Ochando , Louis-Philippe Morency , Luca Moschella , Lucas Lam , Lucy Noble , Ludwig Schmidt , Luheng He , Luis Oliveros Colón , Luke Metz , Lütfi Kerem Şenel , Maarten Bosma , Maarten Sap , Maartje ter Hoeve , Maheen Farooqi , Manaal Faruqui , Mantas Mazeika , Marco Baturan , Marco Marelli , Marco Maru , Maria Jose Ramírez Quintana , Marie Tolkiehn , Mario Giulianelli , Martha Lewis , Martin Potthast , Matthew L. Leavitt , Matthias Hagen , Mátyás Schubert , Medina Orduna Baitemirova , Melody Arnaud , Melvin McElrath , Michael A. Yee , Michael Cohen , Michael Gu , Michael Ivanitskiy , Michael Starritt , Michael Strube , Michał Swędrowski , Michele Bevilacqua , Michihiro Yasunaga , Mihir Kale , Mike Cain , Mimee Xu , Mirac Suzgun , Mitch Walker , Mo Tiwari , Mohit Bansal , Moin Aminnaseri , Mor Geva , Mozhdeh Gheini , Mukund Varma T , Nanyun Peng , Nathan A. Chi , Nayeon Lee , Neta Gur-Ari Krakover , Nicholas Cameron , Nicholas Roberts , Nick Doiron , Nicole Martinez , Nikita Nangia , Niklas Deckers , Niklas Muennighoff , Nitish Shirish Keskar , Niveditha S. Iyer , Noah Constant , Noah Fiedel , Nuan Wen , Oliver Zhang , Omar Agha , Omar Elbaghdadi , Omer Levy , Owain Evans , Pablo Antonio Moreno Casares , Parth Doshi , Pascale Fung , Paul Pu Liang , Paul Vicol , Pegah Alipoormolabashi , Peiyuan Liao , Percy Liang , Peter Chang , Peter Eckersley , Phu Mon Htut , Pinyu Hwang , Piotr Miłkowski , Piyush Patil , Pouya Pezeshkpour , Priti Oli , Qiaozhu Mei , Qing Lyu , Qinlang Chen , Rabin Banjade , Rachel Etta Rudolph , Raefer Gabriel , Rahel Habacker , Ramon Risco , Raphaël Millière , Rhythm Garg , Richard Barnes , Rif A. Saurous , Riku Arakawa , Robbe Raymaekers , Robert Frank , Rohan Sikand , Roman Novak , Roman Sitelew , Ronan LeBras , Rosanne Liu , Rowan Jacobs , Rui Zhang , Ruslan Salakhutdinov , Ryan Chi , Ryan Lee , Ryan Stovall , Ryan Teehan , Rylan Yang , Sahib Singh , Saif M. Mohammad , Sajant Anand , Sam Dillavou , Sam Shleifer , Sam Wiseman , Samuel Gruetter , Samuel R. Bowman , Samuel S. Schoenholz , Sanghyun Han , Sanjeev Kwatra , Sarah A. Rous , Sarik Ghazarian , Sayan Ghosh , Sean Casey , Sebastian Bischoff , Sebastian Gehrmann , Sebastian Schuster , Sepideh Sadeghi , Shadi Hamdan , Sharon Zhou , Shashank Srivastava , Sherry Shi , Shikhar Singh , Shima Asaadi , Shixiang Shane Gu , Shubh Pachchigar , Shubham Toshniwal , Shyam Upadhyay , Shyamolima , Debnath , Siamak Shakeri , Simon Thormeyer , Simone Melzi , Siva Reddy , Sneha Priscilla Makini , Soo-Hwan Lee , Spencer Torene , Sriharsha Hatwar , Stanislas Dehaene , Stefan Divic , Stefano Ermon , Stella Biderman , Stephanie Lin , Stephen Prasad , Steven T. Piantadosi , Stuart M. Shieber , Summer Misherghi , Svetlana Kiritchenko , Swaroop Mishra , Tal Linzen , Tal Schuster , Tao Li , Tao Yu , Tariq Ali , Tatsu Hashimoto , Te-Lin Wu , Théo Desbordes , Theodore Rothschild , Thomas Phan , Tianle Wang , Tiberius Nkinyili , Timo Schick , Timofei Kornev , Titus Tunduny , Tobias Gerstenberg , Trenton Chang , Trishala Neeraj , Tushar Khot , Tyler Shultz , Uri Shaham , Vedant Misra , Vera Demberg , Victoria Nyamai , Vikas Raunak , Vinay Ramasesh , Vinay Uday Prabhu , Vishakh Padmakumar , Vivek Srikumar , William Fedus , William Saunders , William Zhang , Wout Vossen , Xiang Ren , Xiaoyu Tong , Xinran Zhao , Xinyi Wu , Xudong Shen , Yadollah Yaghoobzadeh , Yair Lakretz , Yangqiu Song , Yasaman Bahri , Yejin Choi , Yichi Yang , Yiding Hao , Yifu Chen , Yonatan Belinkov , Yu Hou , Yufang Hou , Yuntao Bai , Zachary Seid , Zhuoye Zhao , Zijian Wang , Zijie J. Wang , Zirui Wang , Ziyi Wu

In this paper, we propose the use of the popular word-based board game Codenames as a suitable benchmark for evaluating the reasoning capabilities of Large Language Models (LLMs). Codenames presents a highly interesting challenge for…

人工智能 · 计算机科学 2025-04-23 Matthew Stephenson , Matthew Sidji , Benoît Ronval

Benchmarks are the primary tool for assessing progress in artificial intelligence (AI), yet current practice evaluates models on isolated test suites and provides little guidance for reasoning about generality or autonomous…

人工智能 · 计算机科学 2025-12-05 Przemyslaw Chojecki

Constructing benchmarks that test the abilities of modern natural language understanding models is difficult - pre-trained language models exploit artifacts in benchmarks to achieve human parity, but still fail on adversarial examples and…

计算与语言 · 计算机科学 2022-01-17 Alon Talmor , Ori Yoran , Ronan Le Bras , Chandra Bhagavatula , Yoav Goldberg , Yejin Choi , Jonathan Berant

Despite rapid technological progress, effective human-machine cooperation remains a significant challenge. Humans tend to cooperate less with machines than with fellow humans, a phenomenon known as the machine penalty. Here, we show that…

人机交互 · 计算机科学 2025-05-29 Zhen Wang , Ruiqi Song , Chen Shen , Shiya Yin , Zhao Song , Balaraju Battu , Lei Shi , Danyang Jia , Talal Rahwan , Shuyue Hu

From the beginning if the history of AI, there has been interest in games as a platform of research. As the field developed, human-level competence in complex games became a target researchers worked to reach. Only relatively recently has…

人工智能 · 计算机科学 2019-08-30 Rodrigo Canaan , Christoph Salge , Julian Togelius , Andy Nealen

From the early days of computing, games have been important testbeds for studying how well machines can do sophisticated decision making. In recent years, machine learning has made dramatic advances with artificial agents reaching…

Multi-agent simulations are versatile tools for exploring interactions among natural and artificial agents, but their development typically demands domain expertise and manual effort. This work introduces the Generative Agents for…

人工智能 · 计算机科学 2025-05-30 Agnieszka Mensfelt , Kostas Stathis , Vince Trencsenyi

As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of…

人工智能 · 计算机科学 2026-05-12 Miles Q. Li , Benjamin C. M. Fung , Martin Weiss , Pulei Xiong , Khalil Al-Hussaeni , Claude Fachkha

Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a score without performing the intended task, emerges…

人工智能 · 计算机科学 2026-05-14 Hao Wang , Hanchen Li , Qiuyang Mang , Alvin Cheung , Koushik Sen , Dawn Song

Given the remarkable performance of Large Language Models (LLMs), an important question arises: Can LLMs conduct human-like scientific research and discover new knowledge, and act as an AI scientist? Scientific discovery is an iterative…

机器学习 · 计算机科学 2025-02-24 Tingting Chen , Srinivas Anumasa , Beibei Lin , Vedant Shah , Anirudh Goyal , Dianbo Liu
‹ 上一页 1 2 3 10 下一页 ›