English
Related papers

Related papers: Evaluating Intelligence via Trial and Error

200 papers

While LLMs have shown impressive capabilities in solving math or coding problems, the ability to make scientific discoveries remains a distinct challenge. This paper proposes a "Turing test for an AI scientist" to assess whether an AI agent…

Artificial Intelligence · Computer Science 2024-05-24 Xiaoxin Yin

Current and foreseeable GenAI models are not capable of achieving artificial general intelligence because they are burdened with anthropogenic debt. They depend heavily on human input to provide well-structured problems, architecture, and…

Neurons and Cognition · Quantitative Biology 2025-02-13 Herbert Roitblat

This paper studies autonomous generative AI agents in multi-echelon supply chains using the MIT Beer Game. We identify four inference-time levers that shape performance: model selection, policies and guardrails, centralized data sharing,…

Artificial Intelligence · Computer Science 2026-05-27 Carol Xuan Long , David Simchi-Levi , Feng Zhu , Huangyuan Su , Andre P. Calmon , Flavio P. Calmon

Effective financial reasoning demands not only textual understanding but also the ability to interpret complex visual data such as charts, tables, and trend graphs. This paper introduces a new benchmark designed to evaluate how well AI…

Artificial Intelligence · Computer Science 2025-06-10 Shuangyan Deng , Haizhou Peng , Jiachen Xu , Chunhou Liu , Ciprian Doru Giurcuaneanu , Jiamou Liu

We develop a taxonomical framework for classifying challenges to the possibility of consciousness in digital artificial intelligence systems. This framework allows us to identify the level of granularity at which a given challenge is…

Artificial Intelligence · Computer Science 2025-11-21 Andres Campero , Derek Shiller , Jaan Aru , Jonathan Simon

A framework is proposed that seeks to identify and establish a set of robust autonomous levels articulating the realm of Artificial Intelligence and Legal Reasoning (AILR). Doing so provides a sound and parsimonious basis for being able to…

Computers and Society · Computer Science 2020-08-18 Lance Eliot

Artificial Intelligence (AI) technology epitomizes the complex challenges posed by human-made artifacts, particularly those widely integrated into society and exerting significant influence, highlighting potential benefits and their…

Artificial Intelligence · Computer Science 2025-10-06 Michael Papademas , Xenia Ziouvelou , Antonis Troumpoukis , Vangelis Karkaletsis

Humans frequently make decisions with the aid of artificially intelligent (AI) systems. A common pattern is for the AI to recommend an action to the human who retains control over the final decision. Researchers have identified ensuring…

Artificial Intelligence · Computer Science 2025-09-26 Ziyang Guo , Yifan Wu , Jason Hartline , Jessica Hullman

The capabilities of artificial intelligence (AI) lie along a jagged frontier, where AI systems surprisingly fail on tasks that humans find easy and succeed on tasks that humans find hard. To investigate user reactions to this phenomenon, we…

Computers and Society · Computer Science 2026-04-07 Jacy Reese Anthis , Hannah Cha , Solon Barocas , Alexandra Chouldechova , Jake Hofman

Large Language Models (LLMs) are recruited in applications that span from clinical assistance and legal support to question answering and education. Their success in specialized tasks has led to the claim that they possess human-like…

Computation and Language · Computer Science 2024-07-10 Vittoria Dentella , Fritz Guenther , Elliot Murphy , Gary Marcus , Evelina Leivada

Recent advancements in algorithms for sequential decision-making under imperfect information have shown remarkable success in large games such as limit- and no-limit poker. These algorithms traditionally formalize the games using the…

Computer Science and Game Theory · Computer Science 2023-12-07 Vojtěch Kovařík , David Milec , Michal Šustr , Dominik Seitz , Viliam Lisý

As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improving, it will become increasingly hard for humans to generate…

Regression or classification? This is perhaps the most basic question faced when tackling a new supervised learning problem. We present an Evolutionary Deep Learning (EDL) algorithm that automatically solves this by identifying the question…

Neural and Evolutionary Computing · Computer Science 2017-07-05 Emmanuel Dufourq , Bruce A. Bassett

Self-assessment rules play an essential role in safe and effective real-world robotic applications, which verify the feasibility of the selected action before actual execution. But how to utilize the self-assessment results to re-choose…

Robotics · Computer Science 2023-02-28 Kechun Xu , Runjian Chen , Shuqi Zhao , Zizhang Li , Hongxiang Yu , Ci Chen , Yue Wang , Rong Xiong

Capture-the-Flag (CTF) competitions are increasingly becoming a testbed for evaluating AI capabilities at solving security tasks, due to the controlled environments and objective success criteria. Existing evaluations have focused on how…

Cryptography and Security · Computer Science 2026-02-25 Tingxuan Tang , Nicolas Janis , Kalyn Asher Montague , Kevin Eykholt , Dhilung Kirat , Youngja Park , Jiyong Jang , Adwait Nadkarni , Yue Xiao

Evaluating AI agents within complex, interactive environments that mirror real-world challenges is critical for understanding their practical capabilities. While existing agent benchmarks effectively assess skills like tool use or…

Artificial Intelligence · Computer Science 2025-08-15 Long Phan , Mantas Mazeika , Andy Zou , Dan Hendrycks

As evaluation designs of large language models may shape our trajectory toward artificial general intelligence, comprehensive and forward-looking assessment is essential. Existing benchmarks primarily assess static knowledge, while…

Computation and Language · Computer Science 2025-08-07 Jiayin Wang , Zhiquang Guo , Weizhi Ma , Min Zhang

Rapid individual cognitive phenotyping holds the potential to revolutionize domains as wide-ranging as personalized learning, employment practices, and precision psychiatry. Going beyond limitations imposed by traditional lab-based…

Strong reciprocity is a fundamental human characteristic associated with our extraordinary sociality and cooperation. Laboratory experiments on social dilemma games and many field studies have quantified well-defined levels of cooperation…

Physics and Society · Physics 2007-11-21 D. Darcet , D. Sornette

Artificial intelligence (AI) is poised to revolutionize military combat systems, but ensuring these AI-enabled capabilities are truly mission-ready presents new challenges. We argue that current technology readiness assessments fail to…

Software Engineering · Computer Science 2025-06-16 S. Tucker Browne , Mark M. Bailey
‹ Prev 1 8 9 10 Next ›