English
Related papers

Related papers: Evaluating Intelligence via Trial and Error

200 papers

Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience across task boundaries. This paper formalizes the…

Artificial Intelligence · Computer Science 2026-05-26 Sihang Jiang , Lipeng Ma , Zhonghua Hong , Keyi Wang , Zhiyu Lu , Tengfei Wang , Shisong Chen , Jinghao Zhang , Tianjun Pan , Weijia Li , Jiaqing Liang , Yanghua Xiao

Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first…

Computers and Society · Computer Science 2025-04-22 Heidy Khlaaf , Sarah Myers West

In the future, powerful AI systems may be deployed in high-stakes settings, where a single failure could be catastrophic. One technique for improving AI safety in high-stakes settings is adversarial training, which uses an adversary to…

The recent development of artificial intelligence enables a machine to achieve a human level of intelligence. Problem-solving and decision-making are two mental abilities to measure human intelligence. Many scholars have proposed different…

Artificial Intelligence · Computer Science 2023-02-23 Caesar Wu , Pascal Bouvry

Current AI alignment methodologies rely on human-provided demonstrations or judgments, and the learned capabilities of AI systems would be upper-bounded by human capabilities as a result. This raises a challenging research question: How can…

Machine Learning · Computer Science 2024-12-11 Zhiqing Sun , Longhui Yu , Yikang Shen , Weiyang Liu , Yiming Yang , Sean Welleck , Chuang Gan

Our understanding of intelligence is directed primarily at the human level. This paper attempts to give a more unifying definition that can be applied to the natural world in general and then Artificial Intelligence. The definition would be…

Artificial Intelligence · Computer Science 2026-05-27 Kieran Greer

Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean that users must critically evaluate AI outputs rather than accept them at face value. The present…

Artificial Intelligence · Computer Science 2026-05-27 Gabriel R. Lau , Wei Yan Low , Louis Tay , Ysabel Guevarra , Dragan Gašević , Andree Hartanto

Recently, there have been several high-profile achievements of agents learning to play games against humans and beat them. In this paper, we study the problem of training intelligent agents in service of game development. Unlike the agents…

Games often incorporate random elements in the form of dice or shuffled card decks. This randomness is a key contributor to the player experience and the variety of game situations encountered. There is a tension between a level of…

Artificial Intelligence · Computer Science 2025-03-05 James Goodman , Diego Perez-Liebana , Simon Lucas

The proliferation of generative AI tools has rendered traditional modular assessments in computing and data-centric education increasingly ineffective, creating a disconnect between academic evaluation and authentic skill measurement. This…

Computers and Society · Computer Science 2026-01-22 Kaihua Ding

Facing the current debate on whether Large Language Models (LLMs) attain near-human intelligence levels (Mitchell & Krakauer, 2023; Bubeck et al., 2023; Kosinski, 2023; Shiffrin & Mitchell, 2023; Ullman, 2023), the current study introduces…

Artificial Intelligence · Computer Science 2024-05-21 Junqi Wang , Chunhui Zhang , Jiapeng Li , Yuxi Ma , Lixing Niu , Jiaheng Han , Yujia Peng , Yixin Zhu , Lifeng Fan

Survival analysis is an essential tool for the study of health data. An inherent component of such data is the presence of missing values. In recent years, researchers proposed new learning algorithms for survival tasks based on neural…

Machine Learning · Statistics 2023-03-27 Paul Dufossé , Sébastien Benzekry

The field of artificial intelligence (AI) is witnessing a recent upsurge in research, tools development, and deployment of applications. Multiple software companies are shifting their focus to developing intelligent systems; and many others…

Software Engineering · Computer Science 2021-08-04 Feras A. Batarseh , Rasika Mohod , Abhinav Kumar , Justin Bui

We introduce a new measure of the discrepancy in strategic games between the social welfare in a Nash equilibrium and in a social optimum, that we call selfishness level. It is the smallest fraction of the social welfare that needs to be…

Computer Science and Game Theory · Computer Science 2014-04-04 Krzysztof R. Apt , Guido Schaefer

Situational awareness, the capacity of an AI system to recognize its own nature, understand its training and deployment context, and reason strategically about its circumstances, is widely considered among the most dangerous emergent…

Artificial Intelligence · Computer Science 2026-03-11 Subramanyam Sahoo , Aman Chadha , Vinija Jain , Divya Chaudhary

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose existential risks. This paper reviews the…

Computers and Society · Computer Science 2023-10-30 Rose Hadshar

LLM-based autonomous agents have demonstrated strong capabilities in reasoning, planning, and tool use, yet remain limited when tasks require sustained coordination across roles, tools, and environments. Multi-agent systems address this…

As AI systems become more advanced and widely deployed, there will likely be increasing debate over whether AI systems could have conscious experiences, desires, or other states of potential moral significance. It is important to inform…

Machine Learning · Computer Science 2023-11-16 Ethan Perez , Robert Long

To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and…

Machine Learning · Statistics 2018-10-23 Geoffrey Irving , Paul Christiano , Dario Amodei

The article proposes a universal dual-axis intelligent systems assessment scale. The scale considers the properties of intelligent systems within the environmental context, which develops over time. In contrast to the frequent consideration…

Artificial Intelligence · Computer Science 2023-08-25 Oleg V. Kubryak , Sergey V. Kovalchuk , Nadezhda G. Bagdasaryan
‹ Prev 1 4 5 6 7 8 10 Next ›