English
Related papers

Related papers: AGI-Elo: How Far Are We From Mastering A Task?

200 papers

Current test and evaluation (T&E) methods for assessing machine learning (ML) system performance often rely on incomplete metrics. Testing is additionally often siloed from the other phases of the ML system lifecycle. Research investigating…

Software Engineering · Computer Science 2022-04-11 Violet Turri , Rachel Dzombak , Eric Heim , Nathan VanHoudnos , Jay Palat , Anusha Sinha

In this perspective paper, we first comprehensively review existing evaluations of Large Language Models (LLMs) using both standardized tests and ability-oriented benchmarks. We pinpoint several problems with current evaluation methods that…

Computation and Language · Computer Science 2026-01-30 Yuxi Ma , Chi Zhang , Song-Chun Zhu

This paper argues that the next generation of AI agent (NGENT) should integrate across-domain abilities to advance toward Artificial General Intelligence (AGI). Although current AI agents are effective in specialized tasks such as robotics,…

Artificial Intelligence · Computer Science 2025-05-01 Zhicong Li , Hangyu Mao , Jiangjin Yin , Mingzhe Xing , Zhiwei Xu , Yuanxing Zhang , Yang Xiao

There is a significant lack of unified approaches to building generally intelligent machines. The majority of current artificial intelligence research operates within a very narrow field of focus, frequently without considering the…

Artificial Intelligence · Computer Science 2016-11-03 Marek Rosa , Jan Feyereisl , The GoodAI Collective

Understanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Yuqi Hu , Longguang Wang , Xian Liu , Ling-Hao Chen , Yuwei Guo , Yukai Shi , Ce Liu , Anyi Rao , Zeyu Wang , Hui Xiong

As AI systems advance and integrate into society, well-designed and transparent evaluations are becoming essential tools in AI governance, informing decisions by providing evidence about system capabilities and risks. Yet there remains a…

A traditional approach to assessing emerging intelligence in the theory of intelligent systems is based on the similarity, "imitation" of human-like actions and behaviors, benchmarking the performance of intelligent systems on the scale of…

Neural and Evolutionary Computing · Computer Science 2025-05-28 Serge Dolgikh

Deep learning yields great results across many fields, from speech recognition, image classification, to translation. But for each problem, getting a deep model to work well involves research into the architecture and a long period of…

Machine Learning · Computer Science 2017-06-19 Lukasz Kaiser , Aidan N. Gomez , Noam Shazeer , Ashish Vaswani , Niki Parmar , Llion Jones , Jakob Uszkoreit

Measuring GUI task difficulty is crucial for user behavior analysis and agent capability evaluation. Yet, existing benchmarks typically quantify difficulty based on motor actions (e.g., step counts), overlooking the cognitive demands…

Human-Computer Interaction · Computer Science 2025-11-13 Yiwen Yin , Zhian Hu , Xiaoxi Xu , Chun Yu , Xintong Wu , Wenyu Fan , Yuanchun Shi

This article discusses some trends and concepts in developing new generation of future Artificial General Intelligence (AGI) systems which relate to complex facets and different types of human intelligence, especially social, emotional,…

Artificial Intelligence · Computer Science 2020-12-14 Andrzej Cichocki , Alexander P. Kuleshov

A multitude of explainability methods and associated fidelity performance metrics have been proposed to help better understand how modern AI systems make decisions. However, much of the current work has remained theoretical -- without much…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Julien Colin , Thomas Fel , Remi Cadene , Thomas Serre

Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know that a reported…

Artificial Intelligence · Computer Science 2026-05-19 Nathanael Jo , Ashia Wilson

One approach to achieving artificial general intelligence (AGI) is through the emergence of complex structures and dynamic properties arising from decentralized networks of interacting artificial intelligence (AI) agents. Understanding the…

Artificial Intelligence · Computer Science 2018-06-20 Anton Kolonin , Ben Goertzel , Deborah Duong , Matt Ikle

In coming years or decades, artificial general intelligence (AGI) may surpass human capabilities across many critical domains. We argue that, without substantial effort to prevent it, AGIs could learn to pursue goals that are in conflict…

Artificial Intelligence · Computer Science 2025-05-06 Richard Ngo , Lawrence Chan , Sören Mindermann

The widespread use of AI and ML models in sensitive areas raises significant concerns about fairness. While the research community has introduced various methods for bias mitigation in binary classification tasks, the issue remains…

Machine Learning · Computer Science 2026-03-24 Maryam Boubekraoui , Giordano d'Aloisio , Antinisca Di Marco

What frameworks and architectures are necessary to create a vision system for AGI? In this paper, we propose a formal model that states the task of perception within AGI. We show the role of discriminative and generative models in achieving…

Computer Vision and Pattern Recognition · Computer Science 2018-07-12 Alexey Potapov , Sergey Rodionov , Maxim Peterson , Oleg Shcherbakov , Innokentii Zhdanov , Nikolai Skorobogatko

With the availability of large databases and recent improvements in deep learning methodology, the performance of AI systems is reaching or even exceeding the human level on an increasing number of complex tasks. Impressive examples of this…

Artificial Intelligence · Computer Science 2017-08-29 Wojciech Samek , Thomas Wiegand , Klaus-Robert Müller

The pursuit of Artificial General Intelligence (AGI) is a central goal in language model development, in which consciousness-like processing could serve as a key facilitator. While current language models are not conscious, they exhibit…

Artificial Intelligence · Computer Science 2026-02-02 Hamid Reza Akbari , Mohammad Hossein Sameti , Amir M. Mansourian , Mohammad Hossein Rohban , Hossein Sameti

The original vision of AI was re-articulated in 2002 via the term 'Artificial General Intelligence' or AGI. This vision is to build 'Thinking Machines' - computer systems that can learn, reason, and solve problems similar to the way humans…

Artificial Intelligence · Computer Science 2023-09-20 Peter Voss , Mladjan Jovanovic

Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence remains elusive- lacking a stable definition and failing to…

Artificial Intelligence · Computer Science 2025-11-18 Ruchira Dhar , Ninell Oldenburg , Anders Soegaard
‹ Prev 1 3 4 5 6 7 10 Next ›