中文
相关论文

相关论文: The MacGyver Test - A Framework for Evaluating Mac…

200 篇论文

To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons…

人工智能 · 计算机科学 2019-11-26 François Chollet

Evaluating learned robot control policies to determine their physical task-level capabilities costs experimenter time and effort. The growing number of policies and tasks exacerbates this issue. It is impractical to test every policy on…

机器人学 · 计算机科学 2025-02-17 Abrar Anwar , Rohan Gupta , Zain Merchant , Sayan Ghosh , Willie Neiswanger , Jesse Thomason

A widely accepted definition of intelligence in the context of Artificial Intelligence (AI) still eludes us. Due to our exceedingly rapid development of AI paradigms, architectures, and tools, the prospect of naturally arising AI…

人工智能 · 计算机科学 2023-07-10 Ira Wolfson

Much work has been done in understanding human creativity and defining measures to evaluate creativity. This is necessary mainly for the reason of having an objective and automatic way of quantifying creative artifacts. In this work, we…

机器学习 · 计算机科学 2017-07-19 Disha Shrivastava , Saneem Ahmed CG , Anirban Laha , Karthik Sankaranarayanan

Intelligence is a crucial trait for species to find solutions within a limited number of trial-and-error attempts. Building on this idea, we introduce Survival Game as a framework to evaluate intelligence based on the number of failed…

人工智能 · 计算机科学 2025-03-06 Jingtao Zhan , Jiahao Zhao , Jiayu Li , Yiqun Liu , Bo Zhang , Qingyao Ai , Jiaxin Mao , Hongning Wang , Min Zhang , Shaoping Ma

As robotic teammates become more common in society, people will assess the robots' roles in their interactions along many dimensions. One such dimension is effectiveness: people will ask whether their robotic partners are trustworthy and…

人工智能 · 计算机科学 2020-10-13 Richard G. Freedman , Steven J. Levine , Brian C. Williams , Shlomo Zilberstein

With the release of ChatGPT and other large language models (LLMs) the discussion about the intelligence, possibilities, and risks, of current and future models have seen large attention. This discussion included much debated scenarios…

人工智能 · 计算机科学 2024-07-31 Nils Körber , Silvan Wehrli , Christopher Irrgang

This paper introduces the Shepherd Test, a new conceptual test for assessing the moral and relational dimensions of superintelligent artificial agents. The test is inspired by human interactions with animals, where ethical considerations…

人工智能 · 计算机科学 2025-09-30 Djallel Bouneffouf , Matthew Riemer , Kush Varshney

As machine learning (ML) systems increasingly permeate high-stakes settings such as healthcare, transportation, military, and national security, concerns regarding their reliability have emerged. Despite notable progress, the performance of…

机器学习 · 计算机科学 2023-08-01 Anthony Corso , David Karamadian , Romeo Valentin , Mary Cooper , Mykel J. Kochenderfer

Generative AI techniques have opened the path for new generations of machines in diverse domains. These machines have various capabilities for example, they can produce images, generate answers or stories, and write codes based on the…

人工智能 · 计算机科学 2023-07-18 Nitisha Aggarwal , Geetika Jain Saxena , Sanjeev Singh , Amit Pundir

Amid mounting concern about the reliability and credibility of machine learning research, we present a principled framework for making robust and generalizable claims: the multiverse analysis. Our framework builds upon the multiverse…

机器学习 · 计算机科学 2022-10-13 Samuel J. Bell , Onno P. Kampman , Jesse Dodge , Neil D. Lawrence

The Turing test may or may not be a valid test of machine intelligence. But in an age of generative AI, the test describes the positions we humans occupy. Judging whether or not something is human or machine produced is an everyday…

计算机与社会 · 计算机科学 2026-01-21 Samuel Gerald Collins

Reproducibility is one of the core dimensions that concur to deliver Trustworthy Artificial Intelligence. Broadly speaking, reproducibility can be defined as the possibility to reproduce the same or a similar experiment or method, thereby…

人工智能 · 计算机科学 2023-02-27 Riccardo Albertoni , Sara Colantonio , Piotr Skrzypczyński , Jerzy Stefanowski

Measuring machine creativity is one of the most fascinating challenges in Artificial Intelligence. This paper explores the possibility of using generative learning techniques for automatic assessment of creativity. The proposed solution…

机器学习 · 计算机科学 2024-05-17 Giorgio Franceschelli , Mirco Musolesi

Social intelligence in natural and artificial systems is usually measured by the evaluation of associated traits or tasks that are deemed to represent some facets of social behaviour. The amalgamation of these traits is then used to…

多智能体系统 · 计算机科学 2014-08-28 Javier Insa-Cabrera , José Hernández-Orallo

Human evaluators provide necessary contributions in evaluating large language models. In the context of Machine Translation (MT) systems for low-resource languages (LRLs), this is made even more apparent since popular automated metrics tend…

计算与语言 · 计算机科学 2025-06-16 Carlos Rafael Catalan

Reliable and robust evaluation methods are a necessary first step towards developing machine learning models that are themselves robust and reliable. Unfortunately, current evaluation protocols typically used to assess classifiers fail to…

机器学习 · 计算机科学 2025-05-26 Michael W. Spratling

Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence remains elusive- lacking a stable definition and failing to…

人工智能 · 计算机科学 2025-11-18 Ruchira Dhar , Ninell Oldenburg , Anders Soegaard

We develop a taxonomical framework for classifying challenges to the possibility of consciousness in digital artificial intelligence systems. This framework allows us to identify the level of granularity at which a given challenge is…

人工智能 · 计算机科学 2025-11-21 Andres Campero , Derek Shiller , Jaan Aru , Jonathan Simon

Due to the diffusion of IoT, modern software systems are often thought to control and coordinate smart devices in order to manage assets and resources, and to guarantee efficient behaviours. For this class of systems, which interact…

计算机科学中的逻辑 · 计算机科学 2024-02-14 Valentina Castiglioni , Michele Loreti , Simone Tini