中文
相关论文

相关论文: The MacGyver Test - A Framework for Evaluating Mac…

200 篇论文

Recent advances in large multimodal models have enabled new opportunities in embodied AI, particularly in robotic manipulation. These models have shown strong potential in generalization and reasoning, but achieving reliable and responsible…

机器人学 · 计算机科学 2025-12-05 Lei Zhang , Ju Dong , Kaixin Bai , Minheng Ni , Zoltan-Csaba Marton , Zhaopeng Chen , Jianwei Zhang

How effectively can LLM-based AI assistants utilize their memory (context) to perform various tasks? Traditional data benchmarks, which are often manually crafted, suffer from several limitations: they are static, susceptible to…

计算与语言 · 计算机科学 2025-06-10 Menglin Xia , Victor Ruehle , Saravan Rajmohan , Reza Shokri

There is a growing interest in the area of machine learning and creativity. This survey presents an overview of the history and the state of the art of computational creativity theories, key machine learning techniques (including generative…

机器学习 · 计算机科学 2025-02-14 Giorgio Franceschelli , Mirco Musolesi

The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emphasize skill-level execution and provide…

机器人学 · 计算机科学 2026-05-29 Chunru Lin , Hongxin Zhang , Fenghao Yu , Zhehuan Chen , Thomas L. Griffiths , Yejin Choi , David Held , Chuang Gan

Automatic question generation (AQG) for mathematics education remains an elusive goal for Intelligent Tutoring Systems and educators. While pre-trained transformer-based language models have significantly advanced natural language…

多智能体系统 · 计算机科学 2025-11-07 Kia Karbasi , Kevin Hong , Mohammad Amin Samadi , Gregory Pottie

In a world of daily emerging scientific inquisition and discovery, the prolific launch of machine learning across industries comes to little surprise for those familiar with the potential of ML. Neither so should the congruent expansion of…

人工智能 · 计算机科学 2021-12-13 Brianna Richardson , Juan E. Gilbert

Multi-objective evaluation is a necessary aspect when managing complex systems, as the intrinsic complexity of a system is generally closely linked to the potential number of optimization objectives. However, an evaluation makes no sense…

物理与社会 · 物理学 2016-08-03 Juste Raimbault

Artificial intelligence (AI) systems are increasingly adopted as tool-using agents that can plan, observe their environment, and take actions over extended time periods. This evolution challenges current evaluation practices where the AI…

密码学与安全 · 计算机科学 2026-03-17 Simone Aonzo , Merve Sahin , Aurélien Francillon , Daniele Perito

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning capabilities, are supported with model performance on narrow…

Machine learning algorithms have become ubiquitous in a number of applications (e.g. image classification). However, due to the insufficient measurement of traditional metrics (e.g. the coarse-grained Accuracy of each classifier),…

机器学习 · 计算机科学 2023-07-17 Qi Liu , Zheng Gong , Zhenya Huang , Chuanren Liu , Hengshu Zhu , Zhi Li , Enhong Chen , Hui Xiong

Generative AI is rapidly transforming how organizations create value and evaluate talent. While large language models enhance baseline output quality, they simultaneously introduce ambiguity in assessing human creativity, as observable…

人机交互 · 计算机科学 2026-04-23 Yigal Rosen , Ilia Rushkin

Performance of NLP systems is typically evaluated by collecting a large-scale dataset by means of crowd-sourcing to train a data-driven model and evaluate it on a held-out portion of the data. This approach has been shown to suffer from…

计算与语言 · 计算机科学 2024-08-12 Viktor Schlegel , Goran Nenadic , Riza Batista-Navarro

A core part of human intelligence is the ability to work flexibly with others to achieve goals. The incorporation of artificial agents into human spaces is making increasing demands on artificial intelligence (AI) to demonstrate and…

人机交互 · 计算机科学 2026-03-30 William J. Bingley , S. Alexander Haslam , Janet Wiles

Traditionally, the way one evaluates the performance of an Artificial Intelligence (AI) system is via a comparison to human performance in specific tasks, treating humans as a reference for high-level cognition. However, these comparisons…

人工智能 · 计算机科学 2019-11-25 Camilo M. Signorelli , Xerxes D. Arsiwalla

Reproducibility crises across sciences highlight the limitations of the paper-centric review system in assessing the rigor and reproducibility of research. AI agents that autonomously design and generate large volumes of research outputs…

计算机与社会 · 计算机科学 2026-02-24 Xiaoyan Bai , Alexander Baumgartner , Haojia Sun , Ari Holtzman , Chenhao Tan

Artificial Intelligence (AI) technology epitomizes the complex challenges posed by human-made artifacts, particularly those widely integrated into society and exerting significant influence, highlighting potential benefits and their…

人工智能 · 计算机科学 2025-10-06 Michael Papademas , Xenia Ziouvelou , Antonis Troumpoukis , Vangelis Karkaletsis

Artificial general intelligence (AGI) refers to research aimed at tackling the full problem of artificial intelligence, that is, create truly intelligent agents. This sets it apart from most AI research which aims at solving relatively…

人工智能 · 计算机科学 2011-09-08 Tom Schaul , Julian Togelius , Jürgen Schmidhuber

The rapid development of artificial intelligence has brought the artificial intelligence threat theory as well as the problem about how to evaluate the intelligence level of intelligent products. Both need to find a quantitative method to…

人工智能 · 计算机科学 2017-12-19 Feng Liu , Yong Shi , Ying Liu

Current learning machines have successfully solved hard application problems, reaching high accuracy and displaying seemingly "intelligent" behavior. Here we apply recent techniques for explaining decisions of state-of-the-art learning…

Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, and first-hand user…