中文
相关论文

相关论文: Hi-Phy: A Benchmark for Hierarchical Physical Reas…

200 篇论文

Scientific reasoning through Large Language Models in heliophysics involves more than just recalling facts: it requires incorporating physical assumptions, maintaining consistent units, and providing clear scientific formats through…

人工智能 · 计算机科学 2026-02-10 Kevin Lee , Russell Spiewak , James Walsh

We present HippoCamp, a new benchmark designed to evaluate agents' capabilities on multimodal file management. Unlike existing agent benchmarks that focus on tasks like web interaction, tool use, or software automation in generic settings,…

High-level reasoning can be defined as the capability to generalize over knowledge acquired via experience, and to exhibit robust behavior in novel situations. Such form of reasoning is a basic skill in humans, who seamlessly use it in a…

人工智能 · 计算机科学 2023-11-15 Alessandro Oltramari

Cognitive Psychology and related disciplines have identified several critical mechanisms that enable intelligent biological agents to learn to solve complex problems. There exists pressing evidence that the cognitive mechanisms that enable…

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have shown great promise in…

计算机视觉与模式识别 · 计算机科学 2025-01-30 Wei Chow , Jiageng Mao , Boyi Li , Daniel Seita , Vitor Guizilini , Yue Wang

Over the past few years the Angry Birds AI competition has been held in an attempt to develop intelligent agents that can successfully and efficiently solve levels for the video game Angry Birds. Many different agents and strategies have…

人工智能 · 计算机科学 2019-05-31 Tommy Liu , Jochen Renz , Peng Zhang , Matthew Stephenson

Humans have an inherent ability to learn novel concepts from only a few samples and generalize these concepts to different situations. Even though today's machine learning models excel with a plethora of training data on standard…

人工智能 · 计算机科学 2021-01-06 Weili Nie , Zhiding Yu , Lei Mao , Ankit B. Patel , Yuke Zhu , Animashree Anandkumar

Evaluating the scientific discovery capabilities of large language model based agents, particularly how they cope with varying environmental complexity and utilize prior knowledge, requires specialized benchmarks currently lacking in the…

机器学习 · 计算机科学 2025-10-28 Yimeng Chen , Piotr Piȩkos , Mateusz Ostaszewski , Firas Laakom , Jürgen Schmidhuber

As Large Language Models (LLMs) gain agentic abilities, they will have to navigate complex multi-agent scenarios, interacting with human users and other agents in cooperative and competitive settings. This will require new reasoning skills,…

人工智能 · 计算机科学 2025-06-26 Andrei Lupu , Timon Willi , Jakob Foerster

Detecting and responding to novel situations in open-world environments is a key capability of human cognition and is a persistent problem for AI systems. In an open-world, novelties can appear in many different forms and may be easy or…

人工智能 · 计算机科学 2023-06-27 Vimukthini Pinto , Cheng Xue , Chathura Nagoda Gamage , Matthew Stephenson , Jochen Renz

While existing benchmarks probe the reasoning abilities of large language models (LLMs) across diverse domains, they predominantly assess passive reasoning, providing models with all the information needed to reach a solution. By contrast,…

机器学习 · 计算机科学 2025-06-11 Zhanke Zhou , Xiao Feng , Zhaocheng Zhu , Jiangchao Yao , Sanmi Koyejo , Bo Han

This paper presents a novel approach to analyze human decision-making that involves comparing the behavior of professional chess players relative to a computational benchmark of cognitively bounded rationality. This benchmark is constructed…

综合经济学 · 经济学 2020-12-03 Dainis Zegners , Uwe Sunde , Anthony Strittmatter

Visual reasoning is central to human cognition, enabling individuals to interpret and abstractly understand their environment. Although recent Multimodal Large Language Models (MLLMs) have demonstrated impressive performance across language…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Jing Bi , Junjia Guo , Susan Liang , Guangyu Sun , Luchuan Song , Yunlong Tang , Jinxi He , Jiarui Wu , Ali Vosoughi , Chen Chen , Chenliang Xu

While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely…

Human-robot interactions (HRI) can be modeled as dynamic or differential games with incomplete information, where each agent holds private reward parameters. Due to the open challenge in finding perfect Bayesian equilibria of such games,…

机器人学 · 计算机科学 2020-11-05 Yi Chen , Lei Zhang , Tanner Merry , Sunny Amatya , Wenlong Zhang , Yi Ren

LLM-driven multi-agent-based simulations have been gaining traction with applications in game-theoretic and social simulations. While most implementations seek to exploit or evaluate LLM-agentic reasoning, they often do so with a weak…

人工智能 · 计算机科学 2026-02-17 Vince Trencsenyi , Agnieszka Mensfelt , Kostas Stathis

We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline,…

To achieve human-like common sense about everyday life, machine learning systems must understand and reason about the goals, preferences, and actions of other agents in the environment. By the end of their first year of life, human infants…

人工智能 · 计算机科学 2022-02-16 Kanishk Gandhi , Gala Stojnic , Brenden M. Lake , Moira R. Dillon

Large language models (LLMs) have rapidly advanced and are increasingly capable of tackling complex scientific problems, including those in physics. Despite this progress, current LLMs often fail to emulate the concise, principle-based…

机器学习 · 计算机科学 2025-06-02 Yinggan Xu , Yue Liu , Zhiqiang Gao , Changnan Peng , Di Luo