中文
相关论文

相关论文: iWorld-Bench: A Benchmark for Interactive World Mo…

200 篇论文

World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face practical limitations in GUI settings, where predicting complex…

人工智能 · 计算机科学 2025-12-17 Shufan Li , Konstantinos Kallidromitis , Akash Gokul , Yusuke Kato , Kazuki Kozuka , Aditya Grover

Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static prompt evaluation misses failures that only appear when agents…

人工智能 · 计算机科学 2026-05-19 Yuxiang Lai , Peng Xia , Haonian Ji , Kaiwen Xiong , Kaide Zeng , Jiaqi Liu , Fang Wu , Jike Zhong , Zeyu Zheng , Cihang Xie , Huaxiu Yao

Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on predicting isolated mental state labels, lacking structured annotations of complex…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Xiaoyu Yuan , Niklas Heikkala , Tiina Törmänen , Hanna Järvenoja , Guoying Zhao , Haoyu Chen

The pursuit of artificial general intelligence necessitates robust methods for evaluating the cognitive capabilities of models beyond narrow task performance. Here, we introduce a psychometric framework to assess the cognitive profiles of…

人工智能 · 计算机科学 2026-05-11 Isaac Galatzer-Levy , Daniel McDuff , Xin Liu , Jed McGiffin

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

While GUI agents have shown impressive capabilities in common computer-use tasks such as OSWorld, current benchmarks mainly focus on isolated and single-application tasks. This overlooks a critical real-world requirement of coordinating…

人工智能 · 计算机科学 2026-05-01 Jinchao Li , Yunxin Li , Chenrui Zhao , Zhenran Xu , Baotian Hu , Min Zhang

Benefiting from the powerful capabilities of large language models (LLMs), agents based on LLMs have shown the potential to address domain-specific tasks and emulate human behaviors. However, the content generated by these agents remains…

人工智能 · 计算机科学 2023-08-29 Shi Jinxin , Zhao Jiabao , Wang Yilei , Wu Xingjiao , Li Jiawen , He Liang

World models - generative models that simulate environment dynamics conditioned on past observations and actions - are gaining prominence in planning, simulation, and embodied AI. However, evaluating their rollouts remains a fundamental…

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a…

机器人学 · 计算机科学 2026-04-14 Liaoyuan Fan , Zetian Xu , Chen Cao , Wenyao Zhang , Mingqi Yuan , Jiayu Chen

To coordinate actions with an interaction partner requires a constant exchange of sensorimotor signals. Humans acquire these skills in infancy and early childhood mostly by imitation learning and active engagement with a skilled partner.…

机器学习 · 计算机科学 2019-10-15 Judith Bütepage , Ali Ghadirzadeh , Özge Öztimur Karadag , Mårten Björkman , Danica Kragic

A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Chuhan Li , Ruilin Han , Joy Hsu , Yongyuan Liang , Rajiv Dhawan , Jiajun Wu , Ming-Hsuan Yang , Xin Eric Wang

The development of generative artificial intelligence for human motion generation has expanded rapidly, necessitating a unified evaluation framework. This paper presents a detailed review of eight evaluation metrics for human motion…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Ali Ismail-Fawaz , Maxime Devanne , Stefano Berretti , Jonathan Weber , Germain Forestier

The rapid advancement of AI-generated video models has created a pressing need for robust and interpretable evaluation frameworks. Existing metrics are limited to producing numerical scores without explanatory comments, resulting in low…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Xiao Liu , Jiawei Zhang

Autonomous agents that address day-to-day digital tasks (e.g., ordering groceries for a household), must not only operate multiple apps (e.g., notes, messaging, shopping app) via APIs, but also generate rich code with complex control flow…

We present ScienceWorld, a benchmark to test agents' scientific reasoning abilities in a new interactive text environment at the level of a standard elementary school science curriculum. Despite the transformer-based progress seen in…

计算与语言 · 计算机科学 2022-11-15 Ruoyao Wang , Peter Jansen , Marc-Alexandre Côté , Prithviraj Ammanabrolu

Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set by existing data. To solve novel problems, agents should acquire skills for exploring and…

Goal changes are a defining feature of real world multi-turn interactions, yet current agent benchmarks primarily evaluate static objectives or one-shot tool use. We introduce AgentChangeBench, a benchmark explicitly designed to measure how…

人工智能 · 计算机科学 2025-10-22 Manik Rana , Calissa Man , Anotida Expected Msiiwa , Jeffrey Paine , Kevin Zhu , Sunishchal Dev , Vasu Sharma , Ahan M R

World models have been emerging as critical components for assessing the consequences of actions generated by interactive agents in online planning and offline evaluation. In text-based environments, world models are typically evaluated and…

Human interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental to a wide range of real-world applications. In this paper,…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Heng Yu , Juze Zhang , Changan Chen , Tiange Xiang , Yusu Fang , Juan Carlos Niebles , Ehsan Adeli

We provide a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives. Interactive narratives -- or text-adventure games -- are partially observable environments…

计算与语言 · 计算机科学 2021-06-18 Prithviraj Ammanabrolu , Mark O. Riedl