English
Related papers

Related papers: JARVIS-1: Open-World Multi-task Agents with Memory…

200 papers

Many everyday tasks rely on external tutorials such as manuals and videos, requiring users to constantly switch between reading instructions and performing actions, which disrupts workflow and increases cognitive load. Augmented reality…

Human-Computer Interaction · Computer Science 2026-05-19 Yusi Sun , Ying Jiang , Jiayin Lu , Yin yang , Yong-Hong Kuo , Chenfanfu Jiang

Due to the dynamic and unpredictable open-world setting, navigating complex environments in Minecraft poses significant challenges for multi-agent systems. Agents must interact with the environment and coordinate their actions with other…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Zhonghan Zhao , Kewei Chen , Dongxu Guo , Wenhao Chai , Tian Ye , Yanting Zhang , Gaoang Wang

We present Game-TARS, a generalist game agent trained with a unified, scalable action space anchored to human-aligned native keyboard-mouse inputs. Unlike API- or GUI-based approaches, this paradigm enables large-scale continual…

It is a long-lasting goal to design a generalist-embodied agent that can follow diverse instructions in human-like ways. However, existing approaches often fail to steadily follow instructions due to difficulties in understanding abstract…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Enshen Zhou , Yiran Qin , Zhenfei Yin , Yuzhou Huang , Ruimao Zhang , Lu Sheng , Yu Qiao , Jing Shao

Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A central challenge is therefore to transform past executions into knowledge that can shape future…

Artificial Intelligence · Computer Science 2026-05-12 Zhengwei Xie , Zhisheng Chen , Ziyan Weng , Jinhan Li , Chenglong Li , Zikai Xiao , Jingwei Song , Jinhao Jing , Vireo Zhang , Kun Wang

Understanding the mechanisms behind decisions taken by large foundation models in sequential decision making tasks is critical to ensuring that such systems operate transparently and safely. In this work, we perform exploratory analysis on…

Real-world visualization tasks involve complex, multi-modal requirements that extend beyond simple text-to-chart generation, requiring reference images, code examples, and iterative refinement. Current systems exhibit fundamental…

Computation and Language · Computer Science 2026-01-27 Jinwei Lu , Yuanfeng Song , Chen Zhang , Raymond Chi-Wing Wong

The Joint Automated Repository for Various Integrated Simulations (JARVIS) is an integrated infrastructure to accelerate materials discovery and design using density functional theory (DFT), classical force-fields (FF), and machine learning…

Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be…

Machine Learning · Computer Science 2026-03-10 Jiajian Li , Qi Wang , Yunbo Wang , Xin Jin , Yang Li , Wenjun Zeng , Xiaokang Yang

Collaboration is a cornerstone of society. In the real world, human teammates make use of multi-sensory data to tackle challenging tasks in ever-changing environments. It is essential for embodied agents collaborating in visually-rich…

Artificial Intelligence · Computer Science 2024-12-09 Qian Long , Zhi Li , Ran Gong , Ying Nian Wu , Demetri Terzopoulos , Xiaofeng Gao

Developing generalist agents capable of solving open-ended tasks in visually rich, dynamic environments remains a core pursuit of embodied AI. While Minecraft has emerged as a compelling benchmark, existing agents often suffer from…

Artificial Intelligence · Computer Science 2026-02-11 Zaijing Li , Yuquan Xie , Rui Shao , Gongwei Chen , Weili Guan , Dongmei Jiang , Yaowei Wang , Liqiang Nie

World models learn general knowledge from videos and simulate experience for training behaviors in imagination, offering a path towards intelligent agents. However, previous world models have been unable to accurately predict object…

Artificial Intelligence · Computer Science 2025-09-30 Danijar Hafner , Wilson Yan , Timothy Lillicrap

Despite recent advances, long-sequence video generation frameworks still suffer from significant limitations: poor assistive capability, suboptimal visual quality, and limited expressiveness. To mitigate these limitations, we propose MAViS,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Qian Wang , Ziqi Huang , Ruoxi Jia , Paul Debevec , Ning Yu

We present MineNPC-Task, a user-authored benchmark and evaluation harness for testing memory-aware, mixed-initiative LLM agents in open-world Minecraft. Rather than relying on synthetic prompts, tasks are elicited through formative and…

Artificial Intelligence · Computer Science 2026-01-12 Tamil Sudaravan Mohan Doss , Michael Xu , Sudha Rao , Andrew D. Wilson , Balasaravanan Thoravi Kumaravel

Object Goal Navigation (ObjectNav) refers to an agent navigating to an object in an unseen environment, which is an ability often required in the accomplishment of complex tasks. While existing methods demonstrate proficiency in isolated…

Robotics · Computer Science 2026-04-15 Jiahua Pei , Yi Liu , Guoping Pan , Yuanhao Jiang , Houde Liu , Xueqian Wang

Designing a generalist scientific agent capable of performing tasks in laboratory settings to assist researchers has become a key goal in recent Artificial Intelligence (AI) research. Unlike everyday tasks, scientific tasks are inherently…

Computation and Language · Computer Science 2026-03-20 Minh Pham Dinh , Munira Syed , Michael G Yankoski , Trenton W. Ford

Agent-based editing models have substantially advanced interactive experiences, processing quality, and creative flexibility. However, two critical challenges persist: (1) instruction hallucination, text-only chain-of-thought (CoT)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yunlong Lin , Linqing Wang , Kunjie Lin , Zixu Lin , Kaixiong Gong , Wenbo Li , Bin Lin , Zhenxi Li , Shiyi Zhang , Yuyang Peng , Wenxun Dai , Xinghao Ding , Chunyu Wang , Qinglin Lu

The choice of action spaces is a critical yet unresolved challenge in developing capable, end-to-end trainable agents. This paper first presents a large-scale, systematic comparison of prominent abstracted action spaces and tokenizers for…

Artificial Intelligence · Computer Science 2025-09-18 Zihao Wang , Muyao Li , Kaichen He , Xiangyu Wang , Zhancun Mu , Anji Liu , Yitao Liang

Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations.…

Artificial Intelligence · Computer Science 2025-06-04 Xinyue Zheng , Haowei Lin , Kaichen He , Zihao Wang , Zilong Zheng , Yitao Liang

Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse feedback, and irreversible mistakes. Video games offer an ideal testbed with rich visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Mingyu Ouyang , Siyuan Hu , Kevin Qinghong Lin , Hwee Tou Ng , Mike Zheng Shou