中文
相关论文

相关论文: iWorld-Bench: A Benchmark for Interactive World Mo…

200 篇论文

During the evolution of large models, performance evaluation is necessarily performed to assess their capabilities and ensure safety before practical application. However, current model evaluations mainly rely on specific tasks and…

人工智能 · 计算机科学 2024-03-07 Youzhi Qu , Chen Wei , Penghui Du , Wenxin Che , Chi Zhang , Wanli Ouyang , Yatao Bian , Feiyang Xu , Bin Hu , Kai Du , Haiyan Wu , Jia Liu , Quanying Liu

We introduce Matrix-Game, an interactive world foundation model for controllable game world generation. Matrix-Game is trained using a two-stage pipeline that first performs large-scale unlabeled pretraining for environment understanding,…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Yifan Zhang , Chunli Peng , Boyang Wang , Puyi Wang , Qingcheng Zhu , Fei Kang , Biao Jiang , Zedong Gao , Eric Li , Yang Liu , Yahui Zhou

With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code synthesis. Existing benchmarks rely on idealized assumptions,…

人工智能 · 计算机科学 2026-05-01 Qiyao Wang , Haoran Hu , Longze Chen , Hongbo Wang , Hamid Alinejad-Rokny , Yuan Lin , Min Yang

Collaboration is a cornerstone of society. In the real world, human teammates make use of multi-sensory data to tackle challenging tasks in ever-changing environments. It is essential for embodied agents collaborating in visually-rich…

人工智能 · 计算机科学 2024-12-09 Qian Long , Zhi Li , Ran Gong , Ying Nian Wu , Demetri Terzopoulos , Xiaofeng Gao

While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they often fail to grasp implicit user intentions. Although…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Jun He , Junyan Ye , Zilong Huang , Dongzhi Jiang , Chenjue Zhang , Leqi Zhu , Renrui Zhang , Xiang Zhang , Weijia Li

Human behaviors in real-world environments are inherently interactive, with an individual's motion shaped by surrounding agents and the scene. Such capabilities are essential for applications in virtual avatars, interactive animation, and…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Yaoqin Ye , Yiteng Xu , Qin Sun , Xinge Zhu , Yujing Sun , Yuexin Ma

We propose a new perspective for approaching artificial general intelligence (AGI) through an intelligence foundation model (IFM). Unlike existing foundation models (FMs), which specialize in pattern learning within specific domains such as…

人工智能 · 计算机科学 2025-12-05 Borui Cai , Yao Zhao

With the rapid advancement of autonomous driving technology, a lack of data has become a major obstacle to enhancing perception model accuracy. Researchers are now exploring controllable data generation using world models to diversify…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Xinqing Li , Ruiqi Song , Qingyu Xie , Ye Wu , Nanxin Zeng , Yunfeng Ai

Inspired by cognitive theories of creativity, this paper introduces a computational model (AIGenC) that lays down the necessary components to enable artificial agents to learn, use and generate transferable representations. Unlike machine…

人工智能 · 计算机科学 2023-06-22 Corina Catarau-Cotutiu , Esther Mondragon , Eduardo Alonso

Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely designed for realistic videos-struggle to evaluate animation-style generation with its…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Leyi Wu , Pengjun Fang , Kai Sun , Yazhou Xing , Yinwei Wu , Songsong Wang , Ziqi Huang , Dan Zhou , Yingqing He , Ying-Cong Chen , Qifeng Chen

Humans are known to have an internal "world model" that enables us to carry out action planning based on world states. AI agents need to have such a world model for action planning as well. It is not clear how current AI models, especially…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Delong Chen , Willy Chung , Yejin Bang , Ziwei Ji , Pascale Fung

Data generation is a data augmentation technique for enhancing the generalization ability for skeleton-based human action recognition. Most existing data generation methods face challenges to ensure the temporal consistency of the dynamic…

计算机视觉与模式识别 · 计算机科学 2024-01-31 Long Liu , Xin Wang , Fangming Li , Jiayu Chen

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Inspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current…

计算机视觉与模式识别 · 计算机科学 2019-12-17 Harshala Gammulle , Simon Denman , Sridha Sridharan , Clinton Fookes

Assessing drivers' interaction capabilities is crucial for understanding human driving behavior and enhancing the interactive abilities of autonomous vehicles. In scenarios involving strong interaction, existing metrics focused on…

机器人学 · 计算机科学 2024-05-07 Jiaqi Liu , Peng Hang , Xiangwang Hu , Jian Sun

Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Jiayu Wang , Yang Jiao , Yue Yu , Tianwen Qian , Shaoxiang Chen , Jingjing Chen , Yu-Gang Jiang

With AsgardBench we aim to evaluate visually grounded, high-level action sequence generation and interactive planning, focusing specifically on plan adaptation during execution based on visual observations rather than navigation or…

人工智能 · 计算机科学 2026-03-20 Andrea Tupini , Lars Liden , Reuben Tan , Yu Wang , Jianfeng Gao

Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Taiye Chen , Xun Hu , Zihan Ding , Chi Jin

We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead the field of lightweight models by offering comprehensive…

We introduce WebGames, a comprehensive benchmark suite designed to evaluate general-purpose web-browsing AI agents through a collection of 50+ interactive challenges. These challenges are specifically crafted to be straightforward for…