English
Related papers

Related papers: PLAICraft: Large-Scale Time-Aligned Vision-Speech-…

200 papers

Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct…

Computation and Language · Computer Science 2026-04-14 Yunzhe Wang , Runhui Xu , Kexin Zheng , Tianyi Zhang , Jayavibhav Niranjan Kogundi , Soham Hans , Volkan Ustun

One of the key challenges in visual imitation learning is collecting large amounts of expert demonstrations for a given task. While methods for collecting human demonstrations are becoming easier with teleoperation methods and the use of…

Robotics · Computer Science 2021-07-20 Sarah Young , Jyothish Pari , Pieter Abbeel , Lerrel Pinto

Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as…

Machine Learning · Computer Science 2022-06-24 Bowen Baker , Ilge Akkaya , Peter Zhokhov , Joost Huizinga , Jie Tang , Adrien Ecoffet , Brandon Houghton , Raul Sampedro , Jeff Clune

Training effective AI agents for multi-turn interactions requires high-quality data that captures realistic human-agent dynamics, yet such data is scarce and expensive to collect manually. We introduce APIGen-MT, a two-phase framework that…

Instructing a robot to complete an everyday task within our homes has been a long-standing challenge for robotics. While recent progress in language-conditioned imitation learning and offline reinforcement learning has demonstrated…

Robotics · Computer Science 2024-07-03 Federico Ceola , Lorenzo Natale , Niko Sünderhauf , Krishan Rana

Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interaction. Such…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Jiawei Mo , Yixuan Chen , Rifen Lin , Yongkang Ni , Min Zeng , Xiping Hu , Min Li

We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environments. We also provide a rich variety of playable entities,…

Artificial Intelligence · Computer Science 2025-08-13 Fangwei Zhong , Kui Wu , Churan Wang , Hao Chen , Hai Ci , Zhoujun Li , Yizhou Wang

Relying on multi-modal observations, embodied robots (e.g., humanoid robots) could perform multiple robotic manipulation tasks in unstructured real-world environments. However, most language-conditioned behavior-cloning agents in robots…

Robotics · Computer Science 2025-12-30 Wenqi Liang , Gan Sun , Yao He , Yu Ren , Jiahua Dong , Yang Cong

World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Jialong Wu , Shaofeng Yin , Ningya Feng , Xu He , Dong Li , Jianye Hao , Mingsheng Long

AI agents have been evaluated in isolation or within small groups, where interactions remain limited in scope and complexity. Large-scale simulations involving many autonomous agents -- reflecting the full spectrum of civilizational…

Effectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception…

Robotics · Computer Science 2025-03-24 Wenbo Cui , Chengyang Zhao , Songlin Wei , Jiazhao Zhang , Haoran Geng , Yaran Chen , Haoran Li , He Wang

A key feature of human collaboration is the ability to iteratively refine the concepts we have communicated. In contrast, while generative AI excels at the \textit{generation} of content, it often struggles to make specific language-guided…

Artificial Intelligence · Computer Science 2025-04-30 William P. McCarthy , Saujas Vaduguru , Karl D. D. Willis , Justin Matejka , Judith E. Fan , Daniel Fried , Yewen Pu

Next generation virtual assistants are envisioned to handle multimodal inputs (e.g., vision, memories of previous interactions, in addition to the user's utterances), and perform multimodal actions (e.g., displaying a route in addition to…

The domain of Embodied AI, in which agents learn to complete tasks through interaction with their environment from egocentric observations, has experienced substantial growth with the advent of deep reinforcement learning and increased…

Computer Vision and Pattern Recognition · Computer Science 2020-08-31 Luca Weihs , Jordi Salvador , Klemen Kotar , Unnat Jain , Kuo-Hao Zeng , Roozbeh Mottaghi , Aniruddha Kembhavi

A fundamental question in cognitive science and AI concerns whether different learning modalities: language, vision, and action, give rise to distinct or shared internal representations. Traditional views assume that models trained on…

Artificial Intelligence · Computer Science 2026-02-02 Nicola Milano , Stefano Nolfi

Metaverse platforms are rapidly evolving to provide immersive spaces for user interaction and content creation. However, the generation of dynamic and interactive 3D objects remains challenging due to the need for advanced 3D modeling and…

Human-Computer Interaction · Computer Science 2025-05-01 Ryutaro Kurai , Takefumi Hiraki , Yuichi Hiroi , Yutaro Hirao , Monica Perusquía-Hernández , Hideaki Uchiyama , Kiyoshi Kiyokawa

Multi-agent behavior modeling aims to understand the interactions that occur between agents. We present a multi-agent dataset from behavioral neuroscience, the Caltech Mouse Social Interactions (CalMS21) Dataset. Our dataset consists of…

In simultaneous interpreting, an interpreter renders a source speech into another language with a very short lag, much sooner than sentences are finished. In order to understand and later reproduce this dynamic and complex task…

Computation and Language · Computer Science 2025-06-06 Dávid Javorský , Ondřej Bojar , François Yvon

Dense pixel-specific representation learning at scale has been bottlenecked due to the unavailability of large-scale multi-view datasets. Current methods for building effective pretraining datasets heavily rely on annotated 3D meshes, point…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Kalyani Marathe , Mahtab Bigverdi , Nishat Khan , Tuhin Kundu , Patrick Howe , Sharan Ranjit S , Anand Bhattad , Aniruddha Kembhavi , Linda G. Shapiro , Ranjay Krishna

Human motion is highly diverse and dynamic, posing challenges for imitation learning algorithms that aim to generalize motor skills for controlling simulated characters. Previous methods typically rely on a universal full-body controller…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Yiming Huang , Zhiyang Dou , Lingjie Liu