中文
相关论文

相关论文: PLAICraft: Large-Scale Time-Aligned Vision-Speech-…

200 篇论文

Embodied agents powered by large language models (LLMs), such as Voyager, promise open-ended competence in worlds such as Minecraft. However, when powered by open-weight LLMs they still falter on elementary tasks after domain-specific…

人工智能 · 计算机科学 2025-12-17 Mircea Lică , Ojas Shirekar , Baptiste Colle , Chirag Raman

Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A central challenge is therefore to transform past executions into knowledge that can shape future…

人工智能 · 计算机科学 2026-05-12 Zhengwei Xie , Zhisheng Chen , Ziyan Weng , Jinhan Li , Chenglong Li , Zikai Xiao , Jingwei Song , Jinhao Jing , Vireo Zhang , Kun Wang

The integration of conversational agents into our daily lives has become increasingly common, yet many of these agents cannot engage in deep interactions with humans. Despite this, there is a noticeable shortage of datasets that capture…

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Yu Shang , Xin Zhang , Yinzhou Tang , Lei Jin , Chen Gao , Wei Wu , Yong Li

Multi-embodiment grasping focuses on developing approaches that exhibit generalist behavior across diverse gripper designs. Existing methods often learn the kinematic structure of the robot implicitly and face challenges due to the…

机器人学 · 计算机科学 2026-04-17 Roman Freiberg , Alexander Qualmann , Ngo Anh Vien , Gerhard Neumann

With the surge in the development of large language models, embodied intelligence has attracted increasing attention. Nevertheless, prior works on embodied intelligence typically encode scene or historical memory in an unimodal manner,…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Yang Liu , Xinshuai Song , Kaixuan Jiang , Weixing Chen , Jingzhou Luo , Guanbin Li , Liang Lin

We release a dataset of 65646 StarCraft replays that contains 1535 million frames and 496 million player actions. We provide full game state data along with the original replays that can be viewed in StarCraft. The game state data was…

人工智能 · 计算机科学 2017-10-03 Zeming Lin , Jonas Gehring , Vasil Khalidov , Gabriel Synnaeve

As technology grows and evolves rapidly, it is increasingly clear that mobile devices are more commonly used for sensitive matters than ever before. A need to authenticate users continuously is sought after as a single-factor or multi…

While Vision-Language Models (VLMs) hold promise for tasks requiring extensive collaboration, traditional multi-agent simulators have facilitated rich explorations of an interactive artificial society that reflects collective behavior.…

计算与语言 · 计算机科学 2024-05-24 Xianhao Yu , Jiaqi Fu , Renjia Deng , Wenjuan Han

Generating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales. Fortunately, most music is also highly structured and can be represented as…

We propose a large scale semantic parsing dataset focused on instruction-driven communication with an agent in Minecraft. We describe the data collection process which yields additional 35K human generated instructions with their semantic…

计算与语言 · 计算机科学 2019-05-07 Yacine Jernite , Kavya Srinet , Jonathan Gray , Arthur Szlam

Detecting and interpreting operator actions, engagement, and object interactions in dynamic industrial workflows remains a significant challenge in human-robot collaboration research, especially within complex, real-world environments.…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Naval Kishore Mehta , Arvind , Himanshu Kumar , Abeer Banerjee , Sumeet Saurav , Sanjay Singh

We introduce IMPACT, a synchronized five-view RGB-D dataset for deployment-oriented industrial procedural understanding, built around real assembly and disassembly of a commercial angle grinder with professional-grade tools. To our…

Studies of human cognition often rely on brief, highly controlled tasks that emphasize group-level effects but poorly capture the rich variability within and between individuals. A suite of minigames built on the novel pixelDOPA platform…

人机交互 · 计算机科学 2026-02-12 Dom CP Marticorena , Zeyu Lu , Chris Wissmann , Yash Agarwal , David Garrison , John M Zempel , Dennis L Barbour

It is a long-lasting goal to design a generalist-embodied agent that can follow diverse instructions in human-like ways. However, existing approaches often fail to steadily follow instructions due to difficulties in understanding abstract…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Enshen Zhou , Yiran Qin , Zhenfei Yin , Yuzhou Huang , Ruimao Zhang , Lu Sheng , Yu Qiao , Jing Shao

Human beings possess the capability to multiply a melange of multisensory cues while actively exploring and interacting with the 3D world. Current multi-modal large language models, however, passively absorb sensory data as inputs, lacking…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Yining Hong , Zishuo Zheng , Peihao Chen , Yian Wang , Junyan Li , Chuang Gan

With the emergence of collaborative robots (cobots), human-robot collaboration in industrial manufacturing is coming into focus. For a cobot to act autonomously and as an assistant, it must understand human actions during assembly. To…

机器人学 · 计算机科学 2023-04-18 Dustin Aganian , Benedict Stephan , Markus Eisenbach , Corinna Stretz , Horst-Michael Gross

Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions. Temporally- and spatially-consistent, long-term world…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Yuta Oshima , Yusuke Iwasawa , Masahiro Suzuki , Yutaka Matsuo , Hiroki Furuta

Safe, agile, and socially compliant multi-robot navigation in cluttered and constrained environments remains a critical challenge. This is especially difficult with self-interested agents with unique, unknown priorities in decentralized…

机器人学 · 计算机科学 2026-05-12 Vagul Mahadevan , Shangtong Zhang , Rohan Chandra

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yu Zhang , Kaiyuan Shen , Yang Li