中文
相关论文

相关论文: PLAICraft: Large-Scale Time-Aligned Vision-Speech-…

200 篇论文

Spatial reasoning tasks in multi-agent environments such as event prediction, agent type identification, or missing data imputation are important for multiple applications (e.g., autonomous surveillance over sensor networks and subtasks for…

计算机视觉与模式识别 · 计算机科学 2024-01-10 Sean Kulinski , Nicholas R. Waytowich , James Z. Hare , David I. Inouye

Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Zile Guo , Zhan Chen , Enze Zhu , Kan Wei , Yongkang Zou , Xiaoxuan Liu , Lei Wang

We introduce SynPlay, a large-scale synthetic human dataset purpose-built for advancing multi-perspective human localization, with a predominant focus on aerial-view perception. SynPlay departs from traditional synthetic datasets by…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Jinsub Yim , Hyungtae Lee , Sungmin Eum , Yi-Ting Shen , Yan Zhang , Heesung Kwon , Shuvra S. Bhattacharyya

Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's evolving understanding of the world. When agents can only partially observe their…

多智能体系统 · 计算机科学 2026-05-19 Vardhan Dongre , Dilek Hakkani-Tür

Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption dataset designed to advance research in human-centric…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Yi-Xing Peng , Qize Yang , Yu-Ming Tang , Shenghao Fu , Kun-Yu Lin , Xihan Wei , Wei-Shi Zheng

For efficient human-agent interaction, an agent should proactively recognize their target user and prepare for upcoming interactions. We formulate this challenging problem as the novel task of jointly forecasting a person's intent to…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Tongfei Bian , Yiming Ma , Mathieu Chollet , Victor Sanchez , Tanaya Guha

Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simultaneous (i) near-perfect preservation of each item's identity,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Gemma Canet Tarrés , Manel Baradad , Francesc Moreno-Noguer , Yumeng Li

This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which include virtual avatars, wearable devices, and robots, are…

Mouse dynamics has grown in popularity as a novel irreproducible behavioral biometric. Datasets which contain general unrestricted mouse movements from users are sparse in the current literature. The Balabit mouse dynamics dataset produced…

机器学习 · 计算机科学 2021-10-22 Nyle Siddiqui , Rushit Dave , Naeem Seliya

We propose a novel approach that uses large language models (LLMs) to generate persona-driven conversations between Players and Non-Player Characters (NPC) in games. Showcasing the application of our methodology, we introduce the Minecraft…

We contend that embodied learning is fundamentally a lifecycle problem rather than a single-stage optimization. Systems that optimize only one link (data collection, simulation, learning, or deployment) rarely sustain improvement or…

Real-world tasks of interest are generally poorly defined by human-readable descriptions and have no pre-defined reward signals unless it is defined by a human designer. Conversely, data-driven algorithms are often designed to solve a…

机器学习 · 计算机科学 2022-05-13 Vinicius G. Goecks , Nicholas Waytowich , David Watkins-Valls , Bharat Prakash

We introduce MABe22, a large-scale, multi-agent video and trajectory benchmark to assess the quality of learned behavior representations. This dataset is collected from a variety of biology experiments, and includes triplets of interacting…

Pushing is a fundamental robotic skill. Existing work has shown how to exploit models of pushing to achieve a variety of tasks, including grasping under uncertainty, in-hand manipulation and clearing clutter. Such models, however, are…

As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Kangsan Kim , Yanlai Yang , Suji Kim , Woongyeong Yeo , Youngwan Lee , Mengye Ren , Sung Ju Hwang

Embodied decision-making enables agents to translate high-level goals into executable actions through continuous interactions within the physical world, forming a cornerstone of general-purpose embodied intelligence. Large language models…

Learning physical interaction skills, such as dancing, handshaking, or sparring, remains a fundamental challenge for agents operating in human environments, particularly when the agent's morphology differs significantly from that of the…

机器人学 · 计算机科学 2025-08-05 Tianyu Li , Hengbo Ma , Sehoon Ha , Kwonjoon Lee

We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Jiaxin Lu , Chun-Hao Paul Huang , Uttaran Bhattacharya , Qixing Huang , Yi Zhou

We introduce Voyager, the first LLM-powered embodied lifelong learning agent in Minecraft that continuously explores the world, acquires diverse skills, and makes novel discoveries without human intervention. Voyager consists of three key…

人工智能 · 计算机科学 2023-10-20 Guanzhi Wang , Yuqi Xie , Yunfan Jiang , Ajay Mandlekar , Chaowei Xiao , Yuke Zhu , Linxi Fan , Anima Anandkumar

As digital virtual assistants become ubiquitous, it becomes increasingly important to understand the situated behaviour of users as they interact with these assistants. To this end, we introduce SIMMC, an extension to ParlAI for multi-modal…

计算与语言 · 计算机科学 2020-02-03 Paul A. Crook , Shivani Poddar , Ankita De , Semir Shafi , David Whitney , Alborz Geramifard , Rajen Subba