English
Related papers

Related papers: SAGE: Scalable Agentic 3D Scene Generation for Emb…

200 papers

We introduce SAGE; a Generative LLM for inferring attribute values for products across world-wide e-Commerce catalogs. We introduce a novel formulation of the attribute-value prediction problem as a Seq2Seq summarization task, across…

Information Retrieval · Computer Science 2023-09-13 Athanasios N. Nikolakopoulos , Swati Kaul , Siva Karthik Gade , Bella Dubrov , Umit Batur , Suleiman Ali Khan

This work presents an embodied agent that can adapt its semantic segmentation network to new indoor environments in a fully autonomous way. Because semantic segmentation networks fail to generalize well to unseen environments, the agent…

Robotics · Computer Science 2022-07-05 René Zurbrügg , Hermann Blum , Cesar Cadena , Roland Siegwart , Lukas Schmid

Interacting with human agents in complex scenarios presents a significant challenge for robotic navigation, particularly in environments that necessitate both collision avoidance and collaborative interaction, such as indoor spaces. Unlike…

Robotics · Computer Science 2024-11-07 Lingfeng Sun , Yixiao Wang , Pin-Yun Hung , Changhao Wang , Xiang Zhang , Zhuo Xu , Masayoshi Tomizuka

We present RoboGen, a generative robotic agent that automatically learns diverse robotic skills at scale via generative simulation. RoboGen leverages the latest advancements in foundation and generative models. Instead of directly using or…

Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with human annotators is prohibitively expensive, error-prone, and…

Artificial Intelligence · Computer Science 2026-04-01 Tim R. Davidson , Benoit Seguin , Enrico Bacis , Cesar Ilharco , Hamza Harkous

Agent development kits (ADKs) provide effective platforms and tooling for constructing agents, and their designs are critical to the constructed agents' performance, especially the functionality for agent topology, tools, and memory.…

Inferring full-body poses from Head Mounted Devices, which capture only 3-joint observations from the head and wrists, is a challenging task with wide AR/VR applications. Previous attempts focus on learning one-stage motion mapping and thus…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Fangyu Du , Yang Yang , Xuehao Gao , Hongye Hou

Despite the growing adoption of mixed reality and interactive AI agents, it remains challenging for these systems to generate high quality 2D/3D scenes in unseen environments. The common practice requires deploying an AI agent to collect…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Qiuyuan Huang , Jae Sung Park , Abhinav Gupta , Paul Bennett , Ran Gong , Subhojit Som , Baolin Peng , Owais Khan Mohammed , Chris Pal , Yejin Choi , Jianfeng Gao

Generalization in robotic manipulation remains a critical challenge, particularly when scaling to new environments with limited demonstrations. This paper introduces CAGE, a novel robotic manipulation policy designed to overcome these…

Robotics · Computer Science 2024-12-09 Shangning Xia , Hongjie Fang , Cewu Lu , Hao-Shu Fang

Relying on multi-modal observations, embodied robots (e.g., humanoid robots) could perform multiple robotic manipulation tasks in unstructured real-world environments. However, most language-conditioned behavior-cloning agents in robots…

Robotics · Computer Science 2025-12-30 Wenqi Liang , Gan Sun , Yao He , Yu Ren , Jiahua Dong , Yang Cong

We address the challenge of generating 3D articulated objects in a controllable fashion. Currently, modeling articulated 3D objects is either achieved through laborious manual authoring, or using methods from prior work that are hard to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jiayi Liu , Hou In Ivan Tam , Ali Mahdavi-Amiri , Manolis Savva

We present a fully automatic system that takes a 3D scene and generates plausible 3D human bodies that are posed naturally in that 3D scene. Given a 3D scene without people, humans can easily imagine how people could interact with the scene…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Yan Zhang , Mohamed Hassan , Heiko Neumann , Michael J. Black , Siyu Tang

This paper describes our research on AI agents embodied in visual, virtual or physical forms, enabling them to interact with both users and their environments. These agents, which include virtual avatars, wearable devices, and robots, are…

This thesis introduces "Embodied Spatial Intelligence" to address the challenge of creating robots that can perceive and act in the real world based on natural language instructions. To bridge the gap between Large Language Models (LLMs)…

Robotics · Computer Science 2025-09-03 Jiading Fang

Autonomous agents powered by Large Language Models are transforming AI, creating an imperative for the visualization field to embrace agentic frameworks. However, our field's focus on a human in the sensemaking loop raises critical…

Human-Computer Interaction · Computer Science 2025-09-17 Vaishali Dhanoa , Anton Wolter , Gabriela Molina León , Hans-Jörg Schulz , Niklas Elmqvist

There has been a significant recent progress in the field of Embodied AI with researchers developing models and algorithms enabling embodied agents to navigate and interact within completely unseen environments. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Luca Weihs , Matt Deitke , Aniruddha Kembhavi , Roozbeh Mottaghi

Spatial intelligence is foundational to AI systems that interact with the physical world, particularly in 3D scene generation and spatial comprehension. Current methodologies for 3D scene generation often rely heavily on predefined…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Libin Liu , Shen Chen , Sen Jia , Jingzhe Shi , Zhongyu Jiang , Can Jin , Wu Zongkai , Jenq-Neng Hwang , Lei Li

3D-aware GANs aim to synthesize realistic 3D scenes such that they can be rendered in arbitrary perspectives to produce images. Although previous methods produce realistic images, they suffer from unstable training or degenerate solutions…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Minjung Shin , Yunji Seo , Jeongmin Bae , Young Sun Choi , Hyunsu Kim , Hyeran Byun , Youngjung Uh

Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap,…

Human-Computer Interaction · Computer Science 2025-08-12 Runlin Duan , Yuzhao Chen , Rahul Jain , Yichen Hu , Jingyu Shi , Karthik Ramani

In recent years, predicting driver's focus of attention has been a very active area of research in the autonomous driving community. Unfortunately, existing state-of-the-art techniques achieve this by relying only on human gaze information,…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Anwesan Pal , Sayan Mondal , Henrik I. Christensen