中文
相关论文

相关论文: FantasyHSI: Video-Generation-Centric 4D Human Synt…

200 篇论文

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Xinhao Cai , Minghang Zheng , Xin Jin , Yang Liu

Three-dimensional scene generation holds significant potential in gaming, film, and virtual reality. However, most existing methods adopt a single-step generation process, making it difficult to balance scene complexity with minimal user…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Jiacheng Hong , Kunzhen Wu , Mingrui Yu , Yichao Gu , Shengze Xue , Shuangjiu Xiao , Deli Dong

Generating 3D scenes from natural language holds great promise for applications in gaming, film, and design. However, existing methods struggle with automation, 3D consistency, and fine-grained control. We present DreamScene, an end-to-end…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Haoran Li , Yuli Tian , Kun Lan , Yong Liao , Lin Wang , Pan Hui , Peng Yuan Zhou

Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privileged state or scripted control, which limits scalability to…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Hang Ye , Xiaoxuan Ma , Fan Lu , Wayne Wu , Kwan-Yee Lin , Yizhou Wang

Synthetic 3D scenes are essential for developing Physical AI and generative models. Existing procedural generation methods often have low output throughput, creating a significant bottleneck in scaling up dataset creation. In this work, we…

机器人学 · 计算机科学 2025-12-19 Jinghuan Shang , Harsh Patel , Ran Gong , Karl Schmeckpeper

Interactive humanoid video generation aims to synthesize lifelike visual agents that can engage with humans through continuous and responsive video. Despite recent advances in video synthesis, existing methods often grapple with the…

Multi-agent trajectory generation is a core problem for autonomous driving and intelligent transportation systems. However, efficiently modeling the dynamic interactions between numerous road users and infrastructures in complex scenes…

机器人学 · 计算机科学 2025-12-25 Xiaoyu Mo , Jintian Ge , Zifan Wang , Chen Lv , Karl Henrik Johansson

Synthesizing 3D human motion plays an important role in many graphics applications as well as understanding human activity. While many efforts have been made on generating realistic and natural human motion, most approaches neglect the…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Jiashun Wang , Huazhe Xu , Jingwei Xu , Sifei Liu , Xiaolong Wang

This paper presents a novel approach to generating the 3D motion of a human interacting with a target object, with a focus on solving the challenge of synthesizing long-range and diverse motions, which could not be fulfilled by existing…

计算机视觉与模式识别 · 计算机科学 2023-10-04 Huaijin Pi , Sida Peng , Minghui Yang , Xiaowei Zhou , Hujun Bao

Synthesizing 3D human avatars interacting realistically with a scene is an important problem with applications in AR/VR, video games and robotics. Towards this goal, we address the task of generating a virtual human -- hands and full body…

机器人学 · 计算机科学 2023-03-30 Purva Tendulkar , Dídac Surís , Carl Vondrick

Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object categories, including complex dexterous manipulations that are difficult to capture…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Hyeonwoo Kim , Jeonghwan Kim , Kyungwon Cho , Hanbyul Joo

Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Our goal is to synthesize humans interacting with a given 3D scene…

计算机视觉与模式识别 · 计算机科学 2022-07-27 Kaifeng Zhao , Shaofei Wang , Yan Zhang , Thabo Beeler , Siyu Tang

Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Bin Li , Ruichi Zhang , Han Liang , Jingyan Zhang , Juze Zhang , Xin Chen , Lan Xu , Jingyi Yu , Jingya Wang

We focus on the human-humanoid interaction task optionally with an object. We propose a new task named online full-body motion reaction synthesis, which generates humanoid reactions based on the human actor's motions. The previous work only…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Yunze Liu , Changxi Chen , Li Yi

In this paper, we introduce SynthAI, a new method for the automated creation of High-Level Synthesis (HLS) designs. SynthAI integrates ReAct agents, Chain-of-Thought (CoT) prompting, web search technologies, and the Retrieval-Augmented…

人工智能 · 计算机科学 2024-09-24 Seyed Arash Sheikholeslam , Andre Ivanov

We present HSImul3R, a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casual captures, including sparse-view images and monocular videos. Existing methods suffer from a perception-simulation…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yukang Cao , Haozhe Xie , Fangzhou Hong , Long Zhuo , Zhaoxi Chen , Liang Pan , Ziwei Liu

Creating scenes for captured motions that achieve realistic human-scene interaction is crucial for 3D animation in movies or video games. As character motion is often captured in a blue-screened studio without real furniture or objects in…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Jianan Li , Tao Huang , Qingxu Zhu , Tien-Tsin Wong

Human-Robot Interaction (HRI) is an emerging subfield of service robotics. While most existing approaches rely on explicit signals (i.e. voice, gesture) to engage, current literature is lacking solutions to address implicit user needs. In…

人工智能 · 计算机科学 2022-02-24 Maëlic Neau , Paulo Santos , Anne-Gwenn Bosser , Nathan Beu , Cédric Buche

The increasing labor shortage and aging population underline the need for assistive robots to support human care recipients. To enable safe and responsive assistance, robots require accurate human motion prediction in physical interaction…

机器人学 · 计算机科学 2025-09-15 Saeed Saadatnejad , Reyhaneh Hosseininejad , Jose Barreiros , Katherine M. Tsui , Alexandre Alahi

We present GATSBI, a generative model that can transform a sequence of raw observations into a structured latent representation that fully captures the spatio-temporal context of the agent's actions. In vision-based decision-making…

计算机视觉与模式识别 · 计算机科学 2021-04-12 Cheol-Hui Min , Jinseok Bae , Junho Lee , Young Min Kim