English
Related papers

Related papers: InterMask: 3D Human Interaction Generation via Col…

200 papers

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Zan Wang , Yixin Chen , Tengyu Liu , Yixin Zhu , Wei Liang , Siyuan Huang

Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Wenxuan Peng , Bharath Hariharan , Hadar Averbuch-Elor

Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce ViewMask-1-to-3, formulating multi-view synthesis as a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Ruishu Zhu , Zhihao Huang , Jiacheng Sun , Ping Luo , Hongyuan Zhang , Xuelong Li

In this work, we address the problem of 4D facial expressions generation. This is usually addressed by animating a neutral 3D face to reach an expression peak, and then get back to the neutral state. In the real world though, people show…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Naima Otberdout , Claudio Ferrari , Mohamed Daoudi , Stefano Berretti , Alberto Del Bimbo

In this paper, we introduce an innovative task focused on human communication, aiming to generate 3D holistic human motions for both speakers and listeners. Central to our approach is the incorporation of factorization to decouple audio…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Mingze Sun , Chao Xu , Xinyu Jiang , Yang Liu , Baigui Sun , Ruqi Huang

This paper addresses the problem of generating 3D interactive human motion from text. Given a textual description depicting the actions of different body parts in contact with static objects, we synthesize sequences of 3D body poses that…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Sihan Ma , Qiong Cao , Jing Zhang , Dacheng Tao

Despite progress in speech-to-video synthesis, existing methods often struggle to capture cross-individual dependencies and provide fine-grained control over reactive behaviors in dyadic settings. To address these challenges, we propose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dongwei Pan , Longwei Guo , Jiazhi Guan , Luying Huang , Yiding Li , Haojie Liu , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou

Recent advancements in large language models (LLMs) have significantly improved their ability to generate natural and contextually relevant text, enabling more human-like AI interactions. However, generating and understanding interactive…

Artificial Intelligence · Computer Science 2025-03-13 Jeongeun Park , Sungjoon Choi , Sangdoo Yun

Text-to-motion generation is a formidable task, aiming to produce human motions that align with the input text while also adhering to human capabilities and physical laws. While there have been advancements in diffusion models, their…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Hanyang Kong , Kehong Gong , Dongze Lian , Michael Bi Mi , Xinchao Wang

Achieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Hung Yu Ling , Yu-Xiong Wang , Liang-Yan Gui

Since 2023, Vector Quantization (VQ)-based discrete generation methods have rapidly dominated human motion generation, primarily surpassing diffusion-based continuous generation methods in standard performance metrics. However, VQ-based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Zichong Meng , Yiming Xie , Xiaogang Peng , Zeyu Han , Huaizu Jiang

We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Angela Dai

Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these…

Graphics · Computer Science 2025-02-21 Gan Chen , Ying He , Mulin Yu , F. Richard Yu , Gang Xu , Fei Ma , Ming Li , Guang Zhou

Human interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental to a wide range of real-world applications. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Heng Yu , Juze Zhang , Changan Chen , Tiange Xiang , Yusu Fang , Juan Carlos Niebles , Ehsan Adeli

Recent progress in video diffusion models has markedly advanced character animation, which synthesizes motioned videos by animating a static identity image according to a driving video. Explicit methods represent motion using skeleton,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Zhufeng Xu , Xuan Gao , Feng-Lin Liu , Haoxian Zhang , Zhixue Fang , Yu-Kun Lai , Xiaoqiang Liu , Pengfei Wan , Lin Gao

The field of generative image inpainting and object insertion has made significant progress with the recent advent of latent diffusion models. Utilizing a precise object mask can greatly enhance these applications. However, due to the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Jaskirat Singh , Jianming Zhang , Qing Liu , Cameron Smith , Zhe Lin , Liang Zheng

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Jiadong Liang , Feng Lu

Generating realistic reactive motions, in which one person reacts to the fixed motions of others, is challenging due to strict interaction constraints and a limited feasible solution space. This paper focuses on a typical scenario: duet…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Xuhai Chen , Zhi Cen , Huaijin Pi , Sida Peng , Xiaowei Zhou , Yong Liu

Interactional synchrony refers to how the speech or behavior of two or more people involved in a conversation become more finely synchronized with each other, and they can appear to behave almost in direct response to one another. Studies…

Social and Information Networks · Computer Science 2018-07-18 Nicholas Watkins , Ifeoma Nwogu

Generating animated 3D objects is at the heart of many applications, yet most advanced works are typically difficult to apply in practice because of their limited setup, their long runtime, or their limited quality. We introduce ActionMesh,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Remy Sabathier , David Novotny , Niloy J. Mitra , Tom Monnier
‹ Prev 1 3 4 5 6 7 10 Next ›