English
Related papers

Related papers: HINT: Hierarchical Interaction Modeling for Autore…

200 papers

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Sirui Xu , Dongting Li , Yucheng Zhang , Xiyan Xu , Qi Long , Ziyin Wang , Yunzhi Lu , Shuchang Dong , Hezi Jiang , Akshat Gupta , Yu-Xiong Wang , Liang-Yan Gui

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

In this work, we address the task of unconditional head motion generation to animate still human faces in a low-dimensional semantic space from a single reference pose. Different from traditional audio-conditioned talking head generation…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Louis Airale , Xavier Alameda-Pineda , Stéphane Lathuilière , Dominique Vaufreydaz

Text-to-motion generation has attracted increasing attention in the research community recently, with potential applications in animation, virtual reality, robotics, and human-computer interaction. Diffusion and autoregressive models are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Kang Ding , Hongsong Wang , Jie Gui , Liang Wang

Controlling physics-based humanoids from natural-language instructions is a critical step toward general-purpose embodied agents. However, existing methods remain constrained by a tension between semantic expressiveness and physical…

Graphics · Computer Science 2026-05-26 Jingyan Zhang , Han Liang , Ruichi Zhang , Bin Li , Juze Zhang , Xin Chen , Jingya Wang , Lan Xu , Jingyi Yu

Our goal is to generate realistic human motion from natural language. Modern methods often face a trade-off between model expressiveness and text-to-motion alignment. Some align text and motion latent spaces but sacrifice expressiveness;…

Computer Vision and Pattern Recognition · Computer Science 2024-10-21 Nefeli Andreou , Xi Wang , Victoria Fernández Abrevaya , Marie-Paule Cani , Yiorgos Chrysanthou , Vicky Kalogeiton

Complex scenes present significant challenges for predicting human behaviour due to the abundance of interaction information, such as human-human and humanenvironment interactions. These factors complicate the analysis and understanding of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Caiyi Sun , Yujing Sun , Xiao Han , Zemin Yang , Jiawei Liu , Xinge Zhu , Siu Ming Yiu , Yuexin Ma

Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or regular motions, significant challenges remain, particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Shen Zheng , Jiaran Cai , Yuansheng Guan , Shenneng Huang , Xingpei Ma , Junjie Cao , Hanfeng Zhao , Qiang Zhang , Shunsi Zhang , Xiao-Ping Zhang

Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the complexities of modeling inter-personal dynamics. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ruihao Xi , Xuekuan Wang , Yongcheng Li , Shuhua Li , Zichen Wang , Yiwei Wang , Feng Wei , Cairong Zhao

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

Navigating dense and dynamic environments poses a significant challenge for autonomous driving systems, owing to the intricate nature of multimodal interaction, wherein the actions of various traffic participants and the autonomous vehicle…

Robotics · Computer Science 2024-08-29 Tong Li , Lu Zhang , Sikang Liu , Shaojie Shen

Efficiently detecting human intent to interact with ubiquitous robots is crucial for effective human-robot interaction (HRI) and collaboration. Over the past decade, deep learning has gained traction in this field, with most existing…

Robotics · Computer Science 2025-09-29 Farida Mohsen , Ali Safa

Generating human motion that satisfies customized zero-shot goal functions, enabling applications such as controllable character animation and behavior synthesis for virtual agents, is a critical capability. While current approaches handle…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Hanchao Liu , Fang-Lue Zhang , Shining Zhang , Tai-Jiang Mu , Shi-Min Hu

Text-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Wanjiang Weng , Xiaofeng Tan , Junbo Wang , Guo-Sen Xie , Pan Zhou , Hongsong Wang

Long-range human movement generation remains a central challenge in computer vision and graphics. Generating coherent transitions across semantically distinct motion domains remains largely unexplored. This capability is particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Haichao Wang , Alexander Okupnik , Yuxing Han , Gene Wen , Johannes Schneider , Kyriakos Flouris

Actions are about how we interact with the environment, including other people, objects, and ourselves. In this paper, we propose a novel multi-modal Holistic Interaction Transformer Network (HIT) that leverages the largely ignored, but…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Gueter Josmy Faure , Min-Hung Chen , Shang-Hong Lai

The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines to understand such complex, context-dependent behaviors, it…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Jeonghyeon Na , Sangwon Baik , Inhee Lee , Junyoung Lee , Hanbyul Joo

Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that drive dexterous behavior, finger articulation, contact timing,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zimu Zhang , Yucheng Zhang , Xiyan Xu , Ziyin Wang , Sirui Xu , Kai Zhou , Bing Zhou , Chuan Guo , Jian Wang , Yu-Xiong Wang , Liang-Yan Gui

Generating 3D scenes from human motion sequences supports numerous applications, including virtual reality and architectural design. However, previous auto-regression-based human-aware 3D scene generation methods have struggled to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Xiaolin Hong , Hongwei Yi , Fazhi He , Qiong Cao

Human-centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human-Object Interaction (HOI) videos. Existing works rely on dense control signals, template videos, or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Jiazhi Guan , Quanwei Yang , Luying Huang , Junhao Liang , Borong Liang , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou , Jingdong Wang