中文
相关论文

相关论文: InterMask: 3D Human Interaction Generation via Col…

200 篇论文

We present a physics-based character control framework for synthesizing human-scene interactions. Recent advances adopt physics simulation to mitigate artifacts produced by data-driven kinematic approaches. However, existing physics-based…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Liang Pan , Jingbo Wang , Buzhen Huang , Junyu Zhang , Haofan Wang , Xu Tang , Yangang Wang

Recent advances in text-to-motion generation using diffusion and autoregressive models have shown promising results. However, these models often suffer from a trade-off between real-time performance, high fidelity, and motion editability.…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Ekkasit Pinyoanuntapong , Pu Wang , Minwoo Lee , Chen Chen

Understanding details of human multimodal interaction can elucidate many aspects of the type of information processing machines must perform to interact with humans. This article gives an overview of recent findings from Linguistics…

计算与语言 · 计算机科学 2020-08-10 João Ranhel , Cacilda Vilela

The rapid progress of Large Multimodal Models (LMMs) and cloud-based AI agents is transforming human-AI collaboration into bidirectional, multimodal interaction. However, existing codecs remain optimized for unimodal, one-way communication,…

人工智能 · 计算机科学 2025-09-29 Qi Mao , Tinghan Yang , Jiahao Li , Bin Li , Libiao Jin , Yan Lu

This paper presents a novel approach to generating the 3D motion of a human interacting with a target object, with a focus on solving the challenge of synthesizing long-range and diverse motions, which could not be fulfilled by existing…

计算机视觉与模式识别 · 计算机科学 2023-10-04 Huaijin Pi , Sida Peng , Minghui Yang , Xiaowei Zhou , Hujun Bao

Human motion generation has made tremendous progress in recent years, with state-of-the-art approaches surpassing ground truth data in leading evaluation benchmarks. However, visual inspection of the generated motions paints a different…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Pascal Herrmann , Maarten Bieshaar , Dennis Mack , Robert Herzog , Juergen Gall

A social interaction is a social exchange between two or more individuals,where individuals modify and adjust their behaviors in response to their interaction partners. Our social interactions are one of most fundamental aspects of our…

计算机视觉与模式识别 · 计算机科学 2018-02-01 Behnaz Nojavanasghari , Yuchi Huang , Saad Khan

Co-speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user-specified gestural style. Existing VQ-VAE based co-speech gesture generation methods improve…

图形学 · 计算机科学 2026-05-11 Junchuan Zhao , Qifan Liang , Ye Wang

Speech-driven 3D facial animation technology has been developed for years, but its practical application still lacks expectations. The main challenges lie in data limitations, lip alignment, and the naturalness of facial expressions.…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Xiangyu Liang , Wenlin Zhuang , Tianyong Wang , Guangxing Geng , Guangyue Geng , Haifeng Xia , Siyu Xia

3D multi-person motion prediction is a highly complex task, primarily due to the dependencies on both individual past movements and the interactions between agents. Moreover, effectively modeling these interactions often incurs substantial…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Yuanhong Zheng , Ruixuan Yu , Jian Sun

Generating 3D scenes from human motion sequences supports numerous applications, including virtual reality and architectural design. However, previous auto-regression-based human-aware 3D scene generation methods have struggled to…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Xiaolin Hong , Hongwei Yi , Fazhi He , Qiong Cao

This paper introduces the first text-guided work for generating the sequence of hand-object interaction in 3D. The main challenge arises from the lack of labeled data where existing ground-truth datasets are nowhere near generalizable in…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Junuk Cha , Jihyeon Kim , Jae Shin Yoon , Seungryul Baek

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

Reconstructing human-object interaction in 3D from a single RGB image is a challenging task and existing data driven methods do not generalize beyond the objects present in the carefully curated 3D interaction datasets. Capturing…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Xianghui Xie , Bharat Lal Bhatnagar , Jan Eric Lenssen , Gerard Pons-Moll

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on training sequences that contain captured human…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Kaifeng Zhao , Yan Zhang , Shaofei Wang , Thabo Beeler , Siyu Tang

This paper addresses a novel task of anticipating 3D human-object interactions (HOIs). Most existing research on HOI synthesis lacks comprehensive whole-body interactions with dynamic objects, e.g., often limited to manipulating small or…

计算机视觉与模式识别 · 计算机科学 2023-09-01 Sirui Xu , Zhengyuan Li , Yu-Xiong Wang , Liang-Yan Gui

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Sirui Xu , Dongting Li , Yucheng Zhang , Xiyan Xu , Qi Long , Ziyin Wang , Yunzhi Lu , Shuchang Dong , Hezi Jiang , Akshat Gupta , Yu-Xiong Wang , Liang-Yan Gui

Social interactions incorporate nonverbal signals to convey emotions alongside speech, including facial expressions and body gestures. Generative models have demonstrated promising results in creating full-body nonverbal animations…

人机交互 · 计算机科学 2026-04-01 Kiran Chhatre , Renan Guarese , Andrii Matviienko , Christopher Peters

Reconstructing a 3D hand mesh from a single RGB image is challenging due to complex articulations, self-occlusions, and depth ambiguities. Traditional discriminative methods, which learn a deterministic mapping from a 2D image to a single…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Muhammad Usama Saleem , Ekkasit Pinyoanuntapong , Mayur Jagdishbhai Patel , Hongfei Xue , Ahmed Helmy , Srijan Das , Pu Wang

Generating realistic human motion with high-level controls is a crucial task for social understanding, robotics, and animation. With high-quality MOCAP data becoming more available recently, a wide range of data-driven approaches have been…

图形学 · 计算机科学 2025-07-29 Wenning Xu , Shiyu Fan , Paul Henderson , Edmond S. L. Ho