English
Related papers

Related papers: ReMoGen: Real-time Human Interaction-to-Reaction G…

200 papers

Recent advances in video diffusion models shows promise for generating robotic decision-making data, with trajectory conditions further enabling fine-grained control. However, existing methods primarily focus on individual object motion and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Xiao Fu , Xintao Wang , Xian Liu , Jianhong Bai , Runsen Xu , Pengfei Wan , Di Zhang , Dahua Lin

The scale and quality of datasets are crucial for training robust perception models. However, obtaining large-scale annotated data is both costly and time-consuming. Generative models have emerged as a powerful tool for data augmentation by…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Haowei Zhu , Tianxiang Pan , Rui Qin , Jun-Hai Yong , Bin Wang

Emotion serves as an essential component in daily human interactions. Existing human motion generation frameworks do not consider the impact of emotions, which reduces naturalness and limits their application in interactive tasks, such as…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Chen Zhu , Buzhen Huang , Zijing Wu , Binghui Zuo , Yangang Wang

To effectively engage in human society, the ability to adapt, filter information, and make informed decisions in ever-changing situations is critical. As robots and intelligent agents become more integrated into human life, there is a…

Artificial Intelligence · Computer Science 2025-11-13 Mingyang Mao , Mariela M. Perez-Cabarcas , Utteja Kallakuri , Nicholas R. Waytowich , Xiaomin Lin , Tinoosh Mohsenin

Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Jianrong Zhang , Hehe Fan , Yi Yang

Effective human-robot interaction requires robots to identify human intentions and generate expressive, socially appropriate motions in real-time. Existing approaches often rely on fixed motion libraries or computationally expensive…

Robotics · Computer Science 2025-09-30 Lingfan Bao , Yan Pan , Tianhu Peng , Dimitrios Kanoulas , Chengxu Zhou

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Jiepeng Wang , Zhaoqing Wang , Hao Pan , Yuan Liu , Dongdong Yu , Changhu Wang , Wenping Wang

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

Can we make virtual characters in a scene interact with their surrounding objects through simple instructions? Is it possible to synthesize such motion plausibly with a diverse set of objects and instructions? Inspired by these questions,…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Anindita Ghosh , Rishabh Dabral , Vladislav Golyanik , Christian Theobalt , Philipp Slusallek

Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archon, a fully…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chong Bao , Shichen Liu , Lijun Yu , David Futschik , Stylianos Moschoglou , Shefali Srivastava , Ziqian Bai , Feitong Tan , Guofeng Zhang , Zhaopeng Cui , Sean Fanello , Yinda Zhang

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xiang Deng , Feng Gao , Yong Zhang , Youxin Pang , Xu Xiaoming , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to work robustly under different spatial distribution of objects, environment and agent itself. To…

Robotics · Computer Science 2026-04-30 Xiuwei Xu , Angyuan Ma , Hankun Li , Bingyao Yu , Zheng Zhu , Jie Zhou , Jiwen Lu

Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, self-occlusions, perspective distortion, and noisy ego-motion. Existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Bohan Zhou , Yi Zhan , Zhongbin Zhang , Zongqing Lu

Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often yields unnatural or implausible outcomes, especially by…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Lee Hsin-Ying , Hanwen Jiang , Yiqun Mei , Jing Shi , Ming-Hsuan Yang , Zhixin Shu

Imitation learning is a promising paradigm for training robot control policies, but these policies can suffer from distribution shift, where the conditions at evaluation time differ from those in the training data. A popular approach for…

Robotics · Computer Science 2024-05-03 Ryan Hoque , Ajay Mandlekar , Caelan Garrett , Ken Goldberg , Dieter Fox

While recent advances in text-to-motion generation have shown promising results, they typically assume all individuals are grouped as a single unit. Scaling these methods to handle larger crowds and ensuring that individuals respond…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Yukang Cao , Xinying Guo , Mingyuan Zhang , Haozhe Xie , Chenyang Gu , Ziwei Liu

Unifying diverse image generation tasks within a single framework remains a fundamental challenge in visual generation. While large language models (LLMs) achieve unification through task-agnostic data and generation, existing visual…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yijing Lin , Mengqi Huang , Shuhan Zhuang , Zhendong Mao

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Nan Jiang , Zimo He , Zi Wang , Hongjie Li , Yixin Chen , Siyuan Huang , Yixin Zhu

This paper proposes a novel deep learning framework for multi-modal motion prediction. The framework consists of three parts: recurrent neural networks to process the target agent's motion process, convolutional neural networks to process…

Robotics · Computer Science 2022-07-05 Zhiyu Huang , Xiaoyu Mo , Chen Lv

Human motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Wentao Zhu , Xiaoxuan Ma , Dongwoo Ro , Hai Ci , Jinlu Zhang , Jiaxin Shi , Feng Gao , Qi Tian , Yizhou Wang
‹ Prev 1 3 4 5 6 7 10 Next ›