中文
相关论文

相关论文: HumanTOMATO: Text-aligned Whole-body Motion Genera…

200 篇论文

This paper explores a cross-modality synthesis task that infers 3D human-object interactions (HOIs) from a given text-based instruction. Existing text-to-HOI synthesis methods mainly deploy a direct mapping from texts to object-specific 3D…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Xuehao Gao , Yang Yang , Shaoyi Du , Yang Wu , Yebin Liu , Guo-Jun Qi

Text-driven person image generation is an emerging and challenging task in cross-modality image generation. Controllable person image generation promotes a wide range of applications such as digital human interaction and virtual try-on.…

计算机视觉与模式识别 · 计算机科学 2022-11-14 Kaiduo Zhang , Muyi Sun , Jianxin Sun , Binghao Zhao , Kunbo Zhang , Zhenan Sun , Tieniu Tan

Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the complexities of modeling inter-personal dynamics. To address…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Ruihao Xi , Xuekuan Wang , Yongcheng Li , Shuhua Li , Zichen Wang , Yiwei Wang , Feng Wei , Cairong Zhao

Text-to-Motion generation has become a fundamental task in human-machine interaction, enabling the synthesis of realistic human motions from natural language descriptions. Although recent advances in large language models and reinforcement…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Runqi Ouyang , Haoyun Li , Zhenyuan Zhang , Xiaofeng Wang , Zeyu Zhang , Zheng Zhu , Guan Huang , Sirui Han , Xingang Wang

A dominant paradigm for teaching humanoid robots complex skills is to retarget human motions as kinematic references to train reinforcement learning (RL) policies. However, existing retargeting pipelines often struggle with the significant…

Our goal is to train a generative model of 3D hand motions, conditioned on natural language descriptions specifying motion characteristics such as handshapes, locations, finger/hand/arm movements. To this end, we automatically build pairs…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Léore Bensabath , Mathis Petrovich , Gül Varol

Motion synthesis for diverse object categories holds great potential for 3D content creation but remains underexplored due to two key challenges: (1) the lack of comprehensive motion datasets that include a wide range of high-quality…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Wonkwang Lee , Jongwon Jeong , Taehong Moon , Hyeon-Jong Kim , Jaehyeon Kim , Gunhee Kim , Byeong-Uk Lee

Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Ruoxi Guo , Huaijin Pi , Zehong Shen , Qing Shuai , Zechen Hu , Zhumei Wang , Yajiao Dong , Ruizhen Hu , Taku Komura , Sida Peng , Xiaowei Zhou

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Wei Yao , Yunlian Sun , Hongwen Zhang , Yebin Liu , Jinhui Tang

Modeling human-human interactions from text remains challenging because it requires not only realistic individual dynamics but also precise, text-consistent spatiotemporal coupling between agents. Currently, progress is hindered by 1)…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Qingxuan Wu , Zhiyang Dou , Chuan Guo , Yiming Huang , Qiao Feng , Bing Zhou , Jian Wang , Lingjie Liu

3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Shanlin Sun , Gabriel De Araujo , Jiaqi Xu , Shenghan Zhou , Hanwen Zhang , Ziheng Huang , Chenyu You , Xiaohui Xie

Humans can produce complex whole-body motions when interacting with their surroundings, by planning, executing and combining individual limb movements. We investigated this fundamental aspect of motor control in the setting of autonomous…

机器人学 · 计算机科学 2023-08-16 Kai Yuan , Noor Sajid , Karl Friston , Zhibin Li

In this paper, we propose a novel framework, Combo, for harmonious co-speech holistic 3D human motion generation and efficient customizable adaption. In particular, we identify that one fundamental challenge as the…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Chao Xu , Mingze Sun , Zhi-Qi Cheng , Fei Wang , Yang Liu , Baigui Sun , Ruqi Huang , Alexander Hauptmann

People may perform diverse gestures affected by various mental and physical factors when speaking the same sentences. This inherent one-to-many relationship makes co-speech gesture generation from audio particularly challenging.…

计算机视觉与模式识别 · 计算机科学 2023-01-18 Jing Li , Di Kang , Wenjie Pei , Xuefei Zhe , Ying Zhang , Linchao Bao , Zhenyu He

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be…

计算机视觉与模式识别 · 计算机科学 2021-05-10 Lincheng Li , Suzhen Wang , Zhimeng Zhang , Yu Ding , Yixing Zheng , Xin Yu , Changjie Fan

Current techniques face difficulties in generating motions from intricate semantic descriptions, primarily due to insufficient semantic annotations in datasets and weak contextual understanding. To address these issues, we present…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Xin He , Shaoli Huang , Xiaohang Zhan , Chao Weng , Ying Shan

Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We…

Recently, text-to-motion models have opened new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Haonan Han , Xiangzuo Wu , Huan Liao , Zunnan Xu , Zhongyuan Hu , Ronghui Li , Yachao Zhang , Xiu Li

Diverse human motion generation is an increasingly important task, having various applications in computer vision, human-computer interaction and animation. While text-to-motion synthesis using diffusion models has shown success in…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Heechang Kim , Gwanghyun Kim , Se Young Chun
‹ 上一页 1 8 9 10 下一页 ›