English
Related papers

Related papers: CoMo: Controllable Motion Generation through Langu…

200 papers

In this paper, we focus on motion discrete tokenization, which converts raw motion into compact discrete tokens--a process proven crucial for efficient motion generation. In this paradigm, increasing the number of tokens is a common…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Sheng Yan , Yong Wang , Xin Du , Junsong Yuan , Mengyuan Liu

Text-to-motion diffusion models can generate realistic animations from text prompts, but do not support fine-grained motion editing controls. In this paper, we present a method for using natural language to iteratively specify local edits…

Graphics · Computer Science 2024-06-04 Purvi Goel , Kuan-Chieh Wang , C. Karen Liu , Kayvon Fatahalian

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counterpart,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xuanke Shi , Boxuan Li , Xiaoyang Han , Zhongang Cai , Lei Yang , Quan Wang , Dahua Lin

When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Gwantae Kim , Seonghyeok Noh , Insung Ham , Hanseok Ko

Human pose estimation traditionally relies on architectures that encode keypoint priors, limiting their generalization to novel poses or unseen keypoints. Recent language-guided approaches like LocLLM reformulate keypoint localization as a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Dewen Zhang , Tahir Hussain , Wangpeng An , Hayaru Shouno

Large Language Models (LLMs) have demonstrated significant potential in code generation. However, in the factory automation sector, particularly motion control, manual programming, alongside inefficient and unsafe debugging practices,…

Artificial Intelligence · Computer Science 2025-07-03 Yin Li , Liangwei Wang , Shiyuan Piao , Boo-Ho Yang , Ziyue Li , Wei Zeng , Fugee Tsung

The task of text2motion is to generate human motion sequences from given textual descriptions, where the model explores diverse mappings from natural language instructions to human body movements. While most existing works are confined to…

Artificial Intelligence · Computer Science 2024-03-27 Kunhang Li , Yansong Feng

While recent advances in text-to-motion generation have shown promising results, they typically assume all individuals are grouped as a single unit. Scaling these methods to handle larger crowds and ensuring that individuals respond…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Yukang Cao , Xinying Guo , Mingyuan Zhang , Haozhe Xie , Chenyang Gu , Ziwei Liu

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Shuling Zhao , Fa-Ting Hong , Xiaoshui Huang , Dan Xu

Character animation in real-world scenarios necessitates a variety of constraints, such as trajectories, key-frames, interactions, etc. Existing methodologies typically treat single or a finite set of these constraint(s) as separate control…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Hanchao Liu , Xiaohang Zhan , Shaoli Huang , Tai-Jiang Mu , Ying Shan

Generating realistic human motion is essential for many computer vision and graphics applications. The wide variety of human body shapes and sizes greatly impacts how people move. However, most existing motion models ignore these…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Shashank Tripathi , Omid Taheri , Christoph Lassner , Michael J. Black , Daniel Holden , Carsten Stoll

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

Human motion is highly expressive and naturally aligned with language, yet prevailing methods relying heavily on joint text-motion embeddings struggle to synthesize temporally accurate, detailed motions and often lack explainability. To…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Junkun Jiang , Ho Yin Au , Jingyu Xiang , Jie Chen

Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise in this task,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Bao Li , Xiaomei Zhang , Miao Xu , Zhaoxin Fan , Xiangyu Zhu , Zhen Lei

Generating 3D human motion from text descriptions remains challenging due to the diverse and complex nature of human motion. While existing methods excel within the training distribution, they often struggle with out-of-distribution…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Zongye Zhang , Bohan Kong , Qingjie Liu , Yunhong Wang

Language plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP's pretraining on…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zhe Li , Weihao Yuan , Yisheng He , Lingteng Qiu , Shenhao Zhu , Xiaodong Gu , Weichao Shen , Yuan Dong , Zilong Dong , Laurence T. Yang

Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub-actions or incoherent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Miaowei Wang , Qingxuan Yan , Zhi Cao , Yayuan Li , Oisin Mac Aodha , Jason J. Corso , Amir Vaxman

Generating animations from natural language sentences finds its applications in a a number of domains such as movie script visualization, virtual human animation and, robot motion planning. These sentences can describe different kinds of…

Computer Vision and Pattern Recognition · Computer Science 2019-12-02 Chaitanya Ahuja , Louis-Philippe Morency

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes,…

Graphics · Computer Science 2025-05-05 Jiefeng Li , Jinkun Cao , Haotian Zhang , Davis Rempe , Jan Kautz , Umar Iqbal , Ye Yuan

Diffusion models have recently gained significant attention in robotics due to their ability to generate multi-modal distributions of system states and behaviors. However, a key challenge remains: ensuring precise control over the generated…

Robotics · Computer Science 2025-10-01 Luobin Wang , Hongzhan Yu , Chenning Yu , Sicun Gao , Henrik Christensen