English
Related papers

Related papers: CoMo: Controllable Motion Generation through Langu…

200 papers

We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In MoMask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Chuan Guo , Yuxuan Mu , Muhammad Gohar Javed , Sen Wang , Li Cheng

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Sohan Anisetty , James Hays

Pretrained Transformer-based language models (LMs) display remarkable natural language generation capabilities. With their immense potential, controlling text generation of such LMs is getting attention. While there are studies that seek to…

Computation and Language · Computer Science 2022-06-13 Alvin Chan , Yew-Soon Ong , Bill Pung , Aston Zhang , Jie Fu

Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during task execution, such as force, which visual data alone cannot…

Robotics · Computer Science 2025-04-21 Chen Wang , Fei Xia , Wenhao Yu , Tingnan Zhang , Ruohan Zhang , C. Karen Liu , Li Fei-Fei , Jie Tan , Jacky Liang

We introduce an approach for detecting and tracking detailed 3D poses of multiple people from a single monocular camera stream. Our system maintains temporally coherent predictions in crowded scenes filled with difficult poses and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Alejandro Newell , Peiyun Hu , Lahav Lipson , Stephan R. Richter , Vladlen Koltun

Aligning multiple modalities in a latent space, such as images and texts, has shown to produce powerful semantic visual representations, fueling tasks like image captioning, text-to-image generation, or image grounding. In the context of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ginger Delmas , Philippe Weinzaepfel , Francesc Moreno-Noguer , Grégory Rogez

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training data. By leveraging a few video samples demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Joanna Materzynska , Josef Sivic , Eli Shechtman , Antonio Torralba , Richard Zhang , Bryan Russell

In recent years, generative artificial intelligence has achieved significant advancements in the field of image generation, spawning a variety of applications. However, video generation still faces considerable challenges in various…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Yuang Zhang , Jiaxi Gu , Li-Wen Wang , Han Wang , Junqi Cheng , Yuefeng Zhu , Fangyuan Zou

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Boyuan Wang , Xiaofeng Wang , Chaojun Ni , Guosheng Zhao , Zhiqin Yang , Zheng Zhu , Muyang Zhang , Yukun Zhou , Xinze Chen , Guan Huang , Lihong Liu , Xingang Wang

As large-scale language model pretraining pushes the state-of-the-art in text generation, recent work has turned to controlling attributes of the text such models generate. While modifying the pretrained models via fine-tuning remains the…

Computation and Language · Computer Science 2021-08-05 Sachin Kumar , Eric Malmi , Aliaksei Severyn , Yulia Tsvetkov

Composed pose retrieval (CPR) enables users to search for human poses by specifying a reference pose and a transition description, but progress in this field is hindered by the scarcity and inconsistency of annotated pose transitions.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Yi-Ting Shen , Sungmin Eum , Doheon Lee , Rohit Shete , Chiao-Yi Wang , Heesung Kwon , Shuvra S. Bhattacharyya

Motion generation is fundamental to computer animation and widely used across entertainment, robotics, and virtual environments. While recent methods achieve impressive results, most rely on fixed skeletal templates, which prevent them from…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Keyi Chen , Mingze Sun , Zhenyu Liu , Zhangquan Chen , Ruqi Huang

Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multimodal fusion. Ideally, replacing a textual question with its…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Feng Han , Zhixiong Zhang , Zheming Liang , Yibin Wang , Jiaqi Wang

We explore the human motion knowledge of Large Language Models (LLMs) through 3D avatar control. Given a motion instruction, we prompt LLMs to first generate a high-level movement plan with consecutive steps (High-level Planning), then…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Kunhang Li , Jason Naradowsky , Yansong Feng , Yusuke Miyao

Inspired by the recent success of LLMs, the field of human motion understanding has increasingly shifted toward developing large motion models. Despite some progress, current efforts remain far from achieving truly generalist models,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Ye Wang , Sipeng Zheng , Bin Cao , Qianshan Wei , Weishuai Zeng , Qin Jin , Zongqing Lu

Text animation, a foundational element in video creation, enables efficient and cost-effective communication, thriving in advertisements, journalism, and social media. However, traditional animation workflows present significant usability…

Human-Computer Interaction · Computer Science 2025-06-13 Bao Zhang , Zihan Li , Zhenglei Liu , Huanchen Wang , Yuxin Ma

Recent virtual try-on approaches have advanced by finetuning pre-trained text-to-image diffusion models to leverage their powerful generative ability. However, the use of text prompts in virtual try-on remains underexplored. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Jeongho Kim , Hoiyeong Jin , Sunghyun Park , Jaegul Choo

Recent research explores optimization using large language models (LLMs) by either iteratively seeking next-step solutions from LLMs or directly prompting LLMs for an optimizer. However, these approaches exhibit inherent limitations,…

Optimization and Control · Mathematics 2024-03-06 Zeyuan Ma , Hongshu Guo , Jiacheng Chen , Guojun Peng , Zhiguang Cao , Yining Ma , Yue-Jiao Gong

Co-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons, resulting in the omission of appearance…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Xu He , Qiaochu Huang , Zhensong Zhang , Zhiwei Lin , Zhiyong Wu , Sicheng Yang , Minglei Li , Zhiyi Chen , Songcen Xu , Xiaofei Wu

This paper investigates controllable generation for large language models (LLMs) with prompt-based control, focusing on Lexically Constrained Generation (LCG). We systematically evaluate the performance of LLMs on satisfying lexical…

Computation and Language · Computer Science 2024-10-08 Bingxuan Li , Yiwei Wang , Tao Meng , Kai-Wei Chang , Nanyun Peng
‹ Prev 1 8 9 10 Next ›