English
Related papers

Related papers: TriC-Motion: Tri-Domain Causal Modeling Grounded T…

200 papers

The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Bingqi Ma , Linlong Lang , Ming Zhang , Dailan He , Xingtong Ge , Yi Zhang , Guanglu Song , Yu Liu

Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Ruoxi Guo , Huaijin Pi , Zehong Shen , Qing Shuai , Zechen Hu , Zhumei Wang , Yajiao Dong , Ruizhen Hu , Taku Komura , Sida Peng , Xiaowei Zhou

This paper aims to model 3D human motion across domains, where a single model is expected to handle multiple modalities, tasks, and datasets. Existing cross-domain models often rely on domain-specific components and multi-stage training,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Mengyuan Liu , Xinshun Wang , Zhongbin Fang , Deheng Ye , Xia Li , Tao Tang , Songtao Wu , Xiangtai Li , Ming-Hsuan Yang

Our goal is to synthesize 3D human motions given textual inputs describing simultaneous actions, for example 'waving hand' while 'walking' at the same time. We refer to generating such simultaneous movements as performing 'spatial…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Nikos Athanasiou , Mathis Petrovich , Michael J. Black , Gül Varol

Video-based world models have recently garnered increasing attention for their ability to synthesize diverse and dynamic visual environments. In this paper, we focus on shared world modeling, where a model generates multiple videos from a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Fan Wu , Jiacheng Wei , Ruibo Li , Yi Xu , Junyou Li , Deheng Ye , Guosheng Lin

While large-scale datasets have driven significant progress in Text-to-Video (T2V) generative models, these models remain highly sensitive to input prompts, demonstrating that prompt design is critical to generation quality. Current methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zillur Rahman , Alex Sheng , Cristian Meo

The task of text2motion is to generate human motion sequences from given textual descriptions, where the model explores diverse mappings from natural language instructions to human body movements. While most existing works are confined to…

Artificial Intelligence · Computer Science 2024-03-27 Kunhang Li , Yansong Feng

Existing text-driven motion generation methods often treat synthesis as a bidirectional mapping between language and motion, but remain limited in capturing the causal logic of action execution and the human intentions that drive behavior.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Junyu Shi , Yong Sun , Zhiyuan Zhang , Lijiang Liu , Zhengjie Zhang , Yuxin He , Qiang Nie

Diffusion models have demonstrated exceptional capability in generating high-quality images, videos, and audio. Due to their adaptiveness in iterative refinement, they provide a strong potential for achieving better non-autoregressive…

Computation and Language · Computer Science 2024-02-26 Yuxuan Liu , Tianchi Yang , Shaohan Huang , Zihan Zhang , Haizhen Huang , Furu Wei , Weiwei Deng , Feng Sun , Qi Zhang

Current state-of-the-art paradigms predominantly treat Text-to-Motion (T2M) generation as a direct translation problem, mapping symbolic language directly to continuous poses. While effective for simple actions, this System 1 approach faces…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yijie Qian , Juncheng Wang , Yuxiang Feng , Chao Xu , Wang Lu , Yang Liu , Baigui Sun , Yiqiang Chen , Yong Liu , Shujun Wang

Controllable text generation concerns two fundamental tasks of wide applications, namely generating text of given attributes (i.e., attribute-conditional generation), and minimally editing existing text to possess desired attributes (i.e.,…

Computation and Language · Computer Science 2022-01-25 Zhiting Hu , Li Erran Li

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired…

Sound · Computer Science 2024-10-08 Han Yang , Kun Su , Yutong Zhang , Jiaben Chen , Kaizhi Qian , Gaowen Liu , Chuang Gan

Motion generation is fundamental to computer animation and widely used across entertainment, robotics, and virtual environments. While recent methods achieve impressive results, most rely on fixed skeletal templates, which prevent them from…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Keyi Chen , Mingze Sun , Zhenyu Liu , Zhangquan Chen , Ruqi Huang

Sound content creation, essential for multimedia works such as video games and films, often involves extensive trial-and-error, enabling creators to semantically reflect their artistic ideas and inspirations, which evolve throughout the…

Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizing compound actions from texts, which integrate multiple concurrent actions into…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Yue Jiang , Mingyu Yang , Liuyuxin Yang , Yang Xu , Bingxin Yun , Yuhe Zhang

Text-to-motion generation holds significant potential for cross-linguistic applications, yet it is hindered by the lack of bilingual datasets and the poor cross-lingual semantic understanding of existing language models. To address these…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Wanjiang Weng , Xiaofeng Tan , Xiangbo Shu , Guo-Sen Xie , Pan Zhou , Hongsong Wang

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Drug discovery is a time-consuming and expensive process, with traditional high-throughput and docking-based virtual screening hampered by low success rates and limited scalability. Recent advances in generative modelling, including…

Artificial Intelligence · Computer Science 2026-03-12 Junkai Ji , Zhangfan Yang , Dong Xu , Ruibin Bai , Jianqiang Li , Tingjun Hou , Zexuan Zhu

3D content inherently encompasses multi-modal characteristics and can be projected into different modalities (e.g., RGB images, RGBD, and point clouds). Each modality exhibits distinct advantages in 3D asset modeling: RGB images contain…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Ziang Cao , Zhaoxi Chen , Liang Pan , Ziwei Liu

Current video generation models usually convert signals indicating appearance and motion received from inputs (e.g., image, text) or latent spaces (e.g., noise vectors) into consecutive frames, fulfilling a stochastic generation process for…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Xue Song , Jingjing Chen , Bin Zhu , Yu-Gang Jiang
‹ Prev 1 3 4 5 6 7 10 Next ›