中文
相关论文

相关论文: MotionMERGE: A Multi-granular Framework for Human …

200 篇论文

Zero-shot Text-to-Video synthesis generates videos based on prompts without any videos. Without motion information from videos, motion priors implied in prompts are vital guidance. For example, the prompt "airplane landing on the runway"…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Sitong Su , Litao Guo , Lianli Gao , Hengtao Shen , Jingkuan Song

We introduce Human Motion Unlearning and motivate it through the concrete task of preventing violent 3D motion synthesis, an important safety requirement given that popular text-to-motion datasets (HumanML3D and Motion-X) contain from 7\%…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Edoardo De Matteis , Matteo Migliarini , Alessio Sampieri , Indro Spinelli , Fabio Galasso

We present a generative model that learns to synthesize human motion from limited training sequences. Our framework provides conditional generation and blending across multiple temporal resolutions. The model adeptly captures human motion…

计算机视觉与模式识别 · 计算机科学 2024-11-26 David Eduardo Moreno-Villamarín , Anna Hilsmann , Peter Eisert

Human motion generation is a critical task with a wide range of applications. Achieving high realism in generated motions requires naturalness, smoothness, and plausibility. Despite rapid advancements in the field, current generation…

计算机视觉与模式识别 · 计算机科学 2025-01-24 Haoru Wang , Wentao Zhu , Luyi Miao , Yishu Xu , Feng Gao , Qi Tian , Yizhou Wang

Human motion analysis tasks, such as temporal 3D pose estimation, motion prediction, and motion in-betweening, play an essential role in computer vision. However, current paradigms suffer from severe fragmentation. First, the field is split…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Xinshun Wang , Peiming Li , Ziyi Wang , Zhongbin Fang , Zhichao Deng , Songtao Wu , Jason Li , Mengyuan Liu

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critical challenges: the…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Wanjiang Weng , Xiaofeng Tan , Hongsong Wang , Pan Zhou

Storytelling video generation (SVG) aims to produce coherent and visually rich multi-scene videos that follow a structured narrative. Existing methods primarily employ LLM for high-level planning to decompose a story into scene-level…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Zun Wang , Jialu Li , Han Lin , Jaehong Yoon , Mohit Bansal

In the realm of motion generation, the creation of long-duration, high-quality motion sequences remains a significant challenge. This paper presents our groundbreaking work on "Infinite Motion", a novel approach that leverages long text to…

计算机视觉与模式识别 · 计算机科学 2024-07-15 Mengtian Li , Chengshuo Zhai , Shengxiang Yao , Zhifeng Xie , Keyu Chen , Yu-Gang Jiang

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptions due to the absence of fine-grained, part-level motion…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Chuqiao Li , Xianghui Xie , Yong Cao , Andreas Geiger , Gerard Pons-Moll

The body movements accompanying speech aid speakers in expressing their ideas. Co-speech motion generation is one of the important approaches for synthesizing realistic avatars. Due to the intricate correspondence between speech and motion,…

多媒体 · 计算机科学 2024-08-28 Sen Wang , Jiangning Zhang , Xin Tan , Zhifeng Xie , Chengjie Wang , Lizhuang Ma

Recent advances in large generative models have greatly enhanced both image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Jay Zhangjie Wu , Xuanchi Ren , Tianchang Shen , Tianshi Cao , Kai He , Yifan Lu , Ruiyuan Gao , Enze Xie , Shiyi Lan , Jose M. Alvarez , Jun Gao , Sanja Fidler , Zian Wang , Huan Ling

Animating virtual characters with holistic co-speech gestures is a challenging but critical task. Previous systems have primarily focused on the weak correlation between audio and gestures, leading to physically unnatural outcomes that…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yongkang Cheng , Shaoli Huang

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Hila Chefer , Uriel Singer , Amit Zohar , Yuval Kirstain , Adam Polyak , Yaniv Taigman , Lior Wolf , Shelly Sheynin

Generative masked transformers have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Yilin Wang , Chuan Guo , Yuxuan Mu , Muhammad Gohar Javed , Xinxin Zuo , Juwei Lu , Hai Jiang , Li Cheng

As the capabilities of large language models (LLMs) continue to expand, aligning these models with human values remains a significant challenge. Recent studies show that reasoning abilities contribute significantly to model safety, while…

计算与语言 · 计算机科学 2025-06-03 Zhili Liu , Yunhao Gou , Kai Chen , Lanqing Hong , Jiahui Gao , Fei Mi , Yu Zhang , Zhenguo Li , Xin Jiang , Qun Liu , James T. Kwok

Conventional methods for human motion synthesis are either deterministic or struggle with the trade-off between motion diversity and motion quality. In response to these limitations, we introduce MoFusion, i.e., a new…

计算机视觉与模式识别 · 计算机科学 2023-05-16 Rishabh Dabral , Muhammad Hamza Mughal , Vladislav Golyanik , Christian Theobalt

Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity of inter-person interactions. We introduce the task of…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Yebin Yang , Di Wen , Lei Qi , Weitong Kong , Junwei Zheng , Ruiping Liu , Yufan Chen , Chengzhi Wu , Kailun Yang , Yuqian Fu , Danda Pani Paudel , Luc Van Gool , Kunyu Peng

Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for visual content…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Rongyao Fang , Chengqi Duan , Kun Wang , Hao Li , Hao Tian , Xingyu Zeng , Rui Zhao , Jifeng Dai , Hongsheng Li , Xihui Liu

Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizing compound actions from texts, which integrate multiple concurrent actions into…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Yue Jiang , Mingyu Yang , Liuyuxin Yang , Yang Xu , Bingxin Yun , Yuhe Zhang