中文
相关论文

相关论文: TM2D: Bimodality Driven 3D Dance Generation via Mu…

200 篇论文

In this work we present a novel, robust transition generation technique that can serve as a new tool for 3D animators, based on adversarial recurrent neural networks. The system synthesizes high-quality motions that use temporally-sparse…

计算机视觉与模式识别 · 计算机科学 2021-02-10 Félix G. Harvey , Mike Yurick , Derek Nowrouzezahrai , Christopher Pal

We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced generative models. Our approach combines the precision of a…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Zhiyuan Zhang , Dongdong Chen , Jing Liao

Text-to-motion generation, which translates textual descriptions into human motions, faces the challenge that users often struggle to precisely convey their intended motions through text alone. To address this issue, this paper introduces…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Tao Wang , Lei Jin , Zhihua Wu , Qiaozhi He , Jiaming Chu , Yu Cheng , Junliang Xing , Jian Zhao , Shuicheng Yan , Li Wang

In the realm of motion generation, the creation of long-duration, high-quality motion sequences remains a significant challenge. This paper presents our groundbreaking work on "Infinite Motion", a novel approach that leverages long text to…

计算机视觉与模式识别 · 计算机科学 2024-07-15 Mengtian Li , Chengshuo Zhai , Shengxiang Yao , Zhifeng Xie , Keyu Chen , Yu-Gang Jiang

Text-guided motion synthesis aims to generate 3D human motion that not only precisely reflects the textual description but reveals the motion details as much as possible. Pioneering methods explore the diffusion model for text-to-motion…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Zhenyu Xie , Yang Wu , Xuehao Gao , Zhongqian Sun , Wei Yang , Xiaodan Liang

Humans perform a variety of interactive motions, among which duet dance is one of the most challenging interactions. However, in terms of human motion generative models, existing works are still unable to generate high-quality interactive…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Ronghui Li , Youliang Zhang , Yachao Zhang , Yuxiang Zhang , Mingyang Su , Jie Guo , Ziwei Liu , Yebin Liu , Xiu Li

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the low-level actions…

机器人学 · 计算机科学 2026-01-01 Karthik Dharmarajan , Wenlong Huang , Jiajun Wu , Li Fei-Fei , Ruohan Zhang

We present X-Dancer, a novel zero-shot music-driven image animation pipeline that creates diverse and long-range lifelike human dance videos from a single static image. As its core, we introduce a unified transformer-diffusion framework,…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Zeyuan Chen , Hongyi Xu , Guoxian Song , You Xie , Chenxu Zhang , Xin Chen , Chao Wang , Di Chang , Linjie Luo

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Text-to-motion generation has gained increasing attention, but most existing methods are limited to generating short-term motions that correspond to a single sentence describing a single action. However, when a text stream describes a…

计算机视觉与模式识别 · 计算机科学 2023-08-04 Zhao Yang , Bing Su , Ji-Rong Wen

This work addresses the problem of generating 3D holistic body motions from human speech. Given a speech recording, we synthesize sequences of 3D body poses, hand gestures, and facial expressions that are realistic and diverse. To achieve…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Hongwei Yi , Hualin Liang , Yifei Liu , Qiong Cao , Yandong Wen , Timo Bolkart , Dacheng Tao , Michael J. Black

Text-based video segmentation aims to segment the target object in a video based on a describing sentence. Incorporating motion information from optical flow maps with appearance and linguistic modalities is crucial yet has been largely…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Wangbo Zhao , Kai Wang , Xiangxiang Chu , Fuzhao Xue , Xinchao Wang , Yang You

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Liangdong Qiu , Chengxing Yu , Yanran Li , Zhao Wang , Haibin Huang , Chongyang Ma , Di Zhang , Pengfei Wan , Xiaoguang Han

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang

Recent breakthroughs in Vision-Language (V&L) joint research have achieved remarkable results in various text-driven tasks. High-quality Text-to-video (T2V), a task that has been long considered mission-impossible, was proven feasible with…

人工智能 · 计算机科学 2022-11-28 Yuxing Qiu , Feng Gao , Minchen Li , Govind Thattai , Yin Yang , Chenfanfu Jiang

Synthesizing natural human motions that enable a 3D human avatar to walk and reach for arbitrary goals in 3D space remains an unsolved problem with many applications. Existing methods (data-driven or using reinforcement learning) are…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Markos Diomataris , Nikos Athanasiou , Omid Taheri , Xi Wang , Otmar Hilliges , Michael J. Black

Diffusion-based video generation can create realistic videos, yet existing image- and text-based conditioning fails to offer precise motion control. Prior methods for motion-conditioned synthesis typically require model-specific…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Assaf Singer , Noam Rotstein , Amir Mann , Ron Kimmel , Or Litany

We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 5.2 hours…

计算机视觉与模式识别 · 计算机科学 2021-08-03 Ruilong Li , Shan Yang , David A. Ross , Angjoo Kanazawa

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Uriel Singer , Adam Polyak , Thomas Hayes , Xi Yin , Jie An , Songyang Zhang , Qiyuan Hu , Harry Yang , Oron Ashual , Oran Gafni , Devi Parikh , Sonal Gupta , Yaniv Taigman
‹ 上一页 1 8 9 10 下一页 ›