English
Related papers

Related papers: EchoMotion: Unified Human Video and Motion Generat…

200 papers

Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when generating intricate,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Xiaomin Li , Xu Jia , Qinghe Wang , Haiwen Diao , Mengmeng Ge , Pengxiang Li , You He , Huchuan Lu

Diffusion models usher a new era of video editing, flexibly manipulating the video contents with text prompts. Despite the widespread application demand in editing human-centered videos, these models face significant challenges in handling…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Xiaojing Zhong , Xinyi Huang , Xiaofeng Yang , Guosheng Lin , Qingyao Wu

Recent advancements utilizing large-scale video data for learning video generation models demonstrate significant potential in understanding complex physical dynamics. It suggests the feasibility of leveraging diverse robot trajectory data…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Youpeng Wen , Junfan Lin , Yi Zhu , Jianhua Han , Hang Xu , Shen Zhao , Xiaodan Liang

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Sohan Anisetty , James Hays

Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames. Existing adaptations,…

Artificial Intelligence · Computer Science 2025-11-03 Zikang Liu , Longteng Guo , Yepeng Tang , Tongtian Yue , Junxian Cai , Kai Ma , Qingbin Liu , Xi Chen , Jing Liu

Text-to-motion generation, which translates textual descriptions into human motions, faces the challenge that users often struggle to precisely convey their intended motions through text alone. To address this issue, this paper introduces…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Tao Wang , Lei Jin , Zhihua Wu , Qiaozhi He , Jiaming Chu , Yu Cheng , Junliang Xing , Jian Zhao , Shuicheng Yan , Li Wang

Autoregressive diffusion enables real-time frame streaming, yet existing sliding-window caches discard past context, causing fidelity degradation, identity drift, and motion stagnation over long horizons. Current approaches preserve a fixed…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Youngrae Kim , Qixin Hu , C. -C. Jay Kuo , Peter A. Beerel

Diffusion models have shown impressive performance in many visual generation and manipulation tasks. Many existing methods focus on training a model for a specific task, especially, text-to-video (T2V) generation, while many other works…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Ruibin Li , Tao Yang , Yangming Shi , Weiguo Feng , Shilei Wen , Bingyue Peng , Lei Zhang

Video-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Moritz Kappel , Vladislav Golyanik , Mohamed Elgharib , Jann-Ole Henningson , Hans-Peter Seidel , Susana Castillo , Christian Theobalt , Marcus Magnor

Lightweight, controllable, and physically plausible human motion synthesis is crucial for animation, virtual reality, robotics, and human-computer interaction applications. Existing methods often compromise between computational efficiency,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Arvin Tashakori , Arash Tashakori , Gongbo Yang , Z. Jane Wang , Peyman Servati

Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datasets, resulting in a lack of visual grounding and restricting…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Yuwei Fang , Willi Menapace , Aliaksandr Siarohin , Tsai-Shien Chen , Kuan-Chien Wang , Ivan Skorokhodov , Graham Neubig , Sergey Tulyakov

Long-term human motion can be represented as a series of motion modes---motion sequences that capture short-term temporal dynamics---with transitions between them. We leverage this structure and present a novel Motion Transformation…

Machine Learning · Computer Science 2018-08-15 Xinchen Yan , Akash Rastogi , Ruben Villegas , Kalyan Sunkavalli , Eli Shechtman , Sunil Hadap , Ersin Yumer , Honglak Lee

We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Yuxin Mao , Jing Zhang , Mochu Xiang , Yiran Zhong , Yuchao Dai

This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autonomous driving. Due to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Jiacheng Lin , Jiajun Chen , Kunyu Peng , Xuan He , Zhiyong Li , Rainer Stiefelhagen , Kailun Yang

This study investigates the use of large language models (LLMs) for human behavior understanding by jointly leveraging motion and video data. We argue that integrating these complementary modalities is essential for capturing both…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Rajan Das Gupta , Lei Wei , Md Yeasin Rahat , Nafiz Fahad , Abir Ahmed , Liew Tze Hui

Despite recent advancements in learning-based motion in-betweening, a key limitation has been overlooked: the requirement for character-specific datasets. In this work, we introduce AnyMoLe, a novel method that addresses this limitation by…

Graphics · Computer Science 2025-03-12 Kwan Yun , Seokhyeon Hong , Chaelin Kim , Junyong Noh

Monocular 3D human performance capture is indispensable for many applications in computer graphics and vision for enabling immersive experiences. However, detailed capture of humans requires tracking of multiple aspects, including the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Yue Jiang , Marc Habermann , Vladislav Golyanik , Christian Theobalt

Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for direct execution on a robot. Still, these demonstrations convey…

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Liangdong Qiu , Chengxing Yu , Yanran Li , Zhao Wang , Haibin Huang , Chongyang Ma , Di Zhang , Pengfei Wan , Xiaoguang Han

Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled…

Robotics · Computer Science 2026-05-06 Zhiyuan Li , Wenyan Yang , Wenshuai Zhao , Yue Ma , Yuanpeng Tu , Pekka Marttinen , Joni Pajarinen