English
Related papers

Related papers: DisMo: Disentangled Motion Representations for Ope…

200 papers

Text-to-motion (T2M) generation with diffusion backbones achieves strong realism and alignment. Safety concerns in T2M methods have been raised in recent years; existing methods replace discrete VQ-VAE codebook entries to steer the model…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Yiling Wang , Zeyu Zhang , Yiran Wang , Hao Tang

Modeling virtual agents with behavior style is one factor for personalizing human agent interaction. We propose an efficient yet effective machine learning approach to synthesize gestures driven by prosodic features and text in the style of…

Sound · Computer Science 2022-08-04 Mireille Fares , Michele Grimaldi , Catherine Pelachaud , Nicolas Obin

In recent years, large-scale pre-trained diffusion transformer models have made significant progress in video generation. While current DiT models can produce high-definition, high-frame-rate, and highly diverse videos, there is a lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Changgu Chen , Xiaoyan Yang , Junwei Shu , Changbo Wang , Yang Li

We present a new approach for representing and reconstructing multidimensional magnetic resonance imaging (MRI) data. Our method builds on a novel, learned feature-based image representation that disentangles different types of features,…

Image and Video Processing · Electrical Eng. & Systems 2026-01-01 Ruiyang Zhao , Fan Lam

Motion transfer is the task of synthesizing future video frames of a single source image according to the motion from a given driving video. In order to solve it, we face the challenging complexity of motion representation and the unknown…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Or Toledano , Yanir Marmor , Dov Gertz

This paper introduces a SSSUMO, semi-supervised deep learning approach for submovement decomposition that achieves state-of-the-art accuracy and speed. While submovement analysis offers valuable insights into motor control, existing methods…

Human-Computer Interaction · Computer Science 2025-07-14 Evgenii Rudakov , Jonathan Shock , Otto Lappi , Benjamin Ultan Cowley

Multimodal representations that enable cross-modal retrieval are widely used. However, these often lack interpretability making it difficult to explain the retrieved results. Solutions such as learning sparse disentangled representations…

Information Retrieval · Computer Science 2025-06-25 Prachi J , Sumit Bhatia , Srikanta Bedathur

This work addresses motion-guided few-shot video object segmentation (FSVOS), which aims to segment dynamic objects in videos based on a few annotated examples with the same motion patterns. Existing FSVOS datasets and methods typically…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Kaining Ying , Hengrui Hu , Henghui Ding

Disentangled representations support a range of downstream tasks including causal reasoning, generative modeling, and fair machine learning. Unfortunately, disentanglement has been shown to be impossible without the incorporation of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-16 Matthew J. Vowels , Necati Cihan Camgoz , Richard Bowden

Video generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Zeqi Xiao , Yifan Zhou , Shuai Yang , Xingang Pan

We present X-UniMotion, a unified and expressive implicit latent representation for whole-body human motion, encompassing facial expressions, body poses, and hand gestures. Unlike prior motion transfer methods that rely on explicit skeletal…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Guoxian Song , Hongyi Xu , Xiaochen Zhao , You Xie , Tianpei Gu , Zenan Li , Chenxu Zhang , Linjie Luo

Text-to-Image (T2I) models excel at synthesizing concepts such as nouns, appearances, and styles. To enable customized content creation based on a few example images of a concept, methods such as Textual Inversion and DreamBooth invert the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 Saman Motamed , Danda Pani Paudel , Luc Van Gool

We present Pix2Gif, a motion-guided diffusion model for image-to-GIF (video) generation. We tackle this problem differently by formulating the task as an image translation problem steered by text and motion magnitude prompts, as shown in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Hitesh Kandala , Jianfeng Gao , Jianwei Yang

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Yixuan Ren , Yang Zhou , Jimei Yang , Jing Shi , Difan Liu , Feng Liu , Mingi Kwon , Abhinav Shrivastava

Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing paradigm of encoding raw pixels into opaque latent spaces and relying on heavy decoders for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Roussel Desmond Nzoyem , Mauro Comi

Recent studies on motion estimation have advocated an optimized motion representation that is globally consistent across the entire video, preferably for every pixel. This is challenging as a uniform representation may not account for the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Rui Li , Dong Liu

Diffusion models generate images with an unprecedented level of quality, but how can we freely rearrange image layouts? Recent works generate controllable scenes via learning spatially disentangled latent codes, but these methods do not…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Jiawei Ren , Mengmeng Xu , Jui-Chieh Wu , Ziwei Liu , Tao Xiang , Antoine Toisoul

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training data. By leveraging a few video samples demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Joanna Materzynska , Josef Sivic , Eli Shechtman , Antonio Torralba , Richard Zhang , Bryan Russell

The technology for Visual Odometry (VO) that estimates the position and orientation of the moving object through analyzing the image sequences captured by on-board cameras, has been well investigated with the rising interest in autonomous…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Ran Zhu , Mingkun Yang , Wang Liu , Rujun Song , Bo Yan , Zhuoling Xiao

Motion retargeting is the long-standing problem in character animation that consists in transferring and adapting the motion of a source character to another target character. A typical application is the creation of motion sequences from…

Graphics · Computer Science 2023-06-16 Lucas Mourot , Ludovic Hoyet , François Le Clerc , Pierre Hellier