English
Related papers

Related papers: Bridging Semantic and Kinematic Conditions with Di…

200 papers

Graph neural networks (GNNs) have revolutionized recommender systems by effectively modeling complex user-item interactions, yet data sparsity and the item cold-start problem significantly impair performance, particularly for new items with…

Machine Learning · Computer Science 2026-03-04 Jialin Liu , Zhaorui Zhang , Ray C. C. Cheung

Video diffusion models achieve strong frame-level fidelity but still struggle with motion coherence, dynamics and realism, often producing jitter, ghosting, or implausible dynamics. A key limitation is that the standard denoising MSE…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Haotian Xue , Qi Chen , Zhonghao Wang , Xun Huang , Eli Shechtman , Jinrong Xie , Yongxin Chen

Text-to-Motion (T2M) generation aims to synthesize realistic and semantically aligned human motion sequences from natural language descriptions. However, current approaches face dual challenges: Generative models (e.g., diffusion models)…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Zhengdao Li , Siheng Wang , Zeyu Zhang , Hao Tang

Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapolate to unseen physical conditions (eg, velocity) due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Wang Lin , Liyu Jia , Wentao Hu , Kaihang Pan , Zhongqi Yue , Wei Zhao , Jingyuan Chen , Fei Wu , Hanwang Zhang

Recent advances in motion diffusion models have enabled spatially controllable text-to-motion generation. However, these models struggle to achieve high-precision control while maintaining high-quality motion generation. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Ekkasit Pinyoanuntapong , Muhammad Usama Saleem , Korrawe Karunratanakul , Pu Wang , Hongfei Xue , Chen Chen , Chuan Guo , Junli Cao , Jian Ren , Sergey Tulyakov

Recent advances in text-to-motion generation using diffusion and autoregressive models have shown promising results. However, these models often suffer from a trade-off between real-time performance, high fidelity, and motion editability.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Ekkasit Pinyoanuntapong , Pu Wang , Minwoo Lee , Chen Chen

We present Diffuse-CLoC, a guided diffusion framework for physics-based look-ahead control that enables intuitive, steerable, and physically realistic motion generation. While existing kinematics motion generation with diffusion models…

Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectively capture the temporal dependencies and redundancies…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Xiang Fan , Xiaohang Sun , Kushan Thakkar , Zhu Liu , Vimal Bhat , Ranjay Krishna , Xiang Hao

Modeling realistic and interactive multi-agent behavior is critical to autonomous driving and traffic simulation. However, existing diffusion and autoregressive approaches are limited by iterative sampling, sequential decoding, or…

Robotics · Computer Science 2025-11-24 Zhiyu Huang , Zewei Zhou , Tianhui Cai , Yun Zhang , Jiaqi Ma

Text-to-motion generation is driven by learning motion representations for semantic alignment with language. Existing methods rely on either continuous or discrete motion representations. However, continuous representations entangle…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Dawei Guan , Di Yang , Chengjie Jin , Jiangtao Wang

Image tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generation. However, mainstream image tokenization methods generally…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Kaiwen Zha , Lijun Yu , Alireza Fathi , David A. Ross , Cordelia Schmid , Dina Katabi , Xiuye Gu

The development of foundation models for functional magnetic resonance imaging (fMRI) time series holds significant promise for predicting phenotypes related to disease and cognition. Current models, however, are often trained using a…

Machine Learning · Computer Science 2026-03-03 Sam Gijsen , Marc-Andre Schulz , Kerstin Ritter

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Canxuan Gang

This paper presents the \textbf{S}emantic-a\textbf{W}ar\textbf{E} spatial-t\textbf{E}mporal \textbf{T}okenizer (SweetTok), a novel video tokenizer to overcome the limitations in current video tokenization methods for compacted yet effective…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zhentao Tan , Ben Xue , Jian Jia , Junhao Wang , Wencai Ye , Shaoyun Shi , Mingjie Sun , Wenjin Wu , Quan Chen , Peng Jiang

Recent advances in diffusion models have endowed talking head synthesis with subtle expressions and vivid head movements, but have also led to slow inference speed and insufficient control over generated results. To address these issues, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Tianqi Li , Ruobing Zheng , Minghui Yang , Jingdong Chen , Ming Yang

Generating human motion guided by conditions such as textual descriptions is challenging due to the need for datasets with pairs of high-quality motion and their corresponding conditions. The difficulty increases when aiming for finer…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Pablo Ruiz-Ponce , German Barquero , Cristina Palmero , Sergio Escalera , José García-Rodríguez

Acquiring aligned visuo-tactile datasets is slow and costly, requiring specialised hardware and large-scale data collection. Synthetic generation is promising, but prior methods are typically single-modality, limiting cross-modal learning.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Sirine Bhouri , Lan Wei , Jian-Qing Zheng , Dandan Zhang

Autoregressive (AR) language models rely on causal tokenization, but extending this paradigm to vision remains non-trivial. Current visual tokenizers either flatten 2D patches into non-causal sequences or enforce heuristic orderings that…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Yitong Chen , Zuxuan Wu , Xipeng Qiu , Yu-Gang Jiang

Diffusion models have emerged as a promising approach for text generation, with recent works falling into two main categories: discrete and continuous diffusion models. Discrete diffusion models apply token corruption independently using…

Computation and Language · Computer Science 2025-05-29 Bocheng Li , Zhujin Gao , Linli Xu

Image tokenizers form the foundation of modern text-to-image generative models but are notoriously difficult to train. Furthermore, most existing text-to-image models rely on large-scale, high-quality private datasets, making them…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Dongwon Kim , Ju He , Qihang Yu , Chenglin Yang , Xiaohui Shen , Suha Kwak , Liang-Chieh Chen