English
Related papers

Related papers: Next-Scale Autoregressive Models for Text-to-Motio…

200 papers

Long-context video modeling is essential for enabling generative models to function as world simulators, as they must maintain temporal coherence over extended time spans. However, most existing models are trained on short clips, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yuchao Gu , Weijia Mao , Mike Zheng Shou

The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Shunlin Lu , Jingbo Wang , Zeyu Lu , Ling-Hao Chen , Wenxun Dai , Junting Dong , Zhiyang Dou , Bo Dai , Ruimao Zhang

Motion forecasting aims to predict the future trajectories of dynamic agents in the scene, enabling autonomous vehicles to effectively reason about scene evolution. Existing approaches operate under the closed-world regime and assume fixed…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Nicolas Schischka , Nikhil Gosala , B Ravi Kiran , Senthil Yogamani , Abhinav Valada

Feature compression is increasingly important for improving the efficiency of downstream tasks, especially in applications involving large-scale or multi-modal data. While existing methods typically rely on dedicated models for achieving…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Yufan Liu , Daoyuan Ren , Zhipeng Zhang , Wenyang Luo , Bing Li , Weiming Hu , Stephen Maybank

Large-scale pre-trained diffusion models have exhibited remarkable capabilities in diverse video generations. Given a set of video clips of the same motion concept, the task of Motion Customization is to adapt existing text-to-video…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Rui Zhao , Yuchao Gu , Jay Zhangjie Wu , David Junhao Zhang , Jiawei Liu , Weijia Wu , Jussi Keppo , Mike Zheng Shou

In this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Gwantae Kim , Youngsuk Ryu , Junyeop Lee , David K. Han , Jeongmin Bae , Hanseok Ko

Modern technological advances have enabled an unprecedented amount of structured data with complex temporal dependence, urging the need for new methods to efficiently model and forecast high-dimensional tensor-valued time series. This paper…

Methodology · Statistics 2023-09-28 Di Wang , Yao Zheng , Guodong Li

Text-driven human motion generation based on diffusion strategies establishes a reliable foundation for multimodal applications in human-computer interactions. However, existing advances face significant efficiency challenges due to the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Mengxian Hu , Minghao Zhu , Xun Zhou , Qingqing Yan , Shu Li , Chengju Liu , Qijun Chen

Autoregressive generative models of images tend to be biased towards capturing local structure, and as a result they often produce samples which are lacking in terms of large-scale coherence. To address this, we propose two methods to learn…

Computer Vision and Pattern Recognition · Computer Science 2019-10-09 Jeffrey De Fauw , Sander Dieleman , Karen Simonyan

Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challenging due to accumulated errors, motion drift, and content…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yifei Yu , Xiaoshan Wu , Xinting Hu , Tao Hu , Yangtian Sun , Xiaoyang Lyu , Bo Wang , Lin Ma , Yuewen Ma , Zhongrui Wang , Xiaojuan Qi

We introduce MotionRL, the first approach to utilize Multi-Reward Reinforcement Learning (RL) for optimizing text-to-motion generation tasks and aligning them with human preferences. Previous works focused on improving numerical performance…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Xiaoyang Liu , Yunyao Mao , Wengang Zhou , Houqiang Li

The scarcity of manipulation data has motivated the use of pretrained large models from other modalities in robotics. In this work, we build upon autoregressive video generation models to propose a Physical Autoregressive Model (PAR), where…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Zijian Song , Sihan Qin , Tianshui Chen , Liang Lin , Guangrun Wang

Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowledge of the motion…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Ekkasit Pinyoanuntapong , Muhammad Usama Saleem , Pu Wang , Minwoo Lee , Srijan Das , Chen Chen

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Learning real-world dynamics from visual observations is crucial for various domains. A common strategy is to calibrate simulators by estimating physical parameters, yet accuracy is ultimately bounded by the underlying physical models,…

Machine Learning · Computer Science 2026-05-22 Jiaxu Wang , Junhao He , Jingkai Sun , Yi Gu , Yunyang Mo , Jiahang Cao , Qiang Zhang , Renjing Xu

As an essential task in autonomous driving (AD), motion prediction aims to predict the future states of surround objects for navigation. One natural solution is to estimate the position of other agents in a step-by-step manner where each…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Xiaosong Jia , Shaoshuai Shi , Zijun Chen , Li Jiang , Wenlong Liao , Tao He , Junchi Yan

Non-autoregressive (NAR) neural machine translation is usually done via knowledge distillation from an autoregressive (AR) model. Under this framework, we leverage large monolingual corpora to improve the NAR model's performance, with the…

Computation and Language · Computer Science 2020-12-01 Jiawei Zhou , Phillip Keung

In real world domains, most graphs naturally exhibit a hierarchical structure. However, data-driven graph generation is yet to effectively capture such structures. To address this, we propose a novel approach that recursively generates…

Machine Learning · Computer Science 2023-06-01 Mahdi Karami , Jun Luo

Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce…

Computation and Language · Computer Science 2025-03-13 Eric Battenberg , RJ Skerry-Ryan , Daisy Stanton , Soroosh Mariooryad , Matt Shannon , Julian Salazar , David Kao

Class-conditional generative models have emerged as accurate and robust classifiers, with diffusion models demonstrating clear advantages over other visual generative paradigms, including autoregressive (AR) models. In this work, we revisit…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Ilia Sudakov , Artem Babenko , Dmitry Baranchuk