中文
相关论文

相关论文: Go to Zero: Towards Zero-shot Motion Generation wi…

200 篇论文

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Shijie Wang , Samaneh Azadi , Rohit Girdhar , Saketh Rambhatla , Chen Sun , Xi Yin

Autoregressive (AR) models offer stable and efficient training, but standard next-token prediction is not well aligned with the temporal structure required for text-conditioned motion generation. We introduce MoScale, a next-scale AR…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Zhiwei Zheng , Shibo Jin , Lingjie Liu , Mingmin Zhao

The success of monocular depth estimation relies on large and diverse training sets. Due to the challenges associated with acquiring dense ground-truth depth across different environments at scale, a number of datasets with distinct…

计算机视觉与模式识别 · 计算机科学 2020-08-26 René Ranftl , Katrin Lasinger , David Hafner , Konrad Schindler , Vladlen Koltun

Zero-shot skeleton-based action recognition aims to recognize actions of unseen categories after training on data of seen categories. The key is to build the connection between visual and semantic space from seen to unseen classes. Previous…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Yujie Zhou , Wenwen Qiang , Anyi Rao , Ning Lin , Bing Su , Jiaqi Wang

Motion understanding is fundamental to physical reasoning, enabling models to infer dynamics and predict future states. However, state-of-the-art models still struggle on recent motion benchmarks, primarily due to the scarcity of…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Yulu Gan , Ligeng Zhu , Dandan Shan , Baifeng Shi , Hongxu Yin , Boris Ivanovic , Song Han , Trevor Darrell , Jitendra Malik , Marco Pavone , Boyi Li

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Boyuan Wang , Xiaofeng Wang , Chaojun Ni , Guosheng Zhao , Zhiqin Yang , Zheng Zhu , Muyang Zhang , Yukun Zhou , Xinze Chen , Guan Huang , Lihong Liu , Xingang Wang

Despite significant advances in quadrupedal robotics, a critical gap persists in foundational motion resources that holistically integrate diverse locomotion, emotionally expressive behaviors, and rich language semantics-essential for…

机器人学 · 计算机科学 2026-03-26 Li Gao , Fuzhi Yang , Jianhui Chen , Liu Liu , Yao Zheng , Yang Cai , Ziqiao Li

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Liangdong Qiu , Chengxing Yu , Yanran Li , Zhao Wang , Haibin Huang , Chongyang Ma , Di Zhang , Pengfei Wan , Xiaoguang Han

Zero-shot action recognition (ZSAR) aims to learn an alignment model between videos and class descriptions of seen actions that is transferable to unseen actions. The text queries (class descriptions) used in existing ZSAR works, however,…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Jiaming Zhou , Junwei Liang , Kun-Yu Lin , Jinrui Yang , Wei-Shi Zheng

Due to recent advances in pose-estimation methods, human motion can be extracted from a common video in the form of 3D skeleton sequences. Despite wonderful application opportunities, effective and efficient content-based access to large…

计算机视觉与模式识别 · 计算机科学 2023-10-05 Nicola Messina , Jan Sedmidubsky , Fabrizio Falchi , Tomáš Rebok

Motion time series collected from mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR due to their…

信号处理 · 电气工程与系统科学 2024-10-29 Xiyuan Zhang , Diyan Teng , Ranak Roy Chowdhury , Shuheng Li , Dezhi Hong , Rajesh K. Gupta , Jingbo Shang

Human-human motion generation is essential for understanding humans as social beings. Current methods fall into two main categories: single-person-based methods and separate modeling-based methods. To delve into this field, we abstract the…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Yabiao Wang , Shuo Wang , Jiangning Zhang , Ke Fan , Jiafu Wu , Zhucun Xue , Yong Liu

Natural and expressive human motion generation is the holy grail of computer animation. It is a challenging task, due to the diversity of possible motion, human perceptual sensitivity to it, and the difficulty of accurately describing it.…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Guy Tevet , Sigal Raab , Brian Gordon , Yonatan Shafir , Daniel Cohen-Or , Amit H. Bermano

This paper presents a method that allows users to design cinematic video shots in the context of image-to-video generation. Shot design, a critical aspect of filmmaking, involves meticulously planning both camera movements and object…

计算机视觉与模式识别 · 计算机科学 2025-02-07 Jinbo Xing , Long Mai , Cusuh Ham , Jiahui Huang , Aniruddha Mahapatra , Chi-Wing Fu , Tien-Tsin Wong , Feng Liu

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Bizhu Wu , Jinheng Xie , Keming Shen , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Human motion synthesis is a fundamental task in computer animation. Despite recent progress in this field utilizing deep learning and motion capture data, existing methods are always limited to specific motion categories, environments, and…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Zhikai Zhang , Yitang Li , Haofeng Huang , Mingxian Lin , Li Yi

Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability in the gestures…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Hassan Ali , Doreen Jirak , Luca Müller , Stefan Wermter

Human motion synthesis in complex scenes presents a fundamental challenge, extending beyond conventional Text-to-Motion tasks by requiring the integration of diverse modalities such as static environments, movable objects, natural language…

图形学 · 计算机科学 2025-05-20 Zichen Geng , Zeeshan Hayder , Wei Liu , Ajmal Mian

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over generated videos.…