中文
相关论文

相关论文: Skeleton2Stage: Reward-Guided Fine-Tuning for Phys…

200 篇论文

Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Tianyi Yan , Wencheng Han , Xia Zhou , Xueyang Zhang , Kun Zhan , Cheng-zhong Xu , Jianbing Shen

Human motion synthesis and editing are essential to many applications like film post-production. However, they often introduce artefacts in motions, which can be detrimental to the perceived realism. In particular, footskating is a frequent…

图形学 · 计算机科学 2022-08-10 Lucas Mourot , Ludovic Hoyet , François Le Clerc , Pierre Hellier

This paper presents Ske2Grid, a new representation learning framework for improved skeleton-based action recognition. In Ske2Grid, we define a regular convolution operation upon a novel grid representation of human skeleton, which is a…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Dongqi Cai , Yangyuxuan Kang , Anbang Yao , Yurong Chen

Human motion synthesis is an important task in computer graphics and computer vision. While focusing on various conditioning signals such as text, action class, or audio to guide the generation process, most existing methods utilize…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Kebing Xue , Hyewon Seo

Fine-tuning foundation models has emerged as a powerful approach for generating objects with specific desired properties. Reinforcement learning (RL) provides an effective framework for this purpose, enabling models to generate outputs that…

机器学习 · 计算机科学 2025-11-04 Pouya M. Ghari , Simone Sciabola , Ye Wang

Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely on distillation-based frameworks to compress the original multi-step denoising process…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Rui Li , Bingyu Li , Yuanzhi Liang , Haibin Huang , Chi Zhang , XueLong Li

Conventionally, supervised fine-tuning (SFT) is treated as a simple imitation learning process that only trains a policy to imitate expert behavior on demonstration datasets. In this work, we challenge this view by establishing a…

机器学习 · 计算机科学 2025-10-06 Jiangnan Li , Thuy-Trang Vu , Ehsan Abbasnejad , Gholamreza Haffari

Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting the generator with additional modules or applying…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Zhaochong An , Orest Kupyn , Théo Uscidda , Andrea Colaco , Karan Ahuja , Serge Belongie , Mar Gonzalez-Franco , Marta Tintore Gazulla

Reward-based fine-tuning of video diffusion models is an effective approach to improve the quality of generated videos, as it can fine-tune models without requiring real-world video datasets. However, it can sometimes be limited to specific…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Takehiro Aoshima , Yusuke Shinohara , Byeongseon Park

We introduce Point2Skeleton, an unsupervised method to learn skeletal representations from point clouds. Existing skeletonization methods are limited to tubular shapes and the stringent requirement of watertight input, while our method aims…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Cheng Lin , Changjian Li , Yuan Liu , Nenglun Chen , Yi-King Choi , Wenping Wang

The success of many RL techniques heavily relies on human-engineered dense rewards, which typically demand substantial domain expertise and extensive trial and error. In our work, we propose DrS (Dense reward learning from Stages), a novel…

机器学习 · 计算机科学 2024-04-26 Tongzhou Mu , Minghua Liu , Hao Su

Training a generative model on a single human skeletal motion sequence without being bound to a specific kinematic tree has drawn significant attention from the animation community. Unlike text-to-motion generation, single-shot models allow…

图形学 · 计算机科学 2025-08-27 Eleni Tselepi , Spyridon Thermos , Gerasimos Potamianos

Recognizing fine-grained actions from temporally corrupted skeleton sequences remains a significant challenge, particularly in real-world scenarios where online pose estimation often yields substantial missing data. Existing methods often…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Dian Shao , Mingfei Shi , Like Liu

Thanks to the powerful generative capacity of diffusion models, recent years have witnessed rapid progress in human motion generation. Existing diffusion-based methods employ disparate network architectures and training strategies. The…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Yiheng Huang , Hui Yang , Chuanchen Luo , Yuxi Wang , Shibiao Xu , Zhaoxiang Zhang , Man Zhang , Junran Peng

Reactive dance generation (RDG), the task of generating a dance conditioned on a lead dancer's motion, holds significant promise for enhancing human-robot interaction and immersive digital entertainment. Despite progress in duet…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Jingzhong Lin , Xinru Li , Yuanyuan Qi , Bohao Zhang , Wenxiang Liu , Kecheng Tang , Wenxuan Huang , Xiangfeng Xu , Bangyan Li , Changbo Wang , Gaoqi He

Video generation models trained on heterogeneous data with likelihood-surrogate objectives can produce visually plausible rollouts that violate physical constraints in embodied manipulation. Although reinforcement-learning post-training…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Zhenyang Ni , Yijiang Li , Ruochen Jiao , Simon Sinong Zhan , Sipeng Chen , Zhenfei Yin , Minshuo Chen , Philip Torr , Zhaoran Wang , Qi Zhu

Markerless motion capture has become an active field of research in computer vision in recent years. Its extensive applications are known in a great variety of fields, including computer animation, human motion analysis, biomedical…

计算机视觉与模式识别 · 计算机科学 2022-01-10 Doan Duy Vo , Russell Butler

Although existing 3D dance generation methods perform well in controlled scenarios, they often struggle to generalize in the wild. When conditioned on unseen music, existing methods often produce unstructured or physically implausible…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Ronghui Li , Zhongyuan Hu , Li Siyao , Youliang Zhang , Haozhe Xie , Mingyuan Zhang , Jie Guo , Xiu Li , Ziwei Liu

We propose a novel iterative approach for crossing the reality gap that utilises live robot rollouts and differentiable physics. Our method, RealityGrad, demonstrates for the first time, an efficient sim2real transfer in combination with a…

机器人学 · 计算机科学 2021-09-13 Jack Collins , Ross Brown , Jürgen Leitner , David Howard

A good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Jiaxu Zhang , Junwu Weng , Di Kang , Fang Zhao , Shaoli Huang , Xuefei Zhe , Linchao Bao , Ying Shan , Jue Wang , Zhigang Tu