中文
相关论文

相关论文: Multi-scale Coarse-to-fine Modeling for Test-time …

200 篇论文

Text-conditioned human motion generation, which allows for user interaction through natural language, has become increasingly popular. Existing methods typically generate short, isolated motions based on a single input sentence. However,…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Kaifeng Zhao , Gen Li , Siyu Tang

Human motion prediction is a necessary component for many applications in robotics and autonomous driving. Recent methods propose using sequence-to-sequence deep learning models to tackle this problem. However, they do not focus on…

计算机视觉与模式识别 · 计算机科学 2020-10-08 Tim Lebailly , Sena Kiciroglu , Mathieu Salzmann , Pascal Fua , Wei Wang

State-of-the-art Text-To-Speech (TTS) models are capable of producing high-quality speech. The generated speech, however, is usually neutral in emotional expression, whereas very often one would want fine-grained emotional control of words…

声音 · 计算机科学 2023-03-14 Shijun Wang , Jón Guðnason , Damian Borth

Despite recent progress, most existing virtual try-on methods still struggle to simultaneously address two core challenges: accurately aligning the garment image with the target human body, and preserving fine-grained garment textures and…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Xianbing Sun , Yan Hong , Jiahui Zhan , Jun Lan , Huijia Zhu , Weiqiang Wang , Liqing Zhang , Jianfu Zhang

We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yannan He , Garvita Tiwari , Xiaohan Zhang , Pankaj Bora , Tolga Birdal , Jan Eric Lenssen , Gerard Pons-Moll

Chain-of-Thought (CoT) reasoning improves multi-step mathematical problem solving in large language models but remains vulnerable to exposure bias and error accumulation, as early mistakes propagate irreversibly through autoregressive…

计算与语言 · 计算机科学 2026-04-21 Shidong Cao , Hongzhan Lin , Yuxuan Gu , Ziyang Luo , Jing Ma

Discrete masked diffusion language models such as LLaDA generate text through iterative denoising, where mask tokens are progressively replaced with predicted tokens. LLaDA2.1 introduced a Token-to-Token (T2T) editing mechanism that…

计算与语言 · 计算机科学 2026-05-27 Lin Yao

Chain-of-Thought (CoT) reasoning successfully enhances the reasoning capabilities of Large Language Models (LLMs), yet it incurs substantial computational overhead for inference. Existing CoT compression methods often suffer from a critical…

In this paper, we propose a cost-matching approach for optimal humanoid locomotion within a Model Predictive Control (MPC)-based Reinforcement Learning (RL) framework. A parameterized MPC formulation with centroidal dynamics is trained to…

机器人学 · 计算机科学 2026-03-31 Wenqi Cai , Kyriakos G. Vamvoudakis , Sébastien Gros , Anthony Tzes

We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence conditioned on an input image. CoMoGen introduces a lightweight MaskAdapter that encodes binary…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Adil Meric , Lin Geng Foo , Mert Kiray , Benjamin Busam , Rishabh Dabral , Christian Theobalt

Text-to-motion generation is driven by learning motion representations for semantic alignment with language. Existing methods rely on either continuous or discrete motion representations. However, continuous representations entangle…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Dawei Guan , Di Yang , Chengjie Jin , Jiangtao Wang

Stochastic human motion prediction aims to generate diverse, plausible futures from observed sequences. Despite advances in generative modeling, existing methods often produce predictions corrupted by high-frequency jitter and temporal…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Wenhan Wu , Zhishuai Guo , Chen Chen , Srijan Das , Hongfei Xue , Pu Wang , Aidong Lu

Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their interactions remains challenging. Existing methods often struggle…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Pengxiang Cai , Mengyang Li

Existing multimodal conditional image synthesis (MCIS) methods generate images conditioned on any combinations of various modalities that require all of them must be exactly conformed, hindering the synthesis controllability and leaving the…

计算机视觉与模式识别 · 计算机科学 2023-05-11 Jianbin Zheng , Daqing Liu , Chaoyue Wang , Minghui Hu , Zuopeng Yang , Changxing Ding , Dacheng Tao

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle…

An adaptive discretization refinement strategy for steady state discrete mesoscale models of coupled mechanics and mass transport in concrete is presented. Coupling is provided by two phenomena: the Biot's theory of poromechanics and an…

计算工程、金融与科学 · 计算机科学 2023-07-03 Jan Mašek , Josef Květon , Jan Eliáš

Trajectory-controlled human motion generation aims to synthesize realistic human motions conditioned on both textual descriptions and spatial trajectories. However, existing methods suffer from two critical limitations: first, the conflict…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Deli Cai , Haoyang Ma , Changxing Ding

Recent progress in text-to-motion has advanced both 3D human motion generation and text-based motion control. Controllable motion generation (CoMo), which enables intuitive control, typically relies on pose code representations, but…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Sukhyun Jeong , Hong-Gi Shin , Yong-Hoon Choi

The embodied learning of human motor control requires whole-body neuro-actuated musculoskeletal dynamics, while the internal muscle-driven processes underlying movement remain inaccessible to direct measurement. Computational modeling…

机器人学 · 计算机科学 2026-04-01 Yunyue Wei , Chenhui Zuo , Shanning Zhuang , Haixin Gong , Yaming Liu , Yanan Sui

Recently, diffusion models (DMs) have made significant strides in high-quality image generation. However, the multi-step denoising process often results in considerable computational overhead, impeding deployment on resource-constrained…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Yu-Shan Tai , An-Yeu , Wu