中文
相关论文

相关论文: Multi-scale Coarse-to-fine Modeling for Test-time …

200 篇论文

The objective of the multi-condition human motion synthesis task is to incorporate diverse conditional inputs, encompassing various forms like text, music, speech, and more. This endows the task with the capability to adapt across multiple…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Zeyu Ling , Bo Han , Yongkang Wong , Mohan Kangkanhalli , Weidong Geng

Text-driven human motion generation based on diffusion strategies establishes a reliable foundation for multimodal applications in human-computer interactions. However, existing advances face significant efficiency challenges due to the…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Mengxian Hu , Minghao Zhu , Xun Zhou , Qingqing Yan , Shu Li , Chengju Liu , Qijun Chen

Recent advancements in large-scale models have showcased remarkable generalization capabilities in various tasks. However, integrating multimodal processing into these models presents a significant challenge, as it often comes with a high…

多媒体 · 计算机科学 2024-07-17 Hao Sun , Yu Song , Xinyao Yu , Jiaqing Liu , Yen-Wei Chen , Lanfen Lin

Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Ruineng Li , Daitao Xing , Huiming Sun , Yuanzhou Ha , Jinglin Shen , Chiuman Ho

As large-scale language model pretraining pushes the state-of-the-art in text generation, recent work has turned to controlling attributes of the text such models generate. While modifying the pretrained models via fine-tuning remains the…

计算与语言 · 计算机科学 2021-08-05 Sachin Kumar , Eric Malmi , Aliaksei Severyn , Yulia Tsvetkov

For dynamic human motion sequences, the original KeyNode-Driven codec often struggles to retain compression efficiency when confronted with rapid movements or strong non-rigid deformations. This paper proposes a novel Bi-modal coding…

信号处理 · 电气工程与系统科学 2025-09-23 Huong Hoang , Keito Suzuki , Truong Nguyen , Pamela Cosman

The compression of real-world scanned 3D human dynamic meshes is an emerging research area, driven by applications such as telepresence, virtual reality, and 3D digital streaming. Unlike synthesized dynamic meshes with fixed topology,…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Huong Hoang , Truong Nguyen , Pamela Cosman

The study of phenomena such as protein folding and conformational changes in molecules is a central theme in chemical physics. Molecular dynamics (MD) simulation is the primary tool for the study of transition processes in biomolecules, but…

计算物理 · 物理学 2022-12-14 Luke Evans , Maria K. Cameron , Pratyush Tiwary

Recent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate global 3D coordinate…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Keming Shen , Bizhu Wu , Junliang Chen , Xiaoqin Wang , Linlin Shen

Measuring the in-context computational effort of language models is a key challenge, as metrics like next-token loss fail to capture reasoning complexity. Prior methods based on latent state compressibility can be invasive and unstable. We…

机器学习 · 计算机科学 2025-12-30 Vincent Herrmann , Eric Alcaide , Michael Wand , Jürgen Schmidhuber

High-quality human motion data is becoming increasingly important for applications in robotics, simulation, and entertainment. Recent generative models offer a potential data source, enabling human motion synthesis through intuitive inputs…

Data-driven and controllable human motion synthesis and prediction are active research areas with various applications in interactive media and social robotics. Challenges remain in these fields for generating diverse motions given past…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Wenjie Yin , Ruibo Tu , Hang Yin , Danica Kragic , Hedvig Kjellström , Mårten Björkman

Accurate representation of the multiscale features in spatiotemporal physical systems using vision transformer (ViT) architectures requires extremely long, computationally prohibitive token sequences. To address this issue, we propose two…

机器学习 · 计算机科学 2024-12-31 Pei Zhang , M. Paul Laiu , Matthew Norman , Doug Stefanski , John Gounley

Irregularly sampled multivariate time series (ISMTS) are prevalent in reality. Most existing methods treat ISMTS as synchronized regularly sampled time series with missing values, neglecting that the irregularities are primarily attributed…

机器学习 · 计算机科学 2024-12-03 Jiexi Liu , Meng Cao , Songcan Chen

Stochastic human motion prediction aims to forecast multiple plausible future motions given a single pose sequence from the past. Most previous works focus on designing elaborate losses to improve the accuracy, while the diversity is…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Dong Wei , Huaijiang Sun , Bin Li , Jianfeng Lu , Weiqing Li , Xiaoning Sun , Shengxiang Hu

The optimal control of a mechanical system is of crucial importance in many realms. Typical examples are the determination of a time-minimal path in vehicle dynamics, a minimal energy trajectory in space mission design, or optimal motion…

最优化与控制 · 数学 2008-10-09 S. Ober-Bloebaum , O. Junge , J. E. Marsden

With the increasing availability of diverse data types, particularly images and time series data from medical experiments, there is a growing demand for techniques designed to combine various modalities of data effectively. Our motivation…

图像与视频处理 · 电气工程与系统科学 2024-05-27 Ali Rasekh , Reza Heidari , Amir Hosein Haji Mohammad Rezaie , Parsa Sharifi Sedeh , Zahra Ahmadi , Prasenjit Mitra , Wolfgang Nejdl

We propose a generative framework for multi-track music source separation (MSS) that reformulates the task as conditional discrete token generation. Unlike conventional approaches that directly estimate continuous signals in the time or…

音频与语音处理 · 电气工程与系统科学 2026-04-20 Pengbo Lyu , Xiangyu Zhao , Chengwei Liu , Haoyin Yan , Xiaotao Liang , Hongyu Wang , Shaofei Xue

We study generalizable policy learning from demonstrations for complex low-level control (e.g., contact-rich object manipulations). We propose a novel hierarchical imitation learning method that utilizes sub-optimal demos. Firstly, we…

机器学习 · 计算机科学 2024-07-09 Zhiwei Jia , Vineet Thumuluri , Fangchen Liu , Linghao Chen , Zhiao Huang , Hao Su

In this paper, we present a novel architecture to realize fine-grained style control on the transformer-based text-to-speech synthesis (TransformerTTS). Specifically, we model the speaking style by extracting a time sequence of local style…

音频与语音处理 · 电气工程与系统科学 2022-03-18 Li-Wei Chen , Alexander Rudnicky