中文
相关论文

相关论文: GenM$^3$: Generative Pretrained Multi-path Motion …

200 篇论文

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

Human mobility traces, often recorded as sequences of check-ins, provide a unique window into both short-term visiting patterns and persistent lifestyle regularities. In this work we introduce GSTM-HMU, a generative spatio-temporal…

机器学习 · 计算机科学 2025-09-24 Wenying Luo , Zhiyuan Lin , Wenhao Xu , Minghao Liu , Zhi Li

Text-to-Motion (T2M) generation aims to synthesize realistic and semantically aligned human motion sequences from natural language descriptions. However, current approaches face dual challenges: Generative models (e.g., diffusion models)…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Zhengdao Li , Siheng Wang , Zeyu Zhang , Hao Tang

This work introduces the M6(GPT)3 composer system, capable of generating complete, multi-minute musical compositions with complex structures in any time signature, in the MIDI domain from input descriptions in natural language. The system…

声音 · 计算机科学 2025-09-30 Jakub Poćwiardowski , Mateusz Modrzejewski , Marek S. Tatara

Human motion prediction from historical pose sequence is at the core of many applications in machine intelligence. However, in current state-of-the-art methods, the predicted future motion is confined within the same activity. One can…

计算机视觉与模式识别 · 计算机科学 2021-03-18 Zhenguang Liu , Kedi Lyu , Shuang Wu , Haipeng Chen , Yanbin Hao , Shouling Ji

We propose to learn a probabilistic motion model from a sequence of images for spatio-temporal registration. Our model encodes motion in a low-dimensional probabilistic space - the motion matrix - which enables various motion analysis tasks…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Julian Krebs , Hervé Delingette , Nicholas Ayache , Tommaso Mansi

We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In MoMask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Chuan Guo , Yuxuan Mu , Muhammad Gohar Javed , Sen Wang , Li Cheng

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Yaqi Zhang , Di Huang , Bin Liu , Shixiang Tang , Yan Lu , Lu Chen , Lei Bai , Qi Chu , Nenghai Yu , Wanli Ouyang

Currently, human-bot symbiosis dialog systems, e.g., pre- and after-sales in E-commerce, are ubiquitous, and the dialog routing component is essential to improve the overall efficiency, reduce human resource cost, and enhance user…

计算与语言 · 计算机科学 2023-04-10 Ziming Huang , Zhuoxuan Jiang , Ke Wang , Juntao Li , Shanshan Feng , Xian-Ling Mao

Diffusion models have become a popular choice for human motion synthesis due to their powerful generative capabilities. However, their high computational complexity and large sampling steps pose challenges for real-time applications.…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Lei Jiang , Ye Wei , Hao Ni

Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities but often struggle with complex, multi-step mathematical reasoning, where minor errors in visual perception or logical deduction can lead to complete failure.…

计算与语言 · 计算机科学 2025-08-08 Jianghangfan Zhang , Yibo Yan , Kening Zheng , Xin Zou , Song Dai , Xuming Hu

Online High-Definition (HD) maps have emerged as the preferred option for autonomous driving, overshadowing the counterpart offline HD maps due to flexible update capability and lower maintenance costs. However, contemporary online HD map…

计算机视觉与模式识别 · 计算机科学 2024-09-16 Siyu Li , Kailun Yang , Hao Shi , Song Wang , You Yao , Zhiyong Li

Background: Machine learning (ML) enhances gait analysis but often lacks the level of interpretability desired for clinical adoption. Large Language Models (LLMs) may offer explanatory capabilities and confidence-aware outputs when applied…

机器学习 · 计算机科学 2026-03-17 Carlo Dindorf , Jonas Dully , Rebecca Keilhauer , Michael Lorenz , Michael Fröhlich

The primary challenges in visible-infrared person re-identification arise from the differences between visible (vis) and infrared (ir) images, including inter-modal and intra-modal variations. These challenges are further complicated by…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Jiarui Li , Zhen Qiu , Yilin Yang , Yuqi Li , Zeyu Dong , Chuanguang Yang

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

Electromyography (EMG)-based gesture recognition has emerged as a promising approach for human-computer interaction. However, its performance is often limited by the scarcity of labeled EMG data, significant cross-user variability, and poor…

人机交互 · 计算机科学 2025-12-11 Nana Wang , Gen Li , Pengfei Ren , Hao Su , Suli Wang

We present LL3M, a multi-agent system that leverages pretrained large language models (LLMs) to generate 3D assets by writing interpretable Python code in Blender. We break away from the typical generative approach that learns from a…

图形学 · 计算机科学 2025-08-12 Sining Lu , Guan Chen , Nam Anh Dinh , Itai Lang , Ari Holtzman , Rana Hanocka

Human motion generation has advanced markedly with the advent of diffusion models. Most recent studies have concentrated on generating motion sequences based on text prompts, commonly referred to as text-to-motion generation. However, the…

计算机视觉与模式识别 · 计算机科学 2025-01-29 Zhongyu Jiang , Wenhao Chai , Zhuoran Zhou , Cheng-Yen Yang , Hsiang-Wei Huang , Jenq-Neng Hwang

Recent studies indicate that the denoising process in deep generative diffusion models implicitly learns and memorizes semantic information from the data distribution. These findings suggest that capturing more complex data distributions…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Yi Tang , Peng Sun , Zhenglin Cheng , Tao Lin

Developing bipedal robots capable of traversing diverse real-world terrains presents a fundamental robotics challenge, as existing methods using predefined height maps and static environments fail to address the complexity of unstructured…

机器人学 · 计算机科学 2025-04-15 Hanwen Wan , Mengkang Li , Donghao Wu , Yebin Zhong , Yixuan Deng , Zhenglong Sun , Xiaoqiang Ji