中文
相关论文

相关论文: Exploring Motion-Language Alignment for Text-drive…

200 篇论文

This paper addresses the problem of generating 3D interactive human motion from text. Given a textual description depicting the actions of different body parts in contact with static objects, we synthesize sequences of 3D body poses that…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Sihan Ma , Qiong Cao , Jing Zhang , Dacheng Tao

Large Language Models (LLMs) have revolutionised the field of Natural Language Processing (NLP) and have achieved state-of-the-art performance in practically every task in this field. However, the prevalent approach used in text generation,…

计算与语言 · 计算机科学 2024-08-12 Nicolo Micheletti , Samuel Belkadi , Lifeng Han , Goran Nenadic

Personalized image generation aims to integrate user-provided concepts into text-to-image models, enabling the generation of customized content based on a given prompt. Recent zero-shot approaches, particularly those leveraging diffusion…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Yiheng Lin , Shifang Zhao , Ting Liu , Xiaochao Qu , Luoqi Liu , Yao Zhao , Yunchao Wei

Recent Multimodal Large Language Models (MLLMs) have demonstrated significant progress in perceiving and reasoning over multimodal inquiries, ushering in a new research era for foundation models. However, vision-language misalignment in…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Wei-Yao Wang , Zhao Wang , Helen Suzuki , Yoshiyuki Kobayashi

Various adaptive abilities are required for robots interacting with humans in daily life. It is difficult to design adaptive algorithms manually; however, by using end-to-end machine learning, labor can be saved during the design process.…

机器人学 · 计算机科学 2019-09-20 Kazuki Fujimoto , Sho Sakaino , Toshiaki Tsuji

Vision-Language-Action (VLA) models improve action generation by conditioning policies on rich vision-language information. However, current auto-regressive policies are constrained by three bottlenecks: (1) architectural bias drives models…

机器人学 · 计算机科学 2026-03-31 Yichi Zhang , Weihao Yuan , Yizhuo Zhang , Xidong Zhang , Jia Wan

In this paper, we address the challenging problem of long-term 3D human motion generation. Specifically, we aim to generate a long sequence of smoothly connected actions from a stream of multiple sentences (i.e., paragraph). Previous…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Taeryung Lee , Fabien Baradel , Thomas Lucas , Kyoung Mu Lee , Gregory Rogez

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Wendong Bu , Kaihang Pan , Yuze Lin , Jiacheng Li , Kai Shen , Wenqiao Zhang , Juncheng Li , Jun Xiao , Siliang Tang

Text-to-motion generation has advanced rapidly, yet two challenges persist. First, existing motion autoencoders compress each frame into a single monolithic latent vector, entangling trajectory and per-joint rotations in an unstructured…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Zeyu Ling , Qing Shuai , Teng Zhang , Shiyang Li , Bo Han , Changqing Zou

Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limitations of pose detection techniques, the extracted pose…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Yuhang Zhang , Yuan Zhou , Zeyu Liu , Yuxuan Cai , Qiuyue Wang , Aidong Men , Huan Yang

Video generation has witnessed remarkable progress with the advent of deep generative models, particularly diffusion models. While existing methods excel in generating high-quality videos from text prompts or single images, personalized…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Yufan Deng , Xun Guo , Yizhi Wang , Jacob Zhiyuan Fang , Angtian Wang , Shenghai Yuan , Yiding Yang , Bo Liu , Haibin Huang , Chongyang Ma

In this work, we present TextHarmony, a unified and versatile multimodal generative model proficient in comprehending and generating visual text. Simultaneously generating images and texts typically results in performance degradation due to…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Zhen Zhao , Jingqun Tang , Binghong Wu , Chunhui Lin , Shu Wei , Hao Liu , Xin Tan , Zhizhong Zhang , Can Huang , Yuan Xie

We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture. Existing unified models handle only restricted…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Ziyi Wang , Xinshun Wang , Shuang Chen , Yang Cong , Mengyuan Liu

In recent years, studies have been actively conducted on combining large language models (LLM) and robotics; however, most have not considered end-to-end feedback in the robot-motion generation phase. The prediction of deep neural networks…

机器人学 · 计算机科学 2024-07-15 Kanata Suzuki , Tetsuya Ogata

Conditional human motion generation is an important topic with many applications in virtual reality, gaming, and robotics. While prior works have focused on generating motion guided by text, music, or scenes, these typically result in…

计算机视觉与模式识别 · 计算机科学 2024-02-26 German Barquero , Sergio Escalera , Cristina Palmero

We introduce a novel method for controlling a motion sequence using an arbitrary temporal control sequence using temporal alignment. Temporal alignment of motion has gained significant attention owing to its applications in motion control…

图形学 · 计算机科学 2025-11-26 Naoki Agata , Takeo Igarashi

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Shijie Wang , Samaneh Azadi , Rohit Girdhar , Saketh Rambhatla , Chen Sun , Xi Yin

Vision-language models (VLMs) have made significant strides in cross-modal understanding through large-scale paired datasets. However, in fashion domain, datasets often exhibit a disparity between the information conveyed in image and text.…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Chull Hwan Song , Taebaek Hwang , Jooyoung Yoon , Shunghyun Choi , Yeong Hyeon Gu

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in descriptive motion…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Anna Deichler , Jim O'Regan , Teo Guichoux , David Johansson , Jonas Beskow

Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Chuan Guo , Inwoo Hwang , Jian Wang , Bing Zhou