English
Related papers

Related papers: Being-M0.5: A Real-Time Controllable Vision-Langua…

200 papers

Text-to-motion (T2M) generation aims to control the behavior of a target character via textual descriptions. Leveraging text-motion paired datasets, existing T2M models have achieved impressive performance in generating high-quality motions…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jiakun Zheng , Ting Xiao , Shiqin Cao , Xinran Li , Zhe Wang , Chenjia Bai

Can Multimodal Large Language Models (MLLMs) develop an intuitive number sense similar to humans? Targeting this problem, we introduce Visual Number Benchmark (VisNumBench) to evaluate the number sense abilities of MLLMs across a wide range…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Tengjin Weng , Jingyi Wang , Wenhao Jiang , Zhong Ming

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Yaqi Zhang , Di Huang , Bin Liu , Shixiang Tang , Yan Lu , Lu Chen , Lei Bai , Qi Chu , Nenghai Yu , Wanli Ouyang

Humanoid whole-body loco-manipulation promises transformative capabilities for daily service and warehouse tasks. While recent advances in general motion tracking (GMT) have enabled humanoids to reproduce diverse human motions, these…

Robotics · Computer Science 2025-10-09 Siheng Zhao , Yanjie Ze , Yue Wang , C. Karen Liu , Pieter Abbeel , Guanya Shi , Rocky Duan

This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video…

Artificial Intelligence · Computer Science 2025-03-27 Lei Li , Sen Jia , Jianhao Wang , Zhongyu Jiang , Feng Zhou , Ju Dai , Tianfang Zhang , Zongkai Wu , Jenq-Neng Hwang

We present HaoMo Vision-Language Model (HMVLM), an end-to-end driving framework that implements the slow branch of a cognitively inspired fast-slow architecture. A fast controller outputs low-level steering, throttle, and brake commands,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Daming Wang , Yuhao Song , Zijian He , Kangliang Chen , Xing Pan , Lu Deng , Weihao Gu

Vision-language models (VLMs) have demonstrated remarkable performance across a wide range of computer-vision tasks, sparking interest in their potential for digital health applications. Here, we apply VLMs to two fundamental challenges in…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Victor Li , Naveenraj Kamalakannan , Avinash Parnandi , Heidi Schambra , Carlos Fernandez-Granda

High-quality human motion data is becoming increasingly important for applications in robotics, simulation, and entertainment. Recent generative models offer a potential data source, enabling human motion synthesis through intuitive inputs…

Human motion generation stands as a significant pursuit in generative computer vision, while achieving long-sequence and efficient motion generation remains challenging. Recent advancements in state space models (SSMs), notably Mamba, have…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Zeyu Zhang , Akide Liu , Ian Reid , Richard Hartley , Bohan Zhuang , Hao Tang

High-fidelity motion tracking serves as the ultimate litmus test for generalizable, human-level motor skills. However, current policies often hit a "generality barrier": as motion libraries scale in diversity, tracking fidelity inevitably…

Extracting human motion from large-scale web videos offers a scalable solution to the data scarcity issue in character animation. However, some human parts in many video frames cannot be seen due to off-screen captures or occlusions. It…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Boyuan Li , Sipeng Zheng , Bin Cao , Ruihua Song , Zongqing Lu

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). Despite their…

Computation and Language · Computer Science 2024-08-27 Florian Schneider , Sunayana Sitaram

Real-time human perception is crucial for effective human-robot interaction (HRI). Large vision-language models (VLMs) offer promising generalizable perceptual capabilities but often suffer from high latency, which negatively impacts user…

Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-13 Yin Wang , Zhiying Leng , Frederick W. B. Li , Shun-Cheng Wu , Xiaohui Liang

Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in this domain is limited by the lack of large-scale datasets with biomechanical annotations,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yujun Huo , He Zhang , Chentao Song , Honglin Song , Zongyu Zuo , Tao Yu

Video data is more cost-effective than motion capture data for learning 3D character motion controllers, yet synthesizing realistic and diverse behaviors directly from videos remains challenging. Previous approaches typically rely on…

Graphics · Computer Science 2025-12-10 Jianan Li , Xiao Chen , Tao Huang , Tien-Tsin Wong

Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human perceptions; 2) the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Xinran Ling , Chen Zhu , Meiqi Wu , Hangyu Li , Xiaokun Feng , Cundian Yang , Aiming Hao , Jiashu Zhu , Jiahong Wu , Xiangxiang Chu

Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Zeyu Zhang , Yiran Wang , Wei Mao , Danning Li , Rui Zhao , Biao Wu , Zirui Song , Bohan Zhuang , Ian Reid , Richard Hartley

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Vision Language Model to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Fan Yang , Zhiyang Chen , Yousong Zhu , Xin Li , Jinqiao Wang

The proliferation of wearable technology enables the generation of vast amounts of sensor data, offering significant opportunities for advancements in health monitoring, activity recognition, and personalized medicine. However, the…

Human-Computer Interaction · Computer Science 2024-08-02 Emilio Ferrara