中文
相关论文

相关论文: Being-M0.5: A Real-Time Controllable Vision-Langua…

200 篇论文

Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Sriram Narayanan , Ziyu Jiang , Srinivasa Narasimhan , Manmohan Chandraker

Recent advances in data-driven reinforcement learning and motion tracking have substantially improved humanoid locomotion, yet critical practical challenges remain. In particular, while low-level motion tracking and trajectory-following…

机器人学 · 计算机科学 2026-02-25 Chenxi Han , Yuheng Min , Zihao Huang , Ao Hong , Hang Liu , Yi Cheng , Houde Liu

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Xindi Yang , Baolu Li , Yiming Zhang , Zhenfei Yin , Lei Bai , Liqian Ma , Zhiyong Wang , Jianfei Cai , Tien-Tsin Wong , Huchuan Lu , Xu Jia

Vision based human motion recognition has fascinated many researchers due to its critical challenges and a variety of applications. The applications range from simple gesture recognition to complicated behaviour understanding in…

计算机视觉与模式识别 · 计算机科学 2016-08-25 Geetanjali Vinayak Kale , Varsha Hemant Patil

In this paper, we introduce DirectorLLM, a novel video generation model that employs a large language model (LLM) to orchestrate human poses within videos. As foundational text-to-video models rapidly evolve, the demand for high-quality…

Text-to-Motion (T2M) generation aims to synthesize realistic human motion sequences from natural language descriptions. While two-stage frameworks leveraging discrete motion representations have advanced T2M research, they often neglect…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Hongsong Wang , Wenjing Yan , Qiuxia Lai , Xin Geng

Millimeter-wave radar offers a privacy-preserving and environment-robust alternative to vision-based sensing, enabling human motion analysis in challenging conditions such as low light, occlusions, rain, or smoke. However, its sparse point…

机器学习 · 计算机科学 2025-11-18 Zengyuan Lai , Jiarui Yang , Songpengcheng Xia , Lizhou Lin , Lan Sun , Renwen Wang , Jianran Liu , Qi Wu , Ling Pei

Humanoid robots are machines built with an anthropomorphic shape. Despite decades of research into the subject, it is still challenging to tackle the robot locomotion problem from an algorithmic point of view. For example, these machines…

机器人学 · 计算机科学 2020-04-28 Stefano Dafarra

Learning motor control for muscle-driven musculoskeletal models is hindered by the computational cost of biomechanically accurate simulation and the scarcity of validated, open full-body models. Here we present MuscleMimic, an open-source…

Integration of diverse data will be a pivotal step towards improving scientific explorations in many disciplines. This work establishes a vision-language model (VLM) that encodes videos with text input in order to classify various behaviors…

机器学习 · 计算机科学 2025-10-23 Paimon Goulart , Jordan Steinhauser , Kylene Shuler , Edward Korzus , Jia Chen , Evangelos E. Papalexakis

We present DiverseMotion, a new approach for synthesizing high-quality human motions conditioned on textual descriptions while preserving motion diversity.Despite the recent significant process in text-based human motion generation,existing…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Yunhong Lou , Linchao Zhu , Yaxiong Wang , Xiaohan Wang , Yi Yang

Effective human-robot interaction requires emotionally rich multimodal expressions, yet most humanoid robots lack coordinated speech, facial expressions, and gestures. Meanwhile, real-world deployment demands on-device solutions that can…

机器人学 · 计算机科学 2026-02-10 Songhua Yang , Xuetao Li , Xuanye Fei , Mengde Li , Miao Li

Learned world models hold significant potential for robotic manipulation, as they can serve as simulator for real-world interactions. While extensive progress has been made in 2D video-based world models, these approaches often lack…

机器人学 · 计算机科学 2025-10-13 Chuanrui Zhang , Zhengxian Wu , Guanxing Lu , Yansong Tang , Ziwei Wang

Large-scale human mobility simulation is critical for many science domains such as urban science, epidemiology, and transportation analysis. Recent works treat large language models (LLMs) as human agents to simulate realistic mobility…

多智能体系统 · 计算机科学 2026-02-20 Hua Yan , Heng Tan , Yu Yang

Due to the several applications on Human-machine interaction (HMI), this area of research has become one of the most popular in recent years. This is the case for instance of advanced training machines, robots for rehabilitation, robotic…

机器人学 · 计算机科学 2020-06-03 Humberto De las Casas , Hanz Richter

Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for visual content…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Rongyao Fang , Chengqi Duan , Kun Wang , Hao Li , Hao Tian , Xingyu Zeng , Rui Zhao , Jifeng Dai , Hongsheng Li , Xihui Liu

Controllable human motion synthesis is essential for applications in AR/VR, gaming and embodied AI. Existing methods often focus solely on either language or full trajectory control, lacking precision in synthesizing motions aligned with…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Weilin Wan , Zhiyang Dou , Taku Komura , Wenping Wang , Dinesh Jayaraman , Lingjie Liu

While Multimodal Large Language Models (MLLMs) excel at many vision tasks, it is unknown if they exhibit human-like perceptual behaviors. To evaluate this, we introduce HVSBench, the first large-scale benchmark with over 85,000 samples…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Jiaying Lin , Shuquan Ye , Dan Xu , Wanli Ouyang , Rynson W. H. Lau

Visual reasoning is a core component of human intelligence and a critical capability for advanced multimodal models. Yet current reasoning evaluations of multimodal large language models (MLLMs) often rely on text descriptions and allow…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Weiye Xu , Jiahao Wang , Weiyun Wang , Zhe Chen , Wengang Zhou , Aijun Yang , Lewei Lu , Houqiang Li , Xiaohua Wang , Xizhou Zhu , Wenhai Wang , Jifeng Dai , Jinguo Zhu

We present a novel approach named OmniControl for incorporating flexible spatial control signals into a text-conditioned human motion generation model based on the diffusion process. Unlike previous methods that can only control the pelvis…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Yiming Xie , Varun Jampani , Lei Zhong , Deqing Sun , Huaizu Jiang
‹ 上一页 1 8 9 10 下一页 ›