English
Related papers

Related papers: Semantic Belief-State World Model for 3D Human Mot…

200 papers

While non-prehensile manipulation (e.g., controlled pushing/poking) constitutes a foundational robotic skill, its learning remains challenging due to the high sensitivity to complex physical interactions involving friction and restitution.…

Machine Learning · Computer Science 2025-05-06 Wenxuan Li , Hang Zhao , Zhiyuan Yu , Yu Du , Qin Zou , Ruizhen Hu , Kai Xu

Markerless motion capture has become an active field of research in computer vision in recent years. Its extensive applications are known in a great variety of fields, including computer animation, human motion analysis, biomedical…

Computer Vision and Pattern Recognition · Computer Science 2022-01-10 Doan Duy Vo , Russell Butler

Accurate human trajectory prediction is one of the most crucial tasks for autonomous driving, ensuring its safety. Yet, existing models often fail to fully leverage the visual cues that humans subconsciously communicate when navigating the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Yang Gao , Saeed Saadatnejad , Alexandre Alahi

Structure-from-Motion (SfM) aims to recover 3D scene structures and camera poses based on the correspondences between input images, and thus the ambiguity caused by duplicate structures (i.e., different structures with strong visual…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Lei Wang , Linlin Ge , Shan Luo , Zihan Yan , Zhaopeng Cui , Jieqing Feng

Coordinated human movement depends on the integration of multisensory inputs, sensorimotor transformation, and motor execution, as well as sensory feedback resulting from body-environment interaction. Building dynamic models of the…

Neurons and Cognition · Quantitative Biology 2025-06-03 Chenhui Zuo , Guohao Lin , Chen Zhang , Shanning Zhuang , Yanan Sui

Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large multimodal models (LMMs) as priors for reconstructing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Sanjay Subramanian , Evonne Ng , Lea Müller , Dan Klein , Shiry Ginosar , Trevor Darrell

Walking-assistive devices require adaptive control methods to ensure smooth transitions between various modes of locomotion. For this purpose, detecting human locomotion modes (e.g., level walking or stair ascent) in advance is crucial for…

Robotics · Computer Science 2023-11-14 Peiwen Fu , Wenjuan Zhong , Yuyang Zhang , Wenxuan Xiong , Yuzhou Lin , Yanlong Tai , Lin Meng , Mingming Zhang

A key step towards understanding human behavior is the prediction of 3D human motion. Successful solutions have many applications in human tracking, HCI, and graphics. Most previous work focuses on predicting a time series of future 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-04-05 Yan Zhang , Michael J. Black , Siyu Tang

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Xingyu Chen

Accurate human motion prediction is crucial for safe human-robot collaboration but remains challenging due to the complexity of modeling intricate and variable human movements. This paper presents Parallel Multi-scale Incremental Prediction…

Robotics · Computer Science 2024-12-17 Juncheng Zou

Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models forecast the external environment, in-cabin intelligence remains strictly…

Robotics · Computer Science 2026-05-07 Haozhuang Chi , Daosheng Qiu , Hao Su , Haochen Liu , Zirui Li , Haoruo Zhang , Chen Lv

End-to-end autonomous driving systems increasingly rely on vision-centric world models to understand and predict their environment. However, a common ineffectiveness in these models is the full reconstruction of future scenes, which expends…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jianbiao Mei , Yu Yang , Xuemeng Yang , Licheng Wen , Jiajun Lv , Botian Shi , Yong Liu

How does the brain predict physical outcomes while acting in the world? Machine learning world models compress visual input into latent spaces, discarding the spatial structure that characterizes sensory cortex. We propose isomorphic world…

Neurons and Cognition · Quantitative Biology 2026-02-24 Joshua Nunley

Theory of Mind (ToM) reasoning with Large Language Models (LLMs) requires inferring how people's implicit, evolving beliefs shape what they seek and how they act under uncertainty -- especially in high-stakes settings such as disaster…

Artificial Intelligence · Computer Science 2026-03-23 Ruxiao Chen , Xilei Zhao , Thomas J. Cova , Frank A. Drews , Susu Xu

In this paper, we tackle the problem of scene-aware 3D human motion forecasting. A key challenge of this task is to predict future human motions that are consistent with the scene by modeling the human-scene interactions. While recent works…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Chaoyue Xing , Wei Mao , Miaomiao Liu

Urban flow prediction is a spatio-temporal modeling task that estimates the throughput of transportation services like buses, taxis, and ride-sharing, where data-driven models have become the most popular solution in the past decade.…

Machine Learning · Computer Science 2024-08-07 Wei Jiang , Tong Chen , Guanhua Ye , Wentao Zhang , Lizhen Cui , Zi Huang , Hongzhi Yin

We introduce Latent-WAM, an efficient end-to-end autonomous driving framework that achieves strong trajectory planning through spatially-aware and dynamics-informed latent world representations. Existing world-model-based planners suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Linbo Wang , Yupeng Zheng , Qiang Chen , Shiwei Li , Yichen Zhang , Zebin Xing , Qichao Zhang , Xiang Li , Deheng Qian , Pengxuan Yang , Yihang Dong , Ce Hao , Xiaoqing Ye , Junyu han , Yifeng Pan , Dongbin Zhao

Although pretrained language models (PTLMs) contain significant amounts of world knowledge, they can still produce inconsistent answers to questions when probed, even after specialized training. As a result, it can be hard to identify what…

Computation and Language · Computer Science 2021-10-01 Nora Kassner , Oyvind Tafjord , Hinrich Schütze , Peter Clark

Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Wei Wei , Shaojie Zhang , Yonghao Dang , Jianqin Yin

Human motion prediction aims to forecast future human poses given a past motion. Whether based on recurrent or feed-forward neural networks, existing methods fail to model the observation that human motion tends to repeat itself, even for…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Wei Mao , Miaomiao Liu , Mathieu Salzmann
‹ Prev 1 8 9 10 Next ›