中文
相关论文

相关论文: MMHU: A Massive-Scale Multimodal Benchmark for Hum…

200 篇论文

Predicting human mobility is crucial for urban planning, traffic control, and emergency response. Mobility behaviors can be categorized into individual and collective, and these behaviors are recorded by diverse mobility data, such as…

机器学习 · 计算机科学 2024-12-23 Qingyue Long , Yuan Yuan , Yong Li

As real-world knowledge continues to evolve, the parametric knowledge acquired by multimodal models during pretraining becomes increasingly difficult to remain consistent with real-world knowledge. Existing research on multimodal knowledge…

计算与语言 · 计算机科学 2026-03-17 Baochen Fu , Yuntao Du , Cheng Chang , Baihao Jin , Wenzhi Deng , Muhao Xu , Hongmei Yan , Weiye Song , Yi Wan

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Bizhu Wu , Jinheng Xie , Meidan Ding , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Recent advances in large language models (LLMs) have increased the demand for comprehensive benchmarks to evaluate their capabilities as human-like agents. Existing benchmarks, while useful, often focus on specific application scenarios,…

Human mobility traces, often recorded as sequences of check-ins, provide a unique window into both short-term visiting patterns and persistent lifestyle regularities. In this work we introduce GSTM-HMU, a generative spatio-temporal…

机器学习 · 计算机科学 2025-09-24 Wenying Luo , Zhiyuan Lin , Wenhao Xu , Minghao Liu , Zhi Li

With the advancement in computer vision deep learning, systems now are able to analyze an unprecedented amount of rich visual information from videos to enable applications such as autonomous driving, socially-aware robot assistant and…

计算机视觉与模式识别 · 计算机科学 2021-07-19 Junwei Liang

While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging…

Rapid development of social robots stimulates active research in human motion modeling, interpretation and prediction, proactive collision avoidance, human-robot interaction and co-habitation in shared spaces. Modern approaches to this end…

Real-world scenes often feature multiple humans interacting with multiple objects in ways that are causal, goal-oriented, or cooperative. Yet existing 3D human-object interaction (HOI) benchmarks consider only a fraction of these complex…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Kaen Kogashi , Anoop Cherian , Meng-Yu Jennifer Kuo

Autonomous driving systems must operate reliably in safety-critical scenarios, particularly those involving unusual or complex behavior by Vulnerable Road Users (VRUs). Identifying these edge cases in driving datasets is essential for…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Stefan Englmeier , Max A. Büttner , Katharina Winter , Fabian B. Flohr

Human detection has witnessed impressive progress in recent years. However, the occlusion issue of detecting human in highly crowded environments is far from solved. To make matters worse, crowd scenarios are still under-represented in…

计算机视觉与模式识别 · 计算机科学 2018-05-02 Shuai Shao , Zijian Zhao , Boxun Li , Tete Xiao , Gang Yu , Xiangyu Zhang , Jian Sun

In light of growing attention of intelligent vehicle systems, we propose developing a driver model that uses a hybrid system formulation to capture the intent of the driver. This model hopes to capture human driving behavior in a way that…

系统与控制 · 计算机科学 2015-05-25 Katherine Driggs-Campbell , Ruzena Bajcsy

Over the years, the separate fields of motion planning, mapping, and human trajectory prediction have advanced considerably. However, the literature is still sparse in providing practical frameworks that enable mobile manipulators to…

机器人学 · 计算机科学 2022-07-27 Mark Nicholas Finean , Luka Petrović , Wolfgang Merkt , Ivan Marković , Ioannis Havoutis

The development of automated vehicles has the potential to revolutionize transportation, but they are currently unable to ensure a safe and time-efficient driving style. Reliable models predicting human behavior are essential for overcoming…

Understanding the short-term motion of vulnerable road users (VRUs) like pedestrians and cyclists is critical for safe autonomous driving, especially in urban scenarios with ambiguous or high-risk behaviors. While vision-language models…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Mihir Godbole , Xiangbo Gao , Zhengzhong Tu

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yufan Deng , Daquan Zhou

Driving Scene understanding is a key ingredient for intelligent transportation systems. To achieve systems that can operate in a complex physical and social environment, they need to understand and learn how humans drive and interact with…

计算机视觉与模式识别 · 计算机科学 2018-11-07 Vasili Ramanishka , Yi-Ting Chen , Teruhisa Misu , Kate Saenko

Multimodal large language models (MLLMs) have been widely applied across various fields due to their powerful perceptual and reasoning capabilities. In the realm of psychology, these models hold promise for a deeper understanding of human…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Jinpeng Hu , Hongchang Shi , Chongyuan Dai , Zhuo Li , Peipei Song , Meng Wang

As a window for urban sensing, human mobility contains rich spatiotemporal information that reflects both residents' behavior preferences and the functions of urban areas. The analysis of human mobility has attracted the attention of many…

社会与信息网络 · 计算机科学 2025-10-29 Liangzhe Han , Leilei Sun , Tongyu Zhu , Tao Tao , Jibin Wang , Weifeng Lv

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks involving both images and videos. However, their capacity to comprehend human-centric video data remains underexplored, primarily…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuxuan Cai , Jiangning Zhang , Zhenye Gan , Qingdong He , Xiaobin Hu , Junwei Zhu , Yabiao Wang , Chengjie Wang , Zhucun Xue , Chaoyou Fu , Xinwei He , Xiang Bai