English
Related papers

Related papers: HuMMan: Multi-Modal 4D Human Dataset for Versatile…

200 papers

Autonomous trucking is a promising technology that can greatly impact modern logistics and the environment. Ensuring its safety on public roads is one of the main duties that requires an accurate perception of the environment. To achieve…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Felix Fent , Fabian Kuttenreich , Florian Ruch , Farija Rizwin , Stefan Juergens , Lorenz Lechermann , Christian Nissler , Andrea Perl , Ulrich Voll , Min Yan , Markus Lienkamp

Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in this domain is limited by the lack of large-scale datasets with biomechanical annotations,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yujun Huo , He Zhang , Chentao Song , Honglin Song , Zongyu Zuo , Tao Yu

To properly assist humans in their needs, human activity recognition (HAR) systems need the ability to fuse information from multiple modalities. Our hypothesis is that multimodal sensors, visual and non-visual tend to provide complementary…

Computer Vision and Pattern Recognition · Computer Science 2022-11-09 Hyeongju Choi , Apoorva Beedu , Harish Haresamudram , Irfan Essa

Advancements in deep neural networks have contributed to near perfect results for many computer vision problems such as object recognition, face recognition and pose estimation. However, human action recognition is still far from…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Asanka G. Perera , Yee Wei Law , Titilayo T. Ogunwa , Javaan Chahl

Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine. However, most publicly available datasets on human body movements cannot be used to study both problems in an…

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Lixiong Qin , Shilong Ou , Miaoxuan Zhang , Jiangning Wei , Yuhang Zhang , Xiaoshuai Song , Yuchen Liu , Mei Wang , Weiran Xu

Multimodal data modeling has emerged as a powerful approach in clinical research, enabling the integration of diverse data types such as imaging, genomics, wearable sensors, and electronic health records. Despite its potential to improve…

Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Jianglin Fu , Shikai Li , Yuming Jiang , Kwan-Yee Lin , Wayne Wu , Ziwei Liu

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

The popularity and diffusion of wearable devices provides new opportunities for sensor-based human activity recognition that leverages deep learning-based algorithms. Although impressive advances have been made, two major challenges remain.…

Signal Processing · Electrical Eng. & Systems 2024-02-16 Mengna Liu , Dong Xiang , Xu Cheng , Xiufeng Liu , Dalin Zhang , Shengyong Chen , Christian S. Jensen

We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes with both wide…

Computer Vision and Pattern Recognition · Computer Science 2020-03-11 Xueyang Wang , Xiya Zhang , Yinheng Zhu , Yuchen Guo , Xiaoyun Yuan , Liuyu Xiang , Zerun Wang , Guiguang Ding , David J Brady , Qionghai Dai , Lu Fang

A lot of real-world phenomena are complex and cannot be captured by single task annotations. This causes a need for subsequent annotations, with interdependent questions and answers describing the nature of the subject at hand. Even in the…

Computation and Language · Computer Science 2020-10-05 Moritz Wolf , Dana Ruiter , Ashwin Geet D'Sa , Liane Reiners , Jan Alexandersson , Dietrich Klakow

Understanding how humans cooperatively rearrange household objects is critical for VR/AR and human-robot interaction. However, in-depth studies on modeling these behaviors are under-researched due to the lack of relevant datasets. We fill…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Yun Liu , Chengwen Zhang , Ruofan Xing , Bingda Tang , Bowen Yang , Li Yi

Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward…

Human sensing, which employs various sensors and advanced deep learning technologies to accurately capture and interpret human body information, has significantly impacted fields like public security and robotics. However, current human…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Xinyan Chen , Jianfei Yang

We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in…

The development of multi-modal learning for Unmanned Aerial Vehicles (UAVs) typically relies on a large amount of pixel-aligned multi-modal image data. However, existing datasets face challenges such as limited modalities, high construction…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Liang Yao , Fan Liu , Shengxiang Xu , Chuanyi Zhang , Xing Ma , Jianyu Jiang , Zequan Wang , Shimin Di , Jun Zhou

Human Action Recognition (HAR), one of the most important tasks in computer vision, has developed rapidly in the past decade and has a wide range of applications in health monitoring, intelligent surveillance, virtual reality, human…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Zhou Shuchang

The ability to grasp objects, signal with gestures, and share emotion through touch all stem from the unique capabilities of human hands. Yet creating high-quality personalized hand avatars from images remains challenging due to complex…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Zicong Fan , Edoardo Remelli , David Dimond , Fadime Sener , Liuhao Ge , Bugra Tekin , Cem Keskin , Shreyas Hampali