English
Related papers

Related papers: COMODO: Cross-Modal Video-to-IMU Distillation for …

200 papers

In the realm of smart sensing with the Internet of Things, earable devices are empowered with the capability of multi-modality sensing and intelligence of context-aware computing, leading to its wide usage in Human Activity Recognition…

Signal Processing · Electrical Eng. & Systems 2024-06-26 Shengzhe Lyu , Yongliang Chen , Di Duan , Renqi Jia , Weitao Xu

In the era of Industry 5.0, monitoring human activity is essential for ensuring both ergonomic safety and overall well-being. While multi-camera centralized setups improve pose estimation accuracy, they often suffer from high computational…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Enrico Martini , Ho Jin Choi , Nadia Figueroa , Nicola Bombieri

Recent advances in imitation learning have shown great promise for developing robust robot manipulation policies from demonstrations. However, this promise is contingent on the availability of diverse, high-quality datasets, which are not…

Robotics · Computer Science 2025-09-24 Omar Rayyan , John Abanes , Mahmoud Hafez , Anthony Tzes , Fares Abu-Dakka

Understanding sensor data can be difficult for non-experts because of the complexity and different semantic meanings of sensor modalities. This leads to a need for intuitive and effective methods to present sensor information. However,…

Human-Computer Interaction · Computer Science 2025-03-26 Yunqi Guo , Kaiyuan Hou , Heming Fu , Hongkai Chen , Zhenyu Yan , Guoliang Xing , Xiaofan Jiang

Multimodal emotion recognition analyzes emotions by combining data from multiple sources. However, real-world noise or sensor failures often cause missing or corrupted data, creating the Incomplete Multimodal Emotion Recognition (IMER)…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Yuehan Jin , Xiaoqing Liu , Yiyuan Yang , Zhiwen Yu , Tong Zhang , Kaixiang Yang

Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Zhen Fan , Peng Dai , Zhuo Su , Xu Gao , Zheng Lv , Jiarui Zhang , Tianyuan Du , Guidong Wang , Yang Zhang

Multimodal egocentric activity recognition integrates visual and inertial cues for robust first-person behavior understanding. However, deploying such systems in open-world environments requires detecting novel activities while continuously…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Wonseon Lim , Hyejeong Im , Dae-Won Kim

Despite LiDAR (Light Detection and Ranging) being an effective privacy-preserving alternative to RGB cameras to perceive human activities, it remains largely underexplored in the context of multi-modal contrastive pre-training for human…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Thomas Kreutz , Max Mühlhäuser , Alejandro Sanchez Guinea

Human Activity Recognition (HAR) systems have been extensively studied by the vision and ubiquitous computing communities due to their practical applications in daily life, such as smart homes, surveillance, and health monitoring.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Hyeongju Choi , Apoorva Beedu , Irfan Essa

Although most existing multi-modal salient object detection (SOD) methods demonstrate effectiveness through training models from scratch, the limited multi-modal data hinders these methods from reaching optimality. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Kunpeng Wang , Danying Lin , Chenglong Li , Zhengzheng Tu , Bin Luo

The primary objective of human activity recognition (HAR) is to infer ongoing human actions from sensor data, a task that finds broad applications in health monitoring, safety protection, and sports analysis. Despite proliferating research,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Hang Xiao , Ying Yu , Jiarui Li , Zhifan Yang , Haotian Tang , Hanyu Liu , Chao Li

Task-based measures of image quality (IQ) are critical for evaluating medical imaging systems, which must account for randomness including anatomical variability. Stochastic object models (SOMs) provide a statistical description of such…

Graphics · Computer Science 2026-03-26 Xiaoning Lei , Jianwei Sun , Wenhao Cai , Xichen Xu , Yanshu Wang , Hu Gao

The research introduces a reproducible framework for transforming raw, heterogeneous sensor streams into aligned, semantically meaningful representations for multimodal human activity recognition. Grounded in the Carnegie Mellon University…

Applications · Statistics 2026-05-05 Yiyao Yang , Yasemin Gulbahar

Existing multi-modal fusion methods typically apply static frame-based image fusion techniques directly to video fusion tasks, neglecting inherent temporal dependencies and leading to inconsistent results across frames. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Meiqi Gong , Hao Zhang , Xunpeng Yi , Linfeng Tang , Jiayi Ma

Multi-modal recommendation has gained traction as items possess rich attributes like text and images. Semantic ID-based approaches effectively discretize this information into compact tokens. However, two challenges persist: (1) Suboptimal…

Artificial Intelligence · Computer Science 2026-05-27 Pingjun Pan , Tingting Zhou , Peiyao Lu , Tingting Fei , Hongxiang Chen , Chuanjiang Luo

Online high-definition (HD) map construction is an important and challenging task in autonomous driving. Recently, there has been a growing interest in cost-effective multi-view camera-based methods without relying on other sensors like…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Xiaoshuai Hao , Ruikai Li , Hui Zhang , Dingzhe Li , Rong Yin , Sangil Jung , Seung-In Park , ByungIn Yoo , Haimei Zhao , Jing Zhang

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous modalities, even…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Yanbei Chen , Yongqin Xian , A. Sophia Koepke , Ying Shan , Zeynep Akata

Wearable devices enable continuous multi-modal physiological and behavioral monitoring, yet analysis of these data streams faces fundamental challenges including the lack of gold-standard labels and incomplete sensor data. While…

Machine Learning · Statistics 2025-09-22 Howon Ryu , Yuliang Chen , Yacun Wang , Andrea Z. LaCroix , Chongzhi Di , Loki Natarajan , Yu Wang , Jingjing Zou

Human Activity Recognition~(HAR) is the classification of human movement, captured using one or more sensors either as wearables or embedded in the environment~(e.g. depth cameras, pressure mats). State-of-the-art methods of HAR rely on…

Computer Vision and Pattern Recognition · Computer Science 2020-06-16 Anjana Wijekoon , Nirmalie Wiratunga

Recently, emotion recognition based on physiological signals has emerged as a field with intensive research. The utilization of multi-modal, multi-channel physiological signals has significantly improved the performance of emotion…

Multimedia · Computer Science 2023-08-22 Xinda Li