English
Related papers

Related papers: X-Fi: A Modality-Invariant Foundation Model for Mu…

200 papers

In this paper, we consider the problem of multimodal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality…

Computer Vision and Pattern Recognition · Computer Science 2022-01-27 Kateryna Chumachenko , Alexandros Iosifidis , Moncef Gabbouj

The correlation of optical measurements with a correct pathology label is often hampered by imprecise registration caused by deformations in histology images. This study explores an automated multi-modal image registration technique…

Image and Video Processing · Electrical Eng. & Systems 2023-11-27 Lianne Feenstra , Maud Lambregts , Theo J. M Ruers , Behdad Dashtbozorg

Human Action Recognition (HAR), one of the most important tasks in computer vision, has developed rapidly in the past decade and has a wide range of applications in health monitoring, intelligent surveillance, virtual reality, human…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Zhou Shuchang

Besides standard cameras, autonomous vehicles typically include multiple additional sensors, such as lidars and radars, which help acquire richer information for perceiving the content of the driving scene. While several recent works focus…

Computer Vision and Pattern Recognition · Computer Science 2023-08-14 Tim Broedermann , Christos Sakaridis , Dengxin Dai , Luc Van Gool

As an effective approach to understanding the human-centric physical world, Wearable Artificial Intelligence (AI), which leverages multimodal wearable sensors to understand human physiology and behavior, has attracted increasing attention…

Signal Processing · Electrical Eng. & Systems 2026-04-14 Yize Cai , Baoshen Guo , Guobin Shen , Zhiqing Hong

Human Activity Recognition (HAR) with wearable sensors is challenged by limited interpretability, which significantly impacts cross-dataset generalization. To address this challenge, we propose Motion-Primitive Transformer (MoPFormer), a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Hao Zhang , Zhan Zhuang , Xuehao Wang , Xiaodong Yang , Yu Zhang

In the realm of geospatial analysis, the diversity of remote sensors, encompassing both optical and microwave technologies, offers a wealth of distinct observational capabilities. Recognizing this, we present msGFM, a multisensor geospatial…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Boran Han , Shuai Zhang , Xingjian Shi , Markus Reichstein

In a human-centered intelligent manufacturing system, sensing and understanding of the worker's activity are the primary tasks. In this paper, we propose a novel multi-modal approach for worker activity recognition by leveraging information…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Wenjin Tao , Ming C. Leu , Zhaozheng Yin

Multimodal federated learning (FL) aims to enrich model training in FL settings where devices are collecting measurements across multiple modalities (e.g., sensors measuring pressure, motion, and other types of data). However, key…

Machine Learning · Computer Science 2024-08-21 Liangqi Yuan , Dong-Jun Han , Vishnu Pandi Chellapandi , Stanislaw H. Żak , Christopher G. Brinton

Being able to explain the prediction to clinical end-users is a necessity to leverage the power of artificial intelligence (AI) models for clinical decision support. For medical images, a feature attribution map, or heatmap, is the most…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Weina Jin , Xiaoxiao Li , Ghassan Hamarneh

The ability to associate touch with other modalities has huge implications for humans and computational systems. However, multimodal learning with touch remains challenging due to the expensive data collection process and non-standardized…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Fengyu Yang , Chao Feng , Ziyang Chen , Hyoungseob Park , Daniel Wang , Yiming Dou , Ziyao Zeng , Xien Chen , Rit Gangopadhyay , Andrew Owens , Alex Wong

Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ruoyu Wang , Wenqian Wang , Jianjun Gao , Dan Lin , Kim-Hui Yap , Bingbing Li

The human somatosensory system integrates multimodal sensory feedback, including tactile, proprioceptive, and thermal signals, to enable comprehensive perception and effective interaction with the environment. Inspired by the biological…

Robotics · Computer Science 2025-09-03 Fengyi Wang , Xiangyu Fu , Nitish Thakor , Gordon Cheng

Multi-modal fusion of sensors is a commonly used approach to enhance the performance of odometry estimation, which is also a fundamental module for mobile robots. However, the question of \textit{how to perform fusion among different…

Robotics · Computer Science 2025-03-20 Leyuan Sun , Guanqun Ding , Yue Qiu , Yusuke Yoshiyasu , Fumio Kanehiro

Multi-modal fusion is crucial for Internet of Things (IoT) perception, widely deployed in smart homes, intelligent transport, industrial automation, and healthcare. However, existing systems often face challenges: high model complexity…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Weiqi Yang , Xu Zhou , Jingfu Guan , Hao Du , Tianyu Bai

We describe a novel metric-based learning approach that introduces a multimodal framework and uses deep audio and geophone encoders in siamese configuration to design an adaptable and lightweight supervised model. This framework eliminates…

Sound · Computer Science 2021-11-16 Muhammad Shakeel , Katsutoshi Itoyama , Kenji Nishida , Kazuhiro Nakadai

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments generate vast amounts of multimodal data, including visual imagery from various…

An important current challenge in Human-Robot Interaction (HRI) is to enable robots to learn on-the-fly from human feedback. However, humans show a great variability in the way they reward robots. We propose to address this issue by…

Robotics · Computer Science 2020-05-11 Rémi Dromnelle , Benoît Girard , Erwan Renaudo , Raja Chatila , Mehdi Khamassi

Multimodal learning involves integrating information from various modalities to enhance learning and comprehension. We compare three modality fusion strategies in person identification and verification by processing two modalities: voice…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Aref Farhadipour , Masoumeh Chapariniya , Teodora Vukovic , Volker Dellwo

We present X-UniMotion, a unified and expressive implicit latent representation for whole-body human motion, encompassing facial expressions, body poses, and hand gestures. Unlike prior motion transfer methods that rely on explicit skeletal…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Guoxian Song , Hongyi Xu , Xiaochen Zhao , You Xie , Tianpei Gu , Zenan Li , Chenxu Zhang , Linjie Luo