English
Related papers

Related papers: Multimodal Foundation Model for Cross-Modal Retrie…

200 papers

This paper presents a novel framework for estimating the position and orientation of flexible manipulators undergoing vertical motion using multiple inertial measurement units (IMUs), optimized and calibrated with ground truth data. The…

Robotics · Computer Science 2025-10-06 Amir Hossein Barjini , Jouni Mattila

Human Activity Recognition (HAR) is a key building block of many emerging applications such as intelligent mobility, sports analytics, ambient-assisted living and human-robot interaction. With robust HAR, systems will become more…

Computer Vision and Pattern Recognition · Computer Science 2019-01-10 Mirco Moencks , Varuna De Silva , Jamie Roche , Ahmet Kondoz

Human activity recognition (HAR) by wearable sensor devices embedded in the Internet of things (IOT) can play a significant role in remote health monitoring and emergency notification, to provide healthcare of higher standards. The purpose…

Machine Learning · Computer Science 2022-01-24 M. Abid , A. Khabou , Y. Ouakrim , H. Watel , S. Chemkhi , A. Mitiche , A. Benazza-Benyahia , N. Mezghani

Multimodal human action recognition (HAR) leverages complementary sensors for activity classification. Beyond recognition, recent advances in large language models (LLMs) enable detailed descriptions and causal reasoning, motivating new…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Siyang Jiang , Mu Yuan , Xiang Ji , Bufang Yang , Zeyu Liu , Lilin Xu , Yang Li , Yuting He , Liran Dong , Wenrui Lu , Zhenyu Yan , Xiaofan Jiang , Wei Gao , Hongkai Chen , Guoliang Xing

This paper proposes a novel Subdivision-Fusion Model (SFM) to recognize human actions. In most action recognition tasks, overlapping feature distribution is a common problem leading to overfitting. In the subdivision stage of the proposed…

Computer Vision and Pattern Recognition · Computer Science 2015-08-19 Hao Zongbo , Lu Linlin , Zhang Qianni , Wu Jie , Izquierdo Ebroul , Yang Juanyu , Zhao Jun

Cross-domain generalization is very important in Time Series Forecasting because similar historical information may lead to distinct future trends due to the domain-specific characteristics. Recent works focus on building unimodal time…

Machine Learning · Computer Science 2026-03-10 Xingjian Wu , Jianxin Jin , Wanghui Qiu , Peng Chen , Yang Shu , Bin Yang , Chenjuan Guo

A central goal of artificial intelligence is to build systems that can understand and predict complex, evolving sequences of events. However, current foundation models, designed for natural language, fail to grasp the holistic nature of…

Machine Learning · Computer Science 2025-09-09 Vignesh Ethiraj , Subhash Talluri

Continuous emotion recognition in terms of valence and arousal under in-the-wild (ITW) conditions remains a challenging problem due to large variations in appearance, head pose, illumination, occlusions, and subject-specific patterns of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Elena Ryumina , Maxim Markitantov , Alexandr Axyonov , Dmitry Ryumin , Mikhail Dolgushin , Denis Dresvyanskiy , Alexey Karpov

The integration of information across multiple modalities and across time is a promising way to enhance the emotion recognition performance of affective systems. Much previous work has focused on instantaneous emotion recognition. The 2018…

Image and Video Processing · Electrical Eng. & Systems 2018-05-07 Didan Deng , Yuqian Zhou , Jimin Pi , Bertram E. Shi

Previous work has demonstrated that virtual accelerometry data, extracted from videos using cross-modality transfer approaches like IMUTube, is beneficial for training complex and effective human activity recognition (HAR) models. Systems…

Computer Vision and Pattern Recognition · Computer Science 2022-11-03 Zikang Leng , Yash Jain , Hyeokhyen Kwon , Thomas Plötz

Utilizing the sensor characteristics of the audio, visible camera, and thermal camera, the robustness of person recognition can be enhanced. Existing multimodal person recognition frameworks are primarily formulated assuming that multimodal…

Multimedia · Computer Science 2022-10-25 Vijay John , Yasutomo Kawanishi

Ensuring the safety and well-being of elderly and vulnerable populations in assisted living environments is a critical concern. Computer vision presents an innovative and powerful approach to predicting health risks through video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yixuan Wang , Paul Stynes , Pramod Pathak , Cristina Muntean

Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Thanh-Dat Truong , Christophe Bobda , Nitin Agarwal , Khoa Luu

Continuous Authentication (CA) using behavioural biometrics is a type of biometric identification that recognizes individuals based on their unique behavioural characteristics, like their typing style. However, the existing systems that use…

Cryptography and Security · Computer Science 2023-07-21 Dilshan Senerath , Sanuja Tharinda , Maduka Vishwajith , Sanka Rasnayaka , Sandareka Wickramanayake , Dulani Meedeniya

Objective: This research aims to develop a lifestyle intervention system, called MoveSense, that forecasts a patient's activity behavior to allow for early and personalized interventions in real-world clinical environments. Methods: We…

Machine Learning · Computer Science 2024-10-15 Abdullah Mamun , Krista S. Leonard , Megan E. Petrov , Matthew P. Buman , Hassan Ghasemzadeh

In this paper, we consider the problem of multimodal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality…

Computer Vision and Pattern Recognition · Computer Science 2022-01-27 Kateryna Chumachenko , Alexandros Iosifidis , Moncef Gabbouj

The main challenge of Multiple Object Tracking (MOT) is the efficiency in associating indefinite number of objects between video frames. Standard motion estimators used in tracking, e.g., Long Short Term Memory (LSTM), only deal with single…

Computer Vision and Pattern Recognition · Computer Science 2019-05-08 Jimuyang Zhang , Sanping Zhou , Jinjun Wang , Dong Huang

Facial Action Unit (AU) detection is a crucial task for emotion analysis from facial movements. The apparent differences of different subjects sometimes mislead changes brought by AUs, resulting in inaccurate results. However, most of the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Jiyuan Cao , Zhilei Liu , Yong Zhang

In this paper, we investigate the reliability of online recognition platforms, Amazon Rekognition and Microsoft Azure, with respect to changes in background, acquisition device, and object orientation. We focus on platforms that are…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Dogancan Temel , Jinsol Lee , Ghassan AlRegib

Facial action unit (AU) intensity is an index to describe all visually discernible facial movements. Most existing methods learn intensity estimator with limited AU data, while they lack generalization ability out of the dataset. In this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-21 Xinhui Song , Tianyang Shi , Zunlei Feng , Mingli Song , Jackie Lin , Chuanjie Lin , Changjie Fan , Yi Yuan