English
Related papers

Related papers: Multimodal Fusion Using Deep Learning Applied to D…

200 papers

Cooperative autonomous driving plays a pivotal role in improving road capacity and safety within intelligent transportation systems, particularly through the deployment of autonomous vehicles on urban streets. By enabling vehicle-to-vehicle…

Robotics · Computer Science 2023-12-13 Ahmed Abdelrahman , Omar M. Shehata , Yarah Basyoni , Elsayed I. Morgan

We introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate…

Computer Vision and Pattern Recognition · Computer Science 2016-03-25 Wael AbdAlmageed , Yue Wua , Stephen Rawlsa , Shai Harel , Tal Hassner , Iacopo Masi , Jongmoo Choi , Jatuporn Toy Leksut , Jungyeon Kim , Prem Natarajan , Ram Nevatia , Gerard Medioni

A fundamental challenge in car-following modeling lies in accurately representing the multi-scale complexity of driving behaviors, particularly the intra-driver heterogeneity where a single driver's actions fluctuate dynamically under…

Machine Learning · Computer Science 2025-06-09 Shirui Zhou , Jiying Yan , Junfang Tian , Tao Wang , Yongfu Li , Shiquan Zhong

This paper presents a novel vehicle motion forecasting method based on multi-head attention. It produces joint forecasts for all vehicles on a road scene as sequences of multi-modal probability density functions of their positions. Its…

Machine Learning · Computer Science 2019-12-23 Jean Mercat , Thomas Gilles , Nicole El Zoghby , Guillaume Sandou , Dominique Beauvois , Guillermo Pita Gil

The dynamic hand gesture recognition task has seen studies on various unimodal and multimodal methods. Previously, researchers have explored depth and 2D-skeleton-based multimodal fusion CRNNs (Convolutional Recurrent Neural Networks) but…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Hasan Mahmud , Mashrur M. Morshed , Md. Kamrul Hasan

As humans, we experience the world with all our senses or modalities (sound, sight, touch, smell, and taste). We use these modalities, particularly sight and touch, to convey and interpret specific meanings. Multimodal expressions are…

Machine Learning · Computer Science 2022-05-17 Anirudh Sundar , Larry Heck

The use of multi-modal data for deep machine learning has shown promise when compared to uni-modal approaches with fusion of multi-modal features resulting in improved performance in several applications. However, most state-of-the-art…

Machine Learning · Computer Science 2020-10-26 Darshana Priyasad , Tharindu Fernando , Simon Denman , Sridha Sridharan , Clinton Fookes

A motion-based control interface promises flexible robot operations in dangerous environments by combining user intuitions with the robot's motor capabilities. However, designing a motion interface for non-humanoid robots, such as…

Robotics · Computer Science 2022-04-29 Sunwoo Kim , Maks Sorokin , Jehee Lee , Sehoon Ha

We study the problem of learning physical object representations for robot manipulation. Understanding object physics is critical for successful object manipulation, but also challenging because physical object properties can rarely be…

Robotics · Computer Science 2019-06-13 Zhenjia Xu , Jiajun Wu , Andy Zeng , Joshua B. Tenenbaum , Shuran Song

Mental rotation -- the ability to compare objects seen from different viewpoints -- is a fundamental example of mental simulation and spatial world modeling in humans. Here we propose a mechanistic model of human mental rotation, leveraging…

Neurons and Cognition · Quantitative Biology 2026-05-29 Raymond Khazoum , Daniela Fernandes , Aleksandr Krylov , Qin Li , Stephane Deny

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Wangyuan Zhu , Jun Yu

Attention (and distraction) recognition is a key factor in improving human-robot collaboration. We present an assembly scenario where a human operator and a cobot collaborate equally to piece together a gearbox. The setup provides multiple…

Human-Computer Interaction · Computer Science 2023-04-03 Pooja Prajod , Matteo Lavit Nicora , Matteo Malosio , Elisabeth André

Learning methods for relative camera pose estimation have been developed largely in isolation from classical geometric approaches. The question of how to integrate predictions from deep neural networks (DNNs) and solutions from geometric…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Bingbing Zhuang , Manmohan Chandraker

In the context of deep learning, this article presents an original deep network, namely CentralNet, for the fusion of information coming from different sensors. This approach is designed to efficiently and automatically balance the…

Computer Vision and Pattern Recognition · Computer Science 2018-11-07 Valentin Vielzeuf , Alexis Lechervy , Stéphane Pateux , Frédéric Jurie

Machine learning is finding increasingly broad application in the physical sciences. This most often involves building a model relationship between a dependent, measurable output and an associated set of controllable, but complicated,…

Computational Physics · Physics 2018-08-29 Brian K. Spears

One of the most universal ways that people communicate is through facial expressions. In this paper, we take a deep dive, implementing multiple deep learning models for facial expression recognition (FER). Our goals are twofold: we aim not…

Computer Vision and Pattern Recognition · Computer Science 2020-04-27 Amil Khanzada , Charles Bai , Ferhat Turker Celepcikay

This work presents a probabilistic deep neural network that combines LiDAR point clouds and RGB camera images for robust, accurate 3D object detection. We explicitly model uncertainties in the classification and regression tasks, and…

Robotics · Computer Science 2020-02-04 Di Feng , Yifan Cao , Lars Rosenbaum , Fabian Timm , Klaus Dietmayer

Image-text multimodal representation learning aligns data across modalities and enables important medical applications, e.g., image classification, visual grounding, and cross-modal retrieval. In this work, we establish a connection between…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Peiqi Wang , William M. Wells , Seth Berkowitz , Steven Horng , Polina Golland

This paper explores the capability of deep neural networks to capture key characteristics of vehicle dynamics, and their ability to perform coupled longitudinal and lateral control of a vehicle. To this extent, two different artificial…

Machine Learning · Computer Science 2018-10-23 Guillaume Devineau , Philip Polack , Florent Altché , Fabien Moutarde

Modern industrial recommendation systems improve recommendation performance by integrating multimodal representations from pre-trained models into ID-based Click-Through Rate (CTR) prediction frameworks. However, existing approaches…

Information Retrieval · Computer Science 2026-04-17 Alin Fan , Hanqing Li , Sihan Lu , Jingsong Yuan , Jiandong Zhang