English
Related papers

Related papers: TransFusionOdom: Interpretable Transformer-based L…

200 papers

Managing fluid balance in dialysis patients is crucial, as improper management can lead to severe complications. In this paper, we propose a multimodal approach that integrates visual features from lung ultrasound images with clinical data…

Image and Video Processing · Electrical Eng. & Systems 2024-10-04 Tianqi Yang , Nantheera Anantrasirichai , Oktay Karakuş , Marco Allinovi , Alin Achim

The main idea of multimodal recommendation is the rational utilization of the item's multimodal information to improve the recommendation performance. Previous works directly integrate item multimodal features with item ID embeddings,…

Information Retrieval · Computer Science 2023-04-25 Yan Zhou , Jie Guo , Hao Sun , Bin Song , Fei Richard Yu

Multi-modal learning has been intensified in recent years, especially for applications in facial analysis and action unit detection whilst there still exist two main challenges in terms of 1) relevant feature learning for representation and…

Computer Vision and Pattern Recognition · Computer Science 2022-03-23 Xiang Zhang , Lijun Yin

Current tomographic imaging systems need major improvements, especially when multi-dimensional, multi-scale, multi-temporal and multi-parametric phenomena are under investigation. Both preclinical and clinical imaging now depend on in vivo…

Fusion technique is a key research topic in multimodal sentiment analysis. The recent attention-based fusion demonstrates advances over simple operation-based fusion. However, these fusion works adopt single-scale, i.e., token-level or…

Computation and Language · Computer Science 2021-12-03 Huaishao Luo , Lei Ji , Yanyong Huang , Bin Wang , Shenggong Ji , Tianrui Li

Compared with real-time multi-object tracking (MOT), offline multi-object tracking (OMOT) has the advantages to perform 2D-3D detection fusion, erroneous link correction, and full track optimization but has to deal with the challenges from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Kemiao Huang , Yinqi Chen , Meiying Zhang , Qi Hao

Multi-sensor fusion of multi-modal measurements from commodity inertial, visual and LiDAR sensors to provide robust and accurate 6DOF pose estimation holds great potential in robotics and beyond. In this paper, building upon our prior work…

Robotics · Computer Science 2020-08-18 Xingxing Zuo , Yulin Yang , Patrick Geneva , Jiajun Lv , Yong Liu , Guoquan Huang , Marc Pollefeys

This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the complex interactions…

Computation and Language · Computer Science 2025-01-15 Hui Lee , Singh Suniljit , Yong Siang Ong

Multimodal fusion focuses on integrating information from multiple modalities with the goal of more accurate prediction, which has achieved remarkable progress in a wide range of scenarios, including autonomous driving and medical…

Machine Learning · Computer Science 2024-11-04 Qingyang Zhang , Yake Wei , Zongbo Han , Huazhu Fu , Xi Peng , Cheng Deng , Qinghua Hu , Cai Xu , Jie Wen , Di Hu , Changqing Zhang

In this letter, we propose a robust, real-time tightly-coupled multi-sensor fusion framework, which fuses measurement from LiDAR, inertial sensor, and visual camera to achieve robust and accurate state estimation. Our proposed framework is…

Robotics · Computer Science 2021-02-25 Jiarong Lin , Chunran Zheng , Wei Xu , Fu Zhang

The primary value of infrared and visible image fusion technology lies in applying the fusion results to downstream tasks. However, existing methods face challenges such as increased training complexity and significantly compromised…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Zengyi Yang , Yafei Zhang , Huafeng Li , Yu Liu

Humanoid robot technology is advancing rapidly, with manufacturers introducing diverse heterogeneous visual perception modules tailored to specific scenarios. Among various perception paradigms, occupancy-based representation has become…

Despite the number of works published in recent years, vehicle localization remains an open, challenging problem. While map-based localization and SLAM algorithms are getting better and better, they remain a single point of failure in…

Robotics · Computer Science 2024-03-21 Luca Mozzarelli , Luca Cattaneo , Matteo Corno , Sergio Matteo Savaresi

Researchers have increasingly adopted Transformer-based models for inertial odometry. While Transformers excel at modeling long-range dependencies, their limited sensitivity to local, fine-grained motion variations and lack of inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Shanshan Zhang , Qi Zhang , Siyue Wang , Tianshui Wen , Liqin Wu , Ziheng Zhou , Xuemin Hong , Ao Peng , Lingxiang Zheng , Yu Yang

Sensor fusion is an essential topic in many perception systems, such as autonomous driving and robotics. Transformers-based detection head and CNN-based feature encoder to extract features from raw sensor-data has emerged as one of the best…

Computer Vision and Pattern Recognition · Computer Science 2023-02-23 Apoorv Singh

Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ruoyu Wang , Wenqian Wang , Jianjun Gao , Dan Lin , Kim-Hui Yap , Bingbing Li

In this paper we present an on-manifold sequence-to-sequence learning approach to motion estimation using visual and inertial sensors. It is to the best of our knowledge the first end-to-end trainable method for visual-inertial odometry…

Computer Vision and Pattern Recognition · Computer Science 2017-04-04 Ronald Clark , Sen Wang , Hongkai Wen , Andrew Markham , Niki Trigoni

LiDAR-camera 3D multi-object tracking (MOT) combines rich visual semantics with accurate depth cues to improve trajectory consistency and tracking reliability. In practice, however, LiDAR and cameras operate at different sampling rates. To…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Xian Wu , Yitao Wu , Xiaoyu Li , Zijia Li , Lijun Zhao , Lining Sun

Unsupervised object discovery (UOD) has recently shown encouraging progress with the adoption of pre-trained Transformer features. However, current methods based on Transformers mainly focus on designing the localization head (e.g., seed…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Zhiwei Lin , Zengyu Yang , Yongtao Wang

Besides standard cameras, autonomous vehicles typically include multiple additional sensors, such as lidars and radars, which help acquire richer information for perceiving the content of the driving scene. While several recent works focus…

Computer Vision and Pattern Recognition · Computer Science 2023-08-14 Tim Broedermann , Christos Sakaridis , Dengxin Dai , Luc Van Gool