English
Related papers

Related papers: Geometric Cross-Modal Comparison of Heterogeneous …

200 papers

Multi-modal face anti-spoofing (FAS) aims to detect genuine human presence by extracting discriminative liveness cues from multiple modalities, such as RGB, infrared (IR), and depth images, to enhance the robustness of biometric…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Jun-Xiong Chong , Fang-Yu Hsu , Ming-Tsung Hsu , Yi-Ting Lin , Kai-Heng Chien , Chiou-Ting Hsu , Pei-Kai Huang

People are increasingly concerned with understanding their personal environment, including possible exposure to harmful air pollutants. In order to make informed decisions on their day-to-day activities, they are interested in real-time…

Despite the growing adoption of radar in robotics, the majority of research has been confined to homogeneous sensor types, overlooking the integration and cross-modality challenges inherent in heterogeneous radar technologies. This leads to…

Robotics · Computer Science 2025-10-13 Hanjun Kim , Minwoo Jung , Wooseong Yang , Ayoung Kim

Merge trees, a type of topological descriptor, serve to identify and summarize the topological characteristics associated with scalar fields. They present a great potential for the analysis and visualization of time-varying data. First,…

Human-Computer Interaction · Computer Science 2021-08-02 Lin Yan , Talha Bin Masood , Farhan Rasheed , Ingrid Hotz , Bei Wang

3D multi-object tracking and trajectory prediction are two crucial modules in autonomous driving systems. Generally, the two tasks are handled separately in traditional paradigms and a few methods have started to explore modeling these two…

Computer Vision and Pattern Recognition · Computer Science 2024-07-01 Jiaheng Zhuang , Guoan Wang , Siyu Zhang , Xiyang Wang , Hangning Zhou , Ziyao Xu , Chi Zhang , Zhiheng Li

Optical-flow-based and kernel-based approaches have been extensively explored for temporal compensation in satellite Video Super-Resolution (VSR). However, these techniques are less generalized in large-scale or complex scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Yi Xiao , Qiangqiang Yuan , Kui Jiang , Xianyu Jin , Jiang He , Liangpei Zhang , Chia-Wen Lin

Vision-based localization is a cost-effective and thus attractive solution for many intelligent mobile platforms. However, its accuracy and especially robustness still suffer from low illumination conditions, illumination changes, and…

Robotics · Computer Science 2024-01-17 Yi-Fan Zuo , Wanting Xu , Xia Wang , Yifu Wang , Laurent Kneip

Most of the existing self-supervised feature learning methods for 3D data either learn 3D features from point cloud data or from multi-view images. By exploring the inherent multi-modality attributes of 3D objects, in this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2020-05-29 Longlong Jing , Yucheng Chen , Ling Zhang , Mingyi He , Yingli Tian

Mobile mapping, in particular, Mobile Lidar Scanning (MLS) is increasingly widespread to monitor and map urban scenes at city scale with unprecedented resolution and accuracy. The resulting point cloud sampling of the scene geometry can be…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Teng Wu , Bruno Vallet , Cédric Demonceaux

Traffic forecasting is a fundamental task in transportation research, however the scope of current research has mainly focused on a single data modality of loop detectors. Recently, the advances in Artificial Intelligence and drone…

Machine Learning · Computer Science 2025-04-29 Weijiang Xiong , Robert Fonod , Alexandre Alahi , Nikolas Geroliminis

Cross-Modal Retrieval (CMR), which retrieves relevant items from one modality (e.g., audio) given a query in another modality (e.g., visual), has undergone significant advancements in recent years. This capability is crucial for robots to…

Robotics · Computer Science 2024-07-31 Jagoda Wojcik , Jiaqi Jiang , Jiacheng Wu , Shan Luo

Current cross-modal retrieval systems are evaluated using R@K measure which does not leverage semantic relationships rather strictly follows the manually marked image text query pairs. Therefore, current systems do not generalize well for…

Computer Vision and Pattern Recognition · Computer Science 2019-09-06 Shah Nawaz , Muhammad Kamran Janjua , Ignazio Gallo , Arif Mahmood , Alessandro Calefati , Faisal Shafait

Accurate online multiple-camera vehicle tracking is essential for intelligent transportation systems, autonomous driving, and smart city applications. Like single-camera multiple-object tracking, it is commonly formulated as a graph problem…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Fabian Herzog , Johannes Gilg , Philipp Wolters , Torben Teepe , Gerhard Rigoll

Geodesic distance serves as a reliable means of measuring distance in nonlinear spaces, and such nonlinear manifolds are prevalent in the current multimodal learning. In these scenarios, some samples may exhibit high similarity, yet they…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Shibin Mei , Hang Wang , Bingbing Ni

Modeling multi-modal time-series data is critical for capturing system-level dynamics, particularly in biosignals where modalities such as ECG, PPG, EDA, and accelerometry provide complementary perspectives on interconnected physiological…

Machine Learning · Computer Science 2025-10-14 Wanting Mao , Maxwell A Xu , Harish Haresamudram , Mithun Saha , Santosh Kumar , James Matthew Rehg

Gesture recognition is one of the most intuitive ways of interaction and has gathered particular attention for human computer interaction. Radar sensors possess multiple intrinsic properties, such as their ability to work in low…

Signal Processing · Electrical Eng. & Systems 2022-05-20 Souvik Hazra , Hao Feng , Gamze Naz Kiprit , Michael Stephan , Lorenzo Servadei , Robert Wille , Robert Weigel , Avik Santra

The core problem of visual multi-robot simultaneous localization and mapping (MR-SLAM) is how to efficiently and accurately perform multi-robot global localization (MR-GL). The difficulties are two-fold. The first is the difficulty of…

Robotics · Computer Science 2021-02-25 Xiyue Guo , Junjie Hu , Junfeng Chen , Fuqin Deng , Tin Lun Lam

Incorporating multi-modal features as side information has recently become a trend in recommender systems. To elucidate user-item preferences, recent studies focus on fusing modalities via concatenation, element-wise sum, or attention…

Information Retrieval · Computer Science 2024-12-20 Rongqing Kenneth Ong , Andy W. H. Khong

Learning common subspace is prevalent way in cross-modal retrieval to solve the problem of data from different modalities having inconsistent distributions and representations that cannot be directly compared. Previous cross-modal retrieval…

Multimedia · Computer Science 2021-10-27 Donghuo Zeng , Jianming Wu , Gen Hattori , Yi Yu , Rong Xu

We investigate the challenging problem of integrating detection, signal processing, target tracking, and adaptive waveform scheduling with lookahead in urban terrain. We propose a closed-loop active sensing system to address this problem by…

Signal Processing · Electrical Eng. & Systems 2019-03-26 Patricia R. Barbosa , Yugandhar Sarkale , Edwin K. P. Chong , Yun Li , Sofia Suvorova , Bill Moran