English
Related papers

Related papers: MoBind: Motion Binding for Fine-Grained IMU-Video …

200 papers

To achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data often suffer from the position shift problem, i.e., the image…

Computer Vision and Pattern Recognition · Computer Science 2022-04-22 Lu Zhang , Zhiyong Liu , Xiangyu Zhu , Zhan Song , Xu Yang , Zhen Lei , Hong Qiao

We simplify space binding by focusing on two core components, a single encoder per modality and high-quality data; enabling training state-of-the-art models on a single GPU in a few hours as opposed to multiple days. We present EBind, an…

Machine Learning · Computer Science 2025-11-19 Jim Broadbent , Felix Cohen , Frederik Hvilshøj , Eric Landau , Eren Sasoglu

As cameras and inertial sensors are becoming ubiquitous in mobile devices and robots, it holds great potential to design visual-inertial navigation systems (VINS) for efficient versatile 3D motion tracking which utilize any (multiple)…

Robotics · Computer Science 2020-06-30 Kevin Eckenhoff , Patrick Geneva , Guoquan Huang

We propose a new contrastive objective for learning overcomplete pixel-level features that are invariant to motion blur. Other invariances (e.g., pose, illumination, or weather) can be learned by applying the corresponding transformations…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 Leonid Pogorelyuk , Stefan T. Radev

Employing an inertial measurement unit (IMU) as an additional sensor can dramatically improve both reliability and accuracy of visual/Lidar odometry (VO/LO). Different IMU integration models are introduced using different assumptions on the…

Robotics · Computer Science 2019-12-03 John Henawy , Zhengguo Li , Wei Yun Yau , Gerald Seet , Kong Wah Wan

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the inter-modality…

Computer Vision and Pattern Recognition · Computer Science 2023-05-17 Heqing Zou , Meng Shen , Chen Chen , Yuchen Hu , Deepu Rajan , Eng Siong Chng

Visual localization, i.e., determining the position and orientation of a vehicle with respect to a map, is a key problem in autonomous driving. We present a multicamera visual inertial localization algorithm for large scale environments. To…

Robotics · Computer Science 2019-05-16 Marcel Geppert , Peidong Liu , Zhaopeng Cui , Marc Pollefeys , Torsten Sattler

IMU-based gesture interfaces are being increasingly adopted as efficient, accessible, and intuitive alternatives to traditional input methods, such as touchscreens and voice. However, current gesture recognition algorithms are tailored to…

Human-Computer Interaction · Computer Science 2026-03-13 Prerna Khanna , Tanmay Srivastava , Shubham Jain , Aruna Balasubramanian

Bundle adjustment jointly optimizes camera intrinsics and extrinsics and 3D point triangulation to reconstruct a static scene. The triangulation constraint, however, is invalid for moving points captured in multiple unsynchronized videos…

Computer Vision and Pattern Recognition · Computer Science 2020-07-28 Minh Vo , Yaser Sheikh , Srinivasa G. Narasimhan

Human motion prediction is an essential component for enabling closer human-robot collaboration. The task of accurately predicting human motion is non-trivial. It is compounded by the variability of human motion, both at a skeletal level…

Robotics · Computer Science 2021-07-02 Mohammad Samin Yasar , Tariq Iqbal

Unsupervised large-scale vision-language pre-training has shown promising advances on various downstream tasks. Existing methods often model the cross-modal interaction either via the similarity of the global feature of each modality which…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Lewei Yao , Runhui Huang , Lu Hou , Guansong Lu , Minzhe Niu , Hang Xu , Xiaodan Liang , Zhenguo Li , Xin Jiang , Chunjing Xu

The ability to estimate 3D movements of users over edge computing-enabled networks, such as 5G/6G networks, is a key enabler for the new era of extended reality (XR) and Metaverse applications. Recent advancements in deep learning have…

Signal Processing · Electrical Eng. & Systems 2024-09-04 Nguyen Quang Hieu , Dinh Thai Hoang , Diep N. Nguyen

Over the past decade, lidars have become a cornerstone of robotics state estimation and perception thanks to their ability to provide accurate geometric information about their surroundings in the form of 3D scans. Unfortunately, most of…

Robotics · Computer Science 2024-10-08 Cedric Le Gentil , Raphael Falque , Teresa Vidal-Calleja

LiDAR-based 3D mapping suffers from cumulative drift causing global misalignment, particularly in GNSS-constrained environments. To address this, we propose a unified framework that fuses LiDAR, GNSS, and IMU data for high-resolution…

This paper presents an inertial sensor aided technique for beam alignment and tracking in massive multiple-input multiple-output (MIMO) vehicle-to-vehicle (V2V) communications based on millimeter waves (mmWave). Since directional…

Signal Processing · Electrical Eng. & Systems 2019-07-18 Mattia Brambilla , Monica Nicoli , Sergio Savaresi , Umberto Spagnolini

This paper proposes a novel inertial-aided localization approach by fusing information from multiple inertial measurement units (IMUs) and exteroceptive sensors. IMU is a low-cost motion sensor which provides measurements on angular…

Robotics · Computer Science 2020-01-20 Ming Zhang , Yiming Chen , Xiangyu Xu , Mingyang Li

Motion in-betweening (MIB) is a process of generating intermediate skeletal movement between the given start and target poses while preserving the naturalness of the motion, such as periodic footstep motion while walking. Although…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Jihoon Kim , Taehyun Byun , Seungyoun Shin , Jungdam Won , Sungjoon Choi

With the rapid development of wearable technology, devices like smartphones, smartwatches, and headphones equipped with IMUs have become essential for applications such as pedestrian positioning. However, traditional pedestrian dead…

Machine Learning · Computer Science 2024-11-13 Lan Sun , Songpengcheng Xia , Junyuan Deng , Jiarui Yang , Zengyuan Lai , Qi Wu , Ling Pei

In this paper we present an on-manifold sequence-to-sequence learning approach to motion estimation using visual and inertial sensors. It is to the best of our knowledge the first end-to-end trainable method for visual-inertial odometry…

Computer Vision and Pattern Recognition · Computer Science 2017-04-04 Ronald Clark , Sen Wang , Hongkai Wen , Andrew Markham , Niki Trigoni

Modern multi-object tracking (MOT) system usually involves separated modules, such as motion model for location and appearance model for data association. However, the compatible problems within both motion and appearance models are always…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Piao Huang , Shoudong Han , Jun Zhao , Donghaisheng Liu , Hongwei Wang , En Yu , Alex ChiChung Kot