中文
相关论文

相关论文: MoBind: Motion Binding for Fine-Grained IMU-Video …

200 篇论文

We introduce a novel method for controlling a motion sequence using an arbitrary temporal control sequence using temporal alignment. Temporal alignment of motion has gained significant attention owing to its applications in motion control…

图形学 · 计算机科学 2025-11-26 Naoki Agata , Takeo Igarashi

The use of a wide range of computer vision solutions, and more recently high-end Inertial Measurement Units (IMU) have become increasingly popular for assessing human physical activity in clinical and research settings. Nevertheless, to…

Recent advances in multimodal foundation models have achieved state-of-the-art performance across a range of tasks. These breakthroughs are largely driven by new pre-training paradigms that leverage large-scale, unlabeled multimodal data,…

机器学习 · 计算机科学 2025-06-10 Xiaojun Shan , Qi Cao , Xing Han , Haofei Yu , Paul Pu Liang

Deploying depth estimation networks in the real world requires high-level robustness against various adverse conditions to ensure safe and reliable autonomy. For this purpose, many autonomous vehicles employ multi-modal sensor systems,…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Ukcheol Shin , Kyunghyun Lee , Jean Oh

Multi-label image classification is a foundational topic in various domains. Multimodal learning approaches have recently achieved outstanding results in image representation and single-label image classification. For instance, Contrastive…

计算机视觉与模式识别 · 计算机科学 2022-11-11 Fengjun Wang , Sarai Mizrachi , Moran Beladev , Guy Nadav , Gil Amsalem , Karen Lastmann Assaraf , Hadas Harush Boker

The multi-modal perception methods are thriving in the autonomous driving field due to their better usage of complementary data from different sensors. Such methods depend on calibration and synchronization between sensors to get accurate…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Zhihang Song , Lihui Peng , Jianming Hu , Danya Yao , Yi Zhang

Leveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Yuhan Liu , Jingwen Fu , Yang Wu , Kangyi Wu , Pengna Li , Jiayi Wu , Sanping Zhou , Jingmin Xin

Inertial motion capture is a promising approach for capturing motion outside the laboratory. However, as one major drawback, most of the current methods require different quantities to be calibrated or computed offline as part of the setup…

机器人学 · 计算机科学 2025-09-17 Michael Lorenz , Bertram Taetz , Gabriele Bleser-Taetz , Didier Stricker

In this paper, we propose a tightly-coupled, multi-modal simultaneous localization and mapping (SLAM) framework, integrating an extensive set of sensors: IMU, cameras, multiple lidars, and Ultra-wideband (UWB) range measurements, hence…

机器人学 · 计算机科学 2021-10-06 Thien-Minh Nguyen , Shenghai Yuan , Muqing Cao , Thien Hoang Nguyen , Lihua Xie

Understanding neural activity and information representation is crucial for advancing knowledge of brain function and cognition. Neural activity, measured through techniques like electrophysiology and neuroimaging, reflects various aspects…

神经元与认知 · 定量生物学 2024-07-22 Fengyu Yang , Chao Feng , Daniel Wang , Tianye Wang , Ziyao Zeng , Zhiyang Xu , Hyoungseob Park , Pengliang Ji , Hanbin Zhao , Yuanning Li , Alex Wong

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Accurate moving object segmentation is an essential task for autonomous driving. It can provide effective information for many downstream tasks, such as collision avoidance, path planning, and static map construction. How to effectively…

计算机视觉与模式识别 · 计算机科学 2022-07-06 Jiadai Sun , Yuchao Dai , Xianjing Zhang , Jintao Xu , Rui Ai , Weihao Gu , Xieyuanli Chen

We present Reusable Motion prior (ReMP), an effective motion prior that can accurately track the temporal evolution of motion in various downstream tasks. Inspired by the success of foundation models, we argue that a robust spatio-temporal…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Hojun Jang , Young Min Kim

Multi-view inverse rendering aims to recover geometry, materials, and illumination consistently across multiple viewpoints. When applied to multi-view images, existing single-view approaches often ignore cross-view relationships, leading to…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Xiangzuo Wu , Chengwei Ren , Jun Zhou , Xiu Li , Yuan Liu

Visual-inertial sensors have a wide range of applications in robotics. However, good performance often requires different sophisticated motion routines to accurately calibrate camera intrinsics and inter-sensor extrinsics. This work…

机器人学 · 计算机科学 2021-10-01 Yunke Ao , Le Chen , Florian Tschopp , Michel Breyer , Andrei Cramariuc , Roland Siegwart

Recent advancements in foundation models for 2D vision have substantially improved the analysis of dynamic scenes from monocular videos. However, despite their strong generalization capabilities, these models often lack 3D consistency, a…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Haoran Zhou , Gim Hee Lee

Recent adaptive methods for efficient video recognition mostly follow the two-stage paradigm of "preview-then-recognition" and have achieved great success on multiple video benchmarks. However, this two-stage paradigm involves two visits of…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Ye Tian , Mengyu Yang , Lanshan Zhang , Zhizhen Zhang , Yang Liu , Xiaohui Xie , Xirong Que , Wendong Wang

Ego-motion estimation is a fundamental requirement for most mobile robotic applications. By sensor fusion, we can compensate the deficiencies of stand-alone sensors and provide more reliable estimations. We introduce a tightly coupled…

机器人学 · 计算机科学 2019-08-30 Haoyang Ye , Yuying Chen , Ming Liu

KBody is a method for fitting a low-dimensional body model to an image. It follows a predict-and-optimize approach, relying on data-driven model estimates for the constraints that will be used to solve for the body's parameters.…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Nikolaos Zioulis , James F. O'Brien

Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-target pairs across different modalities. Yet, despite its…