中文
相关论文

相关论文: MoBind: Motion Binding for Fine-Grained IMU-Video …

200 篇论文

In remote sensing, we are interested in modeling various modalities for some geographic location. Several works have focused on learning the relationship between a location and type of landscape, habitability, audio, textual descriptions,…

人工智能 · 计算机科学 2024-04-19 Aayush Dhakal , Subash Khanal , Srikumar Sastry , Adeel Ahmad , Nathan Jacobs

Fusing LiDAR and camera information is essential for achieving accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features…

计算机视觉与模式识别 · 计算机科学 2023-03-06 Yang Jiao , Zequn Jie , Shaoxiang Chen , Jingjing Chen , Lin Ma , Yu-Gang Jiang

While Contrastive Language-Image Pretraining (CLIP) excels at zero-shot tasks by aligning image and text embeddings, its performance in few-shot classification is hindered by a critical limitation: intra-modal misalignment. This issue,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Christoph Timmermann , Hyunse Lee , Woojin Lee

Micro-gesture recognition and behavior-based emotion prediction are both highly challenging tasks that require modeling subtle, fine-grained human behaviors, primarily leveraging video and skeletal pose data. In this work, we present two…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Arman Martirosyan , Shahane Tigranyan , Maria Razzhivina , Artak Aslanyan , Nazgul Salikhova , Ilya Makarov , Andrey Savchenko , Aram Avetisyan

Inertial-based Motion capture system has been attracting growing attention due to its wearability and unsconstrained use. However, accurate human joint estimation demands several complex and expertise demanding steps, which leads to…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Sara M. Cerqueira , Manuel Palermo , Cristina P. Santos

The key challenge in unaligned multimodal language sequences lies in effectively integrating information from various modalities to obtain a refined multimodal joint representation. Recently, the disentangle and fuse methods have achieved…

计算与语言 · 计算机科学 2024-09-20 Fan Qian , Jiqing Han , Jianchen Li , Yongjun He , Tieran Zheng , Guibin Zheng

In video analysis, background models have many applications such as background/foreground separation, change detection, anomaly detection, tracking, and more. However, while learning such a model in a video captured by a static camera is a…

计算机视觉与模式识别 · 计算机科学 2022-09-19 Guy Erez , Ron Shapira Weber , Oren Freifeld

We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio,…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Srikumar Sastry , Subash Khanal , Aayush Dhakal , Adeel Ahmad , Nathan Jacobs

Recent advancements in technology have brought forth new forms of interactive applications, such as the social metaverse, where end users interact with each other through their virtual avatars. In such applications, precise full-body…

图形学 · 计算机科学 2023-12-11 Deok-Kyeong Jang , Dongseok Yang , Deok-Yun Jang , Byeoli Choi , Taeil Jin , Sung-Hee Lee

Face recognition is a crucial task in various multimedia applications such as security check, credential access and motion sensing games. However, the task is challenging when an input face is noisy (e.g. poor-condition RGB image) or lacks…

计算机视觉与模式识别 · 计算机科学 2021-12-09 Wenbin Teng , Chongyang Bai

Human motion generation is essential for fields such as animation, robotics, and virtual reality, requiring models that effectively capture motion dynamics from text descriptions. Existing approaches often rely on Contrastive Language-Image…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Gabriel Maldonado , Armin Danesh Pazho , Ghazal Alinezhad Noghre , Vinit Katariya , Hamed Tabkhi

We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SimHand. Pre-training with large-scale images achieves promising results in various tasks, but…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Nie Lin , Takehiko Ohkawa , Yifei Huang , Mingfang Zhang , Minjie Cai , Ming Li , Ryosuke Furuta , Yoichi Sato

Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Jie Zhao , Xin Chen , Yongsheng Yuan , Michael Felsberg , Dong Wang , Huchuan Lu

In many robotics and VR/AR applications, fast camera motions lead to a high level of motion blur, causing existing camera pose estimation methods to fail. In this work, we propose a novel framework that leverages motion blur as a rich cue…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Jerred Chen , Ronald Clark

Temporal 3D human pose estimation from monocular videos is a challenging task in human-centered computer vision due to the depth ambiguity of 2D-to-3D lifting. To improve accuracy and address occlusion issues, inertial sensor has been…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Yiming Bao , Xu Zhao , Dahong Qian

Real-time object pose estimation and tracking is challenging but essential for emerging augmented reality (AR) applications. In general, state-of-the-art methods address this problem using deep neural networks which indeed yield…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Yo-Chung Lau , Kuan-Wei Tseng , I-Ju Hsieh , Hsiao-Ching Tseng , Yi-Ping Hung

Human pose estimation in videos has long been a compelling yet challenging task within the realm of computer vision. Nevertheless, this task remains difficult because of the complex video scenes, such as video defocus and self-occlusion.…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Sifan Wu , Haipeng Chen , Yifang Yin , Sihao Hu , Runyang Feng , Yingying Jiao , Ziqi Yang , Zhenguang Liu

Recent years have witnessed the rapid development of short videos, which usually contain both visual and audio modalities. Background music is important to the short videos, which can significantly influence the emotions of the viewers.…

多媒体 · 计算机科学 2024-05-16 Jiajie Teng , Huiyu Duan , Yucheng Zhu , Sijing Wu , Guangtao Zhai

Hyperkinetic movement disorders (HMDs) such as dystonia, tremor, chorea, myoclonus, and tics are disabling motor manifestations across childhood and adulthood. Their fluctuating, intermittent, and frequently co-occurring expressions hinder…

Multi-modal object Re-IDentification (ReID) is devoted to retrieving specific objects through the exploitation of complementary multi-modal image information. Existing methods mainly concentrate on the fusion of multi-modal features, yet…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yangyang Liu , Yuhao Wang , Pingping Zhang