中文
相关论文

相关论文: A Dual-Masked Auto-Encoder for Robust Motion Captu…

200 篇论文

Building scalable models to learn from diverse, multimodal data remains an open challenge. For vision-language data, the dominant approaches are based on contrastive learning objectives that train a separate encoder for each modality. While…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Xinyang Geng , Hao Liu , Lisa Lee , Dale Schuurmans , Sergey Levine , Pieter Abbeel

Self-supervised learning has been a powerful training paradigm to facilitate representation learning. In this study, we design a masked autoencoder (MAE) to guide deep learning models to learn electroencephalography (EEG) signal…

人机交互 · 计算机科学 2024-09-04 Yifei Zhou , Sitong Liu

Reliable markerless motion tracking of people participating in a complex group activity from multiple moving cameras is challenging due to frequent occlusions, strong viewpoint and appearance variations, and asynchronous video streams. To…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Minh Vo , Ersin Yumer , Kalyan Sunkavalli , Sunil Hadap , Yaser Sheikh , Srinivasa Narasimhan

Skeleton-based action recognition, which classifies human actions based on the coordinates of joints and their connectivity within skeleton data, is widely utilized in various scenarios. While Graph Convolutional Networks (GCNs) have been…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Jeonghyeok Do , Munchurl Kim

Spatiotemporal forecasting techniques are significant for various domains such as transportation, energy, and weather. Accurate prediction of spatiotemporal series remains challenging due to the complex spatiotemporal heterogeneity. In…

机器学习 · 计算机科学 2024-10-01 Haotian Gao , Renhe Jiang , Zheng Dong , Jinliang Deng , Yuxin Ma , Xuan Song

Growing techniques have been emerging to improve the performance of passage retrieval. As an effective representation bottleneck pretraining technique, the contextual masked auto-encoder utilizes contextual embedding to assist in the…

计算与语言 · 计算机科学 2023-04-07 Xing Wu , Guangyuan Ma , Peng Wang , Meng Lin , Zijia Lin , Fuzheng Zhang , Songlin Hu

With the development of robotics, skeleton-based action recognition has become increasingly important, as human-robot interaction requires understanding the actions of humans and humanoid robots. Due to different sources of human skeletons…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Jidong Kuang , Hongsong Wang , Jie Gui

3D human mesh recovery from a 2D pose plays an important role in various applications. However, it is hard for existing methods to simultaneously capture the multiple relations during the evolution from skeleton to mesh, including…

计算机视觉与模式识别 · 计算机科学 2023-03-13 Yingxuan You , Hong Liu , Xia Li , Wenhao Li , Ti Wang , Runwei Ding

Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and disaster response. However, existing methods rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Shuhan Hu , Yiru Li , Yuanyuan Li , Yingying Zhu

Recent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such…

计算机视觉与模式识别 · 计算机科学 2023-08-24 Xiang Zhang , Huiyuan Yang , Taoyue Wang , Xiaotian Li , Lijun Yin

Zero-shot skeleton-based action recognition aims to develop models capable of identifying actions beyond the categories encountered during training. Previous approaches have primarily focused on aligning visual and semantic representations…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Wenhan Wu , Zhishuai Guo , Chen Chen , Hongfei Xue , Aidong Lu

Existing LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Zhiwei Lin , Yongtao Wang , Shengxiang Qi , Nan Dong , Ming-Hsuan Yang

This project addresses the challenge of human motion prediction, a critical area for applications such as au- tonomous vehicle movement detection. Previous works have emphasized the need for low inference times to provide real time…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Edmund Shieh , Joshua Lee Franco , Kang Min Bae , Tej Lalvani

Simultaneous recordings from thousands of neurons across multiple brain areas reveal rich mixtures of activity that are shared between regions and dynamics that are unique to each region. Existing alignment or multi-view methods neglect…

Training deep learning models for three-dimensional (3D) medical imaging, such as Computed Tomography (CT), is fundamentally challenged by the scarcity of labeled data. While pre-training on natural images is common, it results in a…

Recent progress in video diffusion models has markedly advanced character animation, which synthesizes motioned videos by animating a static identity image according to a driving video. Explicit methods represent motion using skeleton,…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Zhufeng Xu , Xuan Gao , Feng-Lin Liu , Haoxian Zhang , Zhixue Fang , Yu-Kun Lai , Xiaoqiang Liu , Pengfei Wan , Lin Gao

Knee Osteoarthritis (KOA) is a common musculoskeletal condition that significantly affects mobility and quality of life, particularly in elderly populations. However, training deep learning models for early KOA classification is often…

图像与视频处理 · 电气工程与系统科学 2025-01-17 Zhe Wang , Aladine Chetouani , Mohamed Jarraya , Yung Hsin Chen , Yuhua Ru , Fang Chen , Fabian Bauer , Liping Zhang , Didier Hans , Rachid Jennane

Current works focus on addressing the remote sensing change detection task using bi-temporal images. Although good performance can be achieved, however, seldom of they consider the motion cues which may also be vital. In this work, we…

计算机视觉与模式识别 · 计算机科学 2024-08-16 Xixi Wang , Zitian Wang , Jingtao Jiang , Lan Chen , Xiao Wang , Bo Jiang

Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framework that leverages…

Vision transformers have demonstrated significant advantages in computer vision tasks due to their ability to capture long-range dependencies and contextual relationships through self-attention. However, existing position encoding…

计算机视觉与模式识别 · 计算机科学 2025-05-15 Xi Chen , Shiyang Zhou , Muqi Huang , Jiaxu Feng , Yun Xiong , Kun Zhou , Biao Yang , Yuhui Zhang , Huishuai Bao , Sijia Peng , Chuan Li , Feng Shi