English
Related papers

Related papers: MoPFormer: Motion-Primitive Transformer for Wearab…

200 papers

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, audio-dense} design --…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Detao Bai , Shimin Yao , Weixuan Chen , Chengen Lai , Yuanming Li , Zhiheng Ma , Xihan Wei

Human activity recognition~(HAR) has attracted significant research interest due to its applications in health monitoring and patient rehabilitation. Recent research on HAR focuses on using smartphones due to their widespread use. However,…

Computer Vision and Pattern Recognition · Computer Science 2019-02-06 Ganapati Bhat , Ranadeep Deb , Vatika Vardhan Chaurasia , Holly Shill , Umit Y. Ogras

The mainstream human activity recognition (HAR) algorithms are developed based on RGB cameras, which are easily influenced by low-quality images (e.g., low illumination, motion blur). Meanwhile, the privacy protection issue caused by…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Haoxiang Yang , Chengguo Yuan , Yabin Zhu , Lan Chen , Xiao Wang , Futian Wang

Detecting the openable parts of articulated objects is crucial for downstream applications in intelligent robotics, such as pulling a drawer. This task poses a multitasking challenge due to the necessity of understanding object categories…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Siqi Li , Xiaoxue Chen , Haoyu Cheng , Guyue Zhou , Hao Zhao , Guanzhong Tian

Change detection (CD) in remote sensing aims to identify semantic differences between satellite images captured at different times. While deep learning has significantly advanced this field, existing approaches based on convolutional neural…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Durgesh Ameta , Ujjwal Mishra , Praful Hambarde , Amit Shukla

Real-time human motion reconstruction from a sparse set of (e.g. six) wearable IMUs provides a non-intrusive and economic approach to motion capture. Without the ability to acquire position information directly from IMUs, recent works took…

Computer Vision and Pattern Recognition · Computer Science 2022-12-12 Yifeng Jiang , Yuting Ye , Deepak Gopinath , Jungdam Won , Alexander W. Winkler , C. Karen Liu

Predicting the future behavior of agents is a fundamental task in autonomous vehicle domains. Accurate prediction relies on comprehending the surrounding map, which significantly regularizes agent behaviors. However, existing methods have…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Chen Feng , Hangning Zhou , Huadong Lin , Zhigang Zhang , Ziyao Xu , Chi Zhang , Boyu Zhou , Shaojie Shen

With the rapid progress of large language models (LLMs), multimodal frameworks that unify understanding and generation have become promising, yet they face increasing complexity as the number of modalities and tasks grows. We observe that…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Bingfan Zhu , Biao Jiang , Sunyi Wang , Shixiang Tang , Tao Chen , Linjie Luo , Youyi Zheng , Xin Chen

Human sensing, which employs various sensors and advanced deep learning technologies to accurately capture and interpret human body information, has significantly impacted fields like public security and robotics. However, current human…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Xinyan Chen , Jianfei Yang

Pretrained foundation models and transformer architectures have driven the success of large language models (LLMs) and other modern AI breakthroughs. However, similar advancements in health data modeling remain limited due to the need for…

Machine Learning · Computer Science 2025-07-01 Franklin Y. Ruan , Aiwei Zhang , Jenny Y. Oh , SouYoung Jin , Nicholas C. Jacobson

Ambient sensor-based human activity recognition (HAR) in smart homes remains challenging due to the need for real-time inference, spatially grounded reasoning, and context-aware temporal modeling. Existing approaches often rely on…

Machine Learning · Computer Science 2025-11-11 Zishuai Liu , Weihang You , Jin Lu , Fei Dou

Vision Transformer and its variants have demonstrated great potential in various computer vision tasks. But conventional vision transformers often focus on global dependency at a coarse level, which suffer from a learning challenge on…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Yunhao Wang , Huixin Sun , Xiaodi Wang , Bin Zhang , Chao Li , Ying Xin , Baochang Zhang , Errui Ding , Shumin Han

The growth of global consumption has motivated important applications of deep learning to smart manufacturing and machine health monitoring. In particular, analyzing vibration data offers great potential to extract meaningful insights into…

Machine Learning · Computer Science 2024-05-30 Anthony Zhou , Amir Barati Farimani

Estimating 3D human poses from monocular videos is a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Wenhao Li , Hong Liu , Hao Tang , Pichao Wang , Luc Van Gool

Current approaches to video analysis of human motion focus on raw pixels or keypoints as the basic units of reasoning. We posit that adding higher-level motion primitives, which can capture natural coarser units of motion such as backswing…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Sumith Kulal , Jiayuan Mao , Alex Aiken , Jiajun Wu

Capturing the dependencies between joints is critical in skeleton-based action recognition task. Transformer shows great potential to model the correlation of important joints. However, the existing Transformer-based methods cannot capture…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Helei Qiu , Biao Hou , Bo Ren , Xiaohua Zhang

Training state-of-the-art models for human pose estimation in videos requires datasets with annotations that are really hard and expensive to obtain. Although transformers have been recently utilized for body pose sequence modeling, related…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Fabien Baradel , Romain Brégier , Thibault Groueix , Philippe Weinzaepfel , Yannis Kalantidis , Grégory Rogez

This paper addresses the design of transmit precoder and receive combiner matrices to support $N_{\rm s}$ independent data streams over a time-division duplex (TDD) point-to-point massive multiple-input multiple-output (MIMO) channel with…

Information Theory · Computer Science 2024-06-10 Tao Jiang , Wei Yu

Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ruoyu Wang , Wenqian Wang , Jianjun Gao , Dan Lin , Kim-Hui Yap , Bingbing Li

Common fully glazed facades and transparent objects present architectural barriers and impede the mobility of people with low vision or blindness, for instance, a path detected behind a glass door is inaccessible unless it is correctly…

Computer Vision and Pattern Recognition · Computer Science 2021-08-23 Jiaming Zhang , Kailun Yang , Angela Constantinescu , Kunyu Peng , Karin Müller , Rainer Stiefelhagen