English
Related papers

Related papers: PSUMNet: Unified Modality Part Streams are All You…

200 papers

We propose a Convolutional Neural Network (CNN)-based model "RotationNet," which takes multi-view images of an object as input and jointly estimates its pose and object category. Unlike previous approaches that use known viewpoint labels…

Computer Vision and Pattern Recognition · Computer Science 2018-03-26 Asako Kanezaki , Yasuyuki Matsushita , Yoshifumi Nishida

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Michael S. Ryoo , AJ Piergiovanni , Juhana Kangaspunta , Anelia Angelova

We introduce a novel state-space model (SSM)-based framework for skeleton-based human action recognition, with an anatomically-guided architecture that improves state-of-the-art performance in both clinical diagnostics and general action…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Niki Martinel , Mariano Serrao , Christian Micheloni

Accurately reconstructing full-body poses from sparse head and hand trajectories is a foundational challenge for immersive AR/VR telepresence. Current methods often struggle with error accumulation and unnatural joint coordination,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Runzhen Liu , Chuhua Xian , Fa-Ting Hong

Most current action recognition methods heavily rely on appearance information by taking an RGB sequence of entire image regions as input. While being effective in exploiting contextual information around humans, e.g., human appearance and…

Computer Vision and Pattern Recognition · Computer Science 2021-04-16 Gyeongsik Moon , Heeseung Kwon , Kyoung Mu Lee , Minsu Cho

We study multi-dataset training (MDT) for pose estimation, where skeletal heterogeneity presents a unique challenge that existing methods have yet to address. In traditional domains, \eg regression and classification, MDT typically relies…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Uyoung Jeong , Jonathan Freer , Seungryul Baek , Hyung Jin Chang , Kwang In Kim

In smart manufacturing environments, accurate and real-time recognition of worker actions is essential for productivity, safety, and human-machine collaboration. While skeleton-based human activity recognition (HAR) offers robustness to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Vinit Hegiste , Vidit Goyal , Tatjana Legler , Martin Ruskowski

Motion retargeting is a fundamental problem in computer graphics and computer vision. Existing approaches usually have many strict requirements, such as the source-target skeletons needing to have the same number of joints or share the same…

Graphics · Computer Science 2023-06-16 Lei Hu , Zihao Zhang , Chongyang Zhong , Boyuan Jiang , Shihong Xia

In many applications, a mobile manipulator robot is required to grasp a set of objects distributed in space. This may not be feasible from a single base pose and the robot must plan the sequence of base poses for grasping all objects,…

Robotics · Computer Science 2025-02-04 Lakshadeep Naik , Sinan Kalkan , Sune L. Sørensen , Mikkel B. Kjærgaard , Norbert Krüger

With the prevalence of RGB-D cameras, multi-modal video data have become more available for human action recognition. One main challenge for this task lies in how to effectively leverage their complementary information. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2020-02-03 Sijie Song , Jiaying Liu , Yanghao Li , Zongming Guo

Deep ConvNets have been shown to be effective for the task of human pose estimation from single images. However, several challenging issues arise in the video-based case such as self-occlusion, motion blur, and uncommon poses with few or no…

Computer Vision and Pattern Recognition · Computer Science 2017-04-03 Jie Song , Limin Wang , Luc Van Gool , Otmar Hilliges

Person re-identification (reID) aims at retrieving a person from images captured by different cameras. For deep-learning-based reID methods, it has been proved that using local features together with global feature could help to give robust…

Computer Vision and Pattern Recognition · Computer Science 2022-01-11 Zhijun He , Hongbo Zhao , Wenquan Feng

Recognizing human actions from point cloud sequence has attracted tremendous attention from both academia and industry due to its wide applications. However, most previous studies on point cloud action recognition typically require complex…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Shenglin He , Xiaoyang Qu , Jiguang Wan , Guokuan Li , Changsheng Xie , Jianzong Wang

Multimodal remote sensing semantic segmentation enhances scene interpretation by exploiting complementary physical cues from heterogeneous data. Although pretrained Vision Foundation Models (VFMs) provide strong general-purpose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Haocheng Li , Juepeng Zheng , Shuangxi Miao , Ruibo Lu , Guosheng Cai , Haohuan Fu , Jianxi Huang

Action recognition has been a heated topic in computer vision for its wide application in vision systems. Previous approaches achieve improvement by fusing the modalities of the skeleton sequence and RGB video. However, such methods have a…

Computer Vision and Pattern Recognition · Computer Science 2022-02-24 Xiaoguang Zhu , Ye Zhu , Haoyu Wang , Honglin Wen , Yan Yan , Peilin Liu

The estimation of the camera poses associated with a set of images commonly relies on feature matches between the images. In contrast, we are the first to address this challenge by using objectness regions to guide the pose estimation…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Matteo Taiana , Matteo Toso , Stuart James , Alessio Del Bue

This work proposes a new end-to-end DCNN based approach for motion segmentation, especially for video sequences captured with such non-static cameras, called MOSNET. While other approaches focus on spatial or temporal context only, the…

Computer Vision and Pattern Recognition · Computer Science 2021-02-23 Markus Bosch

Human body trajectories are a salient cue to identify actions in the video. Such body trajectories are mainly conveyed by hands and face across consecutive frames in sign language. However, current methods in continuous sign language…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Lianyu Hu , Liqing Gao , Zekang Liu , Wei Feng

Automatic surgical gesture recognition is fundamental for improving intelligence in robot-assisted surgery, such as conducting complicated tasks of surgery surveillance and skill evaluation. However, current methods treat each frame…

Artificial Intelligence · Computer Science 2020-02-21 Xiaojie Gao , Yueming Jin , Qi Dou , Pheng-Ann Heng

Skeleton-based action recognition aims to recognize human actions given human joint coordinates with skeletal interconnections. By defining a graph with joints as vertices and their natural connections as edges, previous works successfully…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Yuxuan Zhou , Zhi-Qi Cheng , Chao Li , Yanwen Fang , Yifeng Geng , Xuansong Xie , Margret Keuper