English
Related papers

Related papers: Vision-Language Meets the Skeleton: Progressively …

200 papers

Contrastive learning has been successfully leveraged to learn action representations for addressing the problem of semi-supervised skeleton-based action recognition. However, most contrastive learning-based methods only contrast global…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Binqian Xu , Xiangbo Shu

Human action recognition (HAR) with multi-modal inputs (RGB-D, skeleton, point cloud) can achieve high accuracy but typically relies on large labeled datasets and degrades sharply when sensors fail or are noisy. We present Robust…

Signal Processing · Electrical Eng. & Systems 2025-11-18 Hasan Akgul , Mari Eplik , Javier Rojas , Akira Yamamoto , Rajesh Kumar , Maya Singh

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Jinyu Yang , Jiali Duan , Son Tran , Yi Xu , Sampath Chanda , Liqun Chen , Belinda Zeng , Trishul Chilimbi , Junzhou Huang

Combining gradient-based trajectory optimization with differentiable physics simulation is an efficient technique for solving soft-body manipulation problems. Using a well-crafted optimization objective, the solver can quickly converge onto…

Machine Learning · Computer Science 2023-12-12 Zhiao Huang , Feng Chen , Yewen Pu , Chunru Lin , Hao Su , Chuang Gan

Representation learning for sketch-based image retrieval has mostly been tackled by learning embeddings that discard modality-specific information. As instances from different modalities can often provide complementary information…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Abhra Chaudhuri , Massimiliano Mancini , Yanbei Chen , Zeynep Akata , Anjan Dutta

Human skeleton information is important in skeleton-based action recognition, which provides a simple and efficient way to describe human pose. However, existing skeleton-based methods focus more on the skeleton, ignoring the objects…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Hao Wen , Ziqian Lu , Fengli Shen , Zhe-Ming Lu , Jialin Cui

In this work, we address the problem how a network for action recognition that has been trained on a modality like RGB videos can be adapted to recognize actions for another modality like sequences of 3D human poses. To this end, we extract…

Computer Vision and Pattern Recognition · Computer Science 2019-10-11 Fida Mohammad Thoker , Juergen Gall

The introduction of low-cost RGB-D sensors has promoted the research in skeleton-based human action recognition. Devising a representation suitable for characterising actions on the basis of noisy skeleton sequences remains a challenge,…

Computer Vision and Pattern Recognition · Computer Science 2015-04-21 Ruizhi Qiao , Lingqiao Liu , Chunhua Shen , Anton von den Hengel

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Yang Xing , Jiong Wu , Savas Ozdemir , Ying Zhang , Yang Yang , Wei Shao , Kuang Gong

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Qingqing Zhao , Yao Lu , Moo Jin Kim , Zipeng Fu , Zhuoyang Zhang , Yecheng Wu , Zhaoshuo Li , Qianli Ma , Song Han , Chelsea Finn , Ankur Handa , Ming-Yu Liu , Donglai Xiang , Gordon Wetzstein , Tsung-Yi Lin

Vision-and-Language Navigation (VLN) has gained significant research interest in recent years due to its potential applications in real-world scenarios. However, existing VLN methods struggle with the issue of spurious associations,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Liuyi Wang , Zongtao He , Ronghao Dang , Huiyi Chen , Chengju Liu , Qijun Chen

The alignment of vision-language representations endows current Vision-Language Models (VLMs) with strong multi-modal reasoning capabilities. However, the interpretability of the alignment component remains uninvestigated due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Shufan Shen , Junshu Sun , Qingming Huang , Shuhui Wang

The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across different modalities. Recent research tends to use modality-shared language prompts…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Xiaomei Yang , Xizhan Gao , Antai Liu , Kang Wei , Fa Zhu , Guang Feng , Xiaofeng Qu , Sijie Niu

Action prediction is to recognize the class label of an ongoing activity when only a part of it is observed. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is…

Computer Vision and Pattern Recognition · Computer Science 2019-04-04 Jun Liu , Amir Shahroudy , Gang Wang , Ling-Yu Duan , Alex C. Kot

Sign language is commonly used by deaf or mute people to communicate but requires extensive effort to master. It is usually performed with the fast yet delicate movement of hand gestures, body posture, and even facial expressions. Current…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Songyao Jiang , Bin Sun , Lichen Wang , Yue Bai , Kunpeng Li , Yun Fu

Most existing one-shot skeleton-based action recognition focuses on raw low-level information (e.g., joint location), and may suffer from local information loss and low generalization ability. To alleviate these, we propose to leverage text…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Tingbing Yan , Wenzheng Zeng , Yang Xiao , Xingyu Tong , Bo Tan , Zhiwen Fang , Zhiguo Cao , Joey Tianyi Zhou

Point cloud registration, a fundamental task in 3D computer vision, has remained largely unexplored in cross-source point clouds and unstructured scenes. The primary challenges arise from noise, outliers, and variations in scale and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Kezheng Xiong , Maoji Zheng , Qingshan Xu , Chenglu Wen , Siqi Shen , Cheng Wang

Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Haoxiang Gao , Li Zhang , Yu Zhao , Zhou Yang , Jinghan Cao

A novel skill learning approach is proposed that allows a robot to acquire human-like visuospatial skills for object manipulation tasks. Visuospatial skills are attained by observing spatial relationships among objects through…

Robotics · Computer Science 2017-06-06 S. Reza Ahmadzadeh , Fulvio Mastrogiovanni , Petar Kormushev

Pre-training vision-language representations on human action videos has emerged as a promising approach to reduce reliance on large-scale expert demonstrations for training embodied agents. However, prior methods often employ time…

Robotics · Computer Science 2025-12-19 Zhizhen Zhang , Lei Zhu , Zhen Fang , Zi Huang , Yadan Luo