中文
相关论文

相关论文: SkelCap: Automated Generation of Descriptive Text …

200 篇论文

The choice of the representations is essential for deep gait recognition methods. The binary silhouettes and skeletal coordinates are two dominant representations in recent literature, achieving remarkable advances in many scenarios.…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Chao Fan , Jingzhe Ma , Dongyang Jin , Chuanfu Shen , Shiqi Yu

Detecting aligned 3D keypoints is essential under many scenarios such as object tracking, shape retrieval and robotics. However, it is generally hard to prepare a high-quality dataset for all types of objects due to the ambiguity of…

计算机视觉与模式识别 · 计算机科学 2021-03-22 Ruoxi Shi , Zhengrong Xue , Yang You , Cewu Lu

The lack of fluency in sign language remains a barrier to seamless communication for hearing and speech-impaired communities. In this work, we propose a low-cost, real-time ASL-to-speech translation glove and an exhaustive training dataset…

计算与语言 · 计算机科学 2024-07-22 Aditya Makkar , Divya Makkar , Aarav Patel , Liam Hebert

Sign Language Translation (SLT) is a task that has not been studied relatively much compared to the study of Sign Language Recognition (SLR). However, the SLR is a study that recognizes the unique grammar of sign language, which is…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Youngmin Kim , Minji Kwak , Dain Lee , Yeongeun Kim , Hyeongboo Baek

Recent advances in tracking sensors and pose estimation software enable smart systems to use trajectories of skeleton joint locations for supervised learning. We study the problem of accurately recognizing sign language words, which is key…

计算机视觉与模式识别 · 计算机科学 2022-02-04 Joachim Gudmundsson , Martin P. Seybold , John Pfeifer

Speech Emotion Captioning (SEC) has gradually become an active research task. The emotional content conveyed through human speech are often complex, and classifying them into fixed categories may not be enough to fully capture speech…

计算与语言 · 计算机科学 2024-10-28 Ziqi Liang , Haoxiang Shi , Hanhui Chen

Large language models have revolutionized sign language generation by automatically transforming text into high-quality sign language videos, providing accessible communication for the Deaf community. However, existing LLM-based approaches…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Yanchao Zhao , Jihao Zhu , Yu Liu , Weizhuo Chen , Yuling Yang , Kun Peng

In the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have…

计算机视觉与模式识别 · 计算机科学 2021-12-16 Jingye Chen , Haiyang Yu , Jianqi Ma , Bin Li , Xiangyang Xue

We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that rely on standard next-token prediction, ScaleMoGen frames motion generation as a…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Inwoo Hwang , Hojun Jang , Bing Zhou , Jian Wang , Young Min Kim , Chuan Guo

We present SignAvatars, the first large-scale, multi-prompt 3D sign language (SL) motion dataset designed to bridge the communication gap for Deaf and hard-of-hearing individuals. While there has been an exponentially growing number of…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Zhengdi Yu , Shaoli Huang , Yongkang Cheng , Tolga Birdal

In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Kunyu Peng , Cheng Yin , Junwei Zheng , Ruiping Liu , David Schneider , Jiaming Zhang , Kailun Yang , M. Saquib Sarfraz , Rainer Stiefelhagen , Alina Roitberg

Point cloud registration, a fundamental task in 3D computer vision, has remained largely unexplored in cross-source point clouds and unstructured scenes. The primary challenges arise from noise, outliers, and variations in scale and…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Kezheng Xiong , Maoji Zheng , Qingshan Xu , Chenglu Wen , Siqi Shen , Cheng Wang

Sign language recognition is a challenging gesture sequence recognition problem, characterized by quick and highly coarticulated motion. In this paper we focus on recognition of fingerspelling sequences in American Sign Language (ASL)…

计算机视觉与模式识别 · 计算机科学 2019-08-29 Bowen Shi , Aurora Martinez Del Rio , Jonathan Keane , Diane Brentari , Greg Shakhnarovich , Karen Livescu

We are releasing a dataset containing videos of both fluent and non-fluent signers using American Sign Language (ASL), which were collected using a Kinect v2 sensor. This dataset was collected as a part of a project to develop and evaluate…

计算与语言 · 计算机科学 2022-07-11 Saad Hassan , Matthew Seita , Larwan Berke , Yingli Tian , Elaine Gale , Sooyeon Lee , Matt Huenerfauth

In recent years, deep learning techniques have been used to develop sign language recognition systems, potentially serving as a communication tool for millions of hearing-impaired individuals worldwide. However, there are inherent…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Alvaro Leandro Cavalcante Carneiro , Denis Henrique Pinheiro Salvadeo , Lucas de Brito Silva

Generating natural and linguistically accurate sign language avatars remains a formidable challenge. Current Sign Language Production (SLP) frameworks face a stark trade-off: direct text-to-pose models suffer from regression-to-the-mean…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Jianhe Low , Alexandre Symeonidis-Herzig , Maksym Ivashechkin , Ozge Mercanoglu Sincan , Richard Bowden

Skeleton-based action recognition has gained significant attention for its ability to efficiently represent spatiotemporal information in a lightweight format. Most existing approaches use graph-based models to process skeleton sequences,…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Jushang Qiu , Lei Wang

Sign language representation learning presents unique challenges due to the complex spatio-temporal nature of signs and the scarcity of labeled datasets. Existing methods often rely either on models pre-trained on general visual tasks, that…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Ryan Wong , Necati Cihan Camgoz , Richard Bowden

Skeletonization is a popular shape analysis technique that models an object's interior as opposed to just its boundary. Fitting template-based skeletal models is a time-consuming process requiring much manual parameter tuning. Recently,…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Nicolás Gaggion , Enzo Ferrante , Beatriz Paniagua , Jared Vicory

We propose a novel system for unsupervised skeleton-based action recognition. Given inputs of body keypoints sequences obtained during various movements, our system associates the sequences with actions. Our system is based on an…

计算机视觉与模式识别 · 计算机科学 2019-12-02 Kun Su , Xiulong Liu , Eli Shlizerman