中文
相关论文

相关论文: SkelCap: Automated Generation of Descriptive Text …

200 篇论文

Skeleton generation is essential for animating 3D assets, but current deep learning methods remain limited: they cannot handle the growing structural complexity of modern models and offer minimal controllability, creating a major bottleneck…

Our aim is to develop a unified model for sign language understanding, that performs sign language translation (SLT) and sign-subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signing videos into spoken…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Youngjoon Jang , Liliane Momeni , Zifan Jiang , Joon Son Chung , Gül Varol , Andrew Zisserman

In skeleton-based action recognition, graph convolutional networks (GCNs), which model human body skeletons using graphical components such as nodes and connections, have achieved remarkable performance recently. However, current…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Jongmin Yu , Yongsang Yoon , Moongu Jeon

In this paper, we propose a sequential neural encoder with latent structured description (SNELSD) for modeling sentences. This model introduces latent chunk-level representations into conventional sequential neural encoders, i.e., recurrent…

计算与语言 · 计算机科学 2017-11-16 Yu-Ping Ruan , Qian Chen , Zhen-Hua Ling

Accurate recognition and interpretation of sign language are crucial for enhancing communication accessibility for deaf and hard of hearing individuals. However, current approaches of Isolated Sign Language Recognition (ISLR) often face…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Karina Kvanchiani , Roman Kraynov , Elizaveta Petrova , Petr Surovcev , Aleksandr Nagaev , Alexander Kapitanov

Automated assessment of human motion plays a vital role in rehabilitation, enabling objective evaluation of patient performance and progress. Unlike general human activity recognition, rehabilitation motion assessment focuses on analyzing…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Ali Ismail-Fawaz , Maxime Devanne , Stefano Berretti , Jonathan Weber , Germain Forestier

Sign language is the primary approach of communication for the Deaf and Hard-of-Hearing (DHH) community. While there are numerous benchmarks for high-resource sign languages, low-resource languages like Arabic remain underrepresented.…

计算与语言 · 计算机科学 2026-04-01 Mohammad Amer Khalil , Raghad Nahas , Ahmad Nassar , Khloud Al Jallad

Zero-shot video captioning aims to generate sentences for describing videos without training the model on video-text pairs, which remains underexplored. Existing zero-shot image captioning methods typically adopt a text-only training…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Zeyu Pan , Ping Li , Wenxiao Wang

Animating a newly designed character using motion capture (mocap) data is a long standing problem in computer animation. A key consideration is the skeletal structure that should correspond to the available mocap data, and the shape…

图形学 · 计算机科学 2021-05-07 Peizhuo Li , Kfir Aberman , Rana Hanocka , Libin Liu , Olga Sorkine-Hornung , Baoquan Chen

Generating realistic and structurally plausible human images into existing scenes remains a significant challenge for current generative models, which often produce artifacts like distorted limbs and unnatural poses. We attribute this…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Chuqiao Wu , Jin Song , Yiyun Fei

Detecting anomalous human behaviour is an important visual task in safety-critical applications such as healthcare monitoring, workplace safety, or public surveillance. In these contexts, abnormalities are often reflected with unusual human…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Anja Delić , Matej Grcić , Siniša Šegvić

While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when…

声音 · 计算机科学 2025-01-07 Xiquan Li , Wenxi Chen , Ziyang Ma , Xuenan Xu , Yuzhe Liang , Zhisheng Zheng , Qiuqiang Kong , Xie Chen

Sign languages, used by around 70 million Deaf individuals globally, are visual languages that convey visual and contextual information. Current methods in vision-based sign language recognition (SLR) and translation (SLT) struggle with…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Yuqi Liu , Wenqian Zhang , Sihan Ren , Chengyu Huang , Jingyi Yu , Lan Xu

Skeleton-based action recognition attracts practitioners and researchers due to the lightweight, compact nature of datasets. Compared with RGB-video-based action recognition, skeleton-based action recognition is a safer way to protect the…

计算机视觉与模式识别 · 计算机科学 2023-02-15 Saemi Moon , Myeonghyeon Kim , Zhenyue Qin , Yang Liu , Dongwoo Kim

Accurate sign language understanding serves as a crucial communication channel for individuals with disabilities. Current sign language translation algorithms predominantly rely on RGB frames, which may be limited by fixed frame rates,…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Xiao Wang , Yuehang Li , Fuling Wang , Bo Jiang , Yaowei Wang , Yonghong Tian , Jin Tang , Bin Luo

This work make the first attempt to generate articulated human motion sequence from a single image. On the one hand, we utilize paired inputs including human skeleton information as motion embedding and a single human image as appearance…

计算机视觉与模式识别 · 计算机科学 2017-09-15 Yichao Yan , Jingwei Xu , Bingbing Ni , Xiaokang Yang

Round-the-clock monitoring of human behavior and emotions is required in many healthcare applications which is very expensive but can be automated using machine learning (ML) and sensor technologies. Unfortunately, the lack of…

信号处理 · 电气工程与系统科学 2021-02-17 Bonny Banerjee , Masoumeh Heidari Kapourchali , Murchana Baruah , Mousumi Deb , Kenneth Sakauye , Mette Olufsen

This work dedicates to continuous sign language recognition (CSLR), which is a weakly supervised task dealing with the recognition of continuous signs from videos, without any prior knowledge about the temporal boundaries between…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Fangyun Wei , Yutong Chen

Skeleton-based human action recognition aims to classify human skeletal sequences, which are spatiotemporal representations of actions, into predefined categories. To reduce the reliance on costly annotations of skeletal sequences while…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Zhigang Tu , Zhengbo Zhang , Jia Gong , Junsong Yuan , Bo Du

The task of translating natural language questions into query languages has long been a central focus in semantic parsing. Recent advancements in Large Language Models (LLMs) have significantly accelerated progress in this field. However,…

计算与语言 · 计算机科学 2025-11-25 Yuchen Ji , Bo Xu , Jie Shi , Jiaqing Liang , Deqing Yang , Yu Mao , Hai Chen , Yanghua Xiao