中文
相关论文

相关论文: Neural Sign Actors: A diffusion model for 3D sign …

200 篇论文

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain limits their ability…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Alexandre Symeonidis-Herzig , Özge Mercanoğlu Sincan , Richard Bowden

Latent diffusion models such as Stable Diffusion achieve state-of-the-art results on text-to-image generation tasks. However, the extent to which these models have a semantic understanding of the images they generate is not well understood.…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Cameron Braunstein , Mariya Toneva , Eddy Ilg

This research explores using lightweight deep neural network architectures to enable the humanoid robot Pepper to understand American Sign Language (ASL) and facilitate non-verbal human-robot interaction. First, we introduce a lightweight…

机器人学 · 计算机科学 2023-10-02 JongYoon Lim , Inkyu Sa , Bruce MacDonald , Ho Seok Ahn

Existing Sign Language Learning applications focus on the demonstration of the sign in the hope that the student will copy a sign correctly. In these cases, only a teacher can confirm that the sign was completed correctly, by reviewing a…

机器学习 · 计算机科学 2024-12-25 Nikita Louison , Wayne Goodridge , Koffka Khan

Sign language recognition is a challenging and often underestimated problem comprising multi-modal articulators (handshape, orientation, movement, upper body and face) that integrate asynchronously on multiple streams. Learning powerful…

计算机视觉与模式识别 · 计算机科学 2019-11-22 Hamid Reza Vaezi Joze , Oscar Koller

Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text. Token-level diffusion doesn't model word-order dependencies explicitly and operates on short, fixed…

计算与语言 · 计算机科学 2025-05-27 Xiaochen Zhu , Georgi Karadzhov , Chenxi Whitehouse , Andreas Vlachos

Sign Language (SL) enables two-way communication for the deaf and hard-of-hearing community, yet many sign languages remain under-resourced in the AI space. Sign Language Instruction Generation (SLIG) produces step-by-step textual…

人机交互 · 计算机科学 2025-08-27 Md Tariquzzaman , Md Farhan Ishmam , Saiyma Sittul Muna , Md Kamrul Hasan , Hasan Mahmud

Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent…

Sign language recognition is important for natural and convenient communication between deaf community and hearing majority. We take the highly efficient initial step of automatic fingerspelling recognition system using convolutional neural…

计算机视觉与模式识别 · 计算机科学 2015-10-15 Byeongkeun Kang , Subarna Tripathi , Truong Q. Nguyen

It has always been a rather tough task to communicate with someone possessing a hearing impairment. One of the most tested ways to establish such a communication is through the use of sign based languages. However, not many people are aware…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Sharanya Mukherjee , Md Hishaam Akhtar , Kannadasan R

Gloss-free Sign Language Production (SLP) offers a direct translation of spoken language sentences into sign language, bypassing the need for gloss intermediaries. This paper presents the Sign language Vector Quantization Network, a novel…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Eui Jun Hwang , Huije Lee , Jong C. Park

The count of people suffering from various levels of hearing loss reached 1.57 billion in 2019. This huge number tends to suffer on many personal and professional levels and strictly needs to be included with the rest of society healthily.…

信号处理 · 电气工程与系统科学 2023-12-20 Basma Kalandar , Ziemowit Dworakowski

Recently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models. However, existing methods…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Xiao Han , Yukang Cao , Kai Han , Xiatian Zhu , Jiankang Deng , Yi-Zhe Song , Tao Xiang , Kwan-Yee K. Wong

In this paper, we propose a 3D Convolutional Neural Network (3DCNN) based multi-stream framework to recognize American Sign Language (ASL) manual signs (consisting of movements of the hands, as well as non-manual face movements in some…

计算机视觉与模式识别 · 计算机科学 2019-06-10 Longlong Jing , Elahe Vahdani , Matt Huenerfauth , Yingli Tian

Existing end-to-end sign-language animation systems suffer from low naturalness, limited facial/body expressivity, and no user control. We propose a human-centered, real-time speech-to-sign animation framework that integrates (1) a…

人机交互 · 计算机科学 2025-06-25 Yingchao Li

The absence of effective communication the deaf population represents the main social gap in this community. Furthermore, the sign language, main deaf communication tool, is unlettered, i.e., there is no formal written representation. In…

计算与语言 · 计算机科学 2025-03-26 Fredy Alejandro Mendoza López , Jefferson Rodriguez , Fabio Martínez

Diffusion-based models have gained wide adoption in the virtual human generation due to their outstanding expressiveness. However, their substantial computational requirements have constrained their deployment in real-time interactive…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Haojie Yu , Zhaonian Wang , Yihan Pan , Meng Cheng , Hao Yang , Chao Wang , Tao Xie , Xiaoming Xu , Xiaoming Wei , Xunliang Cai

Recent work have addressed the generation of human poses represented by 2D/3D coordinates of human joints for sign language. We use the state of the art in Deep Learning for motion transfer and evaluate them on How2Sign, an American Sign…

计算机视觉与模式识别 · 计算机科学 2021-01-05 Lucas Ventura , Amanda Duarte , Xavier Giro-i-Nieto

Sign language recognition (SLR) facilitates communication between deaf and hearing individuals. Deep learning is widely used to develop SLR-based systems; however, it is computationally intensive and requires substantial computational…

机器人学 · 计算机科学 2025-12-24 Nitin Kumar Singh , Arie Rachmad Syulistyo , Yuichiro Tanaka , Hakaru Tamukoh

In this paper, we introduce a neural rendering pipeline for transferring the facial expressions, head pose, and body movements of one person in a source video to another in a target video. We apply our method to the challenging case of Sign…

计算机视觉与模式识别 · 计算机科学 2023-05-31 Christina O. Tze , Panagiotis P. Filntisis , Athanasia-Lida Dimou , Anastasios Roussos , Petros Maragos