English
Related papers

Related papers: Sign-IDD: Iconicity Disentangled Diffusion for Sig…

200 papers

Sign language production (SLP) aims to translate spoken language sentences into a sequence of pose frames in a sign language, bridging the communication gap and promoting digital inclusion for deaf and hard-of-hearing communities. Existing…

Computation and Language · Computer Science 2025-09-16 Liqian Feng , Lintao Wang , Kun Hu , Dehui Kong , Zhiyong Wang

The Sign Language Production (SLP) project aims to automatically translate spoken languages into sign sequences. Our approach focuses on the transformation of sign gloss sequences into their corresponding sign pose sequences (G2P). In this…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Pan Xie , Qipeng Zhang , Taiyi Peng , Hao Tang , Yao Du , Zexian Li

The diversity of sign representation is essential for Sign Language Production (SLP) as it captures variations in appearance, facial expressions, and hand movements. However, existing SLP models are often unable to capture diversity while…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Mohamed Ilyes Lakhal , Richard Bowden

Sign Languages (SL) serve as the primary mode of communication for the Deaf and Hard of Hearing communities. Deep learning methods for SL recognition and translation have achieved promising results. However, Sign Language Production (SLP)…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Vasileios Baltatzis , Rolandos Alexandros Potamias , Evangelos Ververas , Guanxiong Sun , Jiankang Deng , Stefanos Zafeiriou

Sign languages are visual languages, with vocabularies as rich as their spoken language counterparts. However, current deep-learning based Sign Language Production (SLP) models produce under-articulated skeleton pose sequences from…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Ben Saunders , Necati Cihan Camgoz , Richard Bowden

Earlier Sign Language Production (SLP) models typically relied on autoregressive methods that generate output tokens one by one, which inherently provide temporal alignment. Although techniques like Teacher Forcing can prevent model…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Maoxiao Ye , Xinfeng Ye , Mano Manoharan

Sign languages are multi-channel visual languages, where signers use a continuous 3D space to communicate.Sign Language Production (SLP), the automatic translation from spoken to sign languages, must embody both the continuous articulation…

Computer Vision and Pattern Recognition · Computer Science 2021-03-15 Ben Saunders , Necati Cihan Camgoz , Richard Bowden

Sign language translation (SLT) is challenging, as it involves converting sign language videos into natural language. Previous studies have prioritized accuracy over diversity. However, diversity is crucial for handling lexical and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 JiHwan Moon , Jihoon Park , Jungeun Kim , Jongseong Bae , Hyeongwoo Jeon , Ha Young Kim

In this paper, we propose a dual-condition diffusion pre-training model named SignDiff that can generate human sign language speakers from a skeleton pose. SignDiff has a novel Frame Reinforcement Network called FR-Net, similar to dense…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Sen Fang , Chunyu Sui , Yanghao Zhou , Xuedong Zhang , Hongbin Zhong , Yapeng Tian , Chen Chen

Drawing on recent advancements in diffusion models for text-to-image generation, identity-preserved personalization has made significant progress in accurately capturing specific identities with just a single reference image. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Yi Wu , Ziqiang Li , Heliang Zheng , Chaoyue Wang , Bin Li

In this work, we propose DARSLP, a simple gloss-free, transformer-based sign language production (SLP) framework that directly maps spoken-language text to sign pose sequences. We first train a pose autoencoder that encodes sign poses into…

Machine Learning · Computer Science 2025-09-24 Sumeyye Meryem Tasyurek , Tugce Kiziltepe , Hacer Yalim Keles

We propose a novel framework for ID-preserving generation using a multi-modal encoding strategy rather than injecting identity features via adapters into pre-trained models. Our method treats identity and text as a unified conditioning…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Zichuan Liu , Liming Jiang , Qing Yan , Yumin Jia , Hao Kang , Xin Lu

Sign Language Production (SLP) is the process of converting the complex input text into a real video. Most previous works focused on the Text2Gloss, Gloss2Pose, Pose2Vid stages, and some concentrated on Prompt2Gloss and Text2Avatar stages.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sen Fang , Yalin Feng , Hongbin Zhong , Yanxin Zhang , Dimitris N. Metaxas

Text-conditioned image generation models have recently achieved astonishing results in image quality and text alignment and are consequently employed in a fast-growing number of applications. Since they are highly data-driven, relying on…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Patrick Schramowski , Manuel Brack , Björn Deiseroth , Kristian Kersting

Contrastive Language-Image Pre-training (CLIP) starts to emerge in many computer vision tasks and has achieved promising performance. However, it remains underexplored whether CLIP can be generalized to 3D hand pose estimation, as bridging…

Multimedia · Computer Science 2023-09-29 Shaoxiang Guo , Qing Cai , Lin Qi , Junyu Dong

Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation methods may not generalize well to synthesized images or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yuanhao Zhai , Kevin Lin , Linjie Li , Chung-Ching Lin , Jianfeng Wang , Zhengyuan Yang , David Doermann , Junsong Yuan , Zicheng Liu , Lijuan Wang

We propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured in-the-wild image of a subject. The foundation of our approach is anchored in…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Francesca Babiloni , Alexandros Lattas , Jiankang Deng , Stefanos Zafeiriou

To be truly understandable and accepted by Deaf communities, an automatic Sign Language Production (SLP) system must generate a photo-realistic signer. Prior approaches based on graphical avatars have proven unpopular, whereas recent neural…

Computer Vision and Pattern Recognition · Computer Science 2020-11-30 Ben Saunders , Necati Cihan Camgoz , Richard Bowden

Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative model of 3D body motion and explore the role of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Rui Hong , Jana Kosecka

There has been significant progress in personalized image synthesis with methods such as Textual Inversion, DreamBooth, and LoRA. Yet, their real-world applicability is hindered by high storage demands, lengthy fine-tuning processes, and…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Qixun Wang , Xu Bai , Haofan Wang , Zekui Qin , Anthony Chen , Huaxia Li , Xu Tang , Yao Hu
‹ Prev 1 2 3 10 Next ›