English
Related papers

Related papers: Adversarial Training for Multi-Channel Sign Langua…

200 papers

In recent years, large-scale pre-trained speech language models (SLMs) have demonstrated remarkable advancements in various generative speech modeling applications, such as text-to-speech synthesis, voice conversion, and speech enhancement.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-19 Yinghao Aaron Li , Cong Han , Nima Mesgarani

Modern sensing systems generate large volumes of unlabeled multivariate time-series data. This abundance of unlabeled data makes self-supervised learning (SSL) a natural approach for learning transferable representations. However, most…

Artificial Intelligence · Computer Science 2026-03-13 Yuliang Chen , Arvind Pillai , Yu Yvonne Wu , Tess Z. Griffin , Lisa Marsch , Michael V. Heinz , Nicholas C. Jacobson , Andrew Campbell

Sign language, which contains hand movements, facial expressions and bodily gestures, is a significant medium for communicating with hard-of-hearing people. A well-trained sign language community communicates easily, but those who don't…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Ajeet Kumar Yadav , Nishant Kumar , Rathna G N

Modern text-to-speech synthesis pipelines typically involve multiple processing stages, each of which is designed or learnt independently from the rest. In this work, we take on the challenging task of learning to synthesise speech from…

Sound · Computer Science 2021-03-18 Jeff Donahue , Sander Dieleman , Mikołaj Bińkowski , Erich Elsen , Karen Simonyan

Sign language recognition and translation technologies have the potential to increase access and inclusion of deaf signing communities, but research progress is bottlenecked by a lack of representative data. We introduce a new resource for…

Computation and Language · Computer Science 2023-10-03 Lee Kezar , Elana Pontecorvo , Adele Daniels , Connor Baer , Ruth Ferster , Lauren Berger , Jesse Thomason , Zed Sevcikova Sehyr , Naomi Caselli

Sign language is the primary communication mode for 72 million hearing-impaired individuals worldwide, necessitating effective bidirectional Sign Language Production and Sign Language Translation systems. However, functional bidirectional…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Yulong Li , Yuxuan Zhang , Feilong Tang , Ming Hu , Zhixiang Lu , Haochen Xue , Jianghao Wu , Mian Zhou , Kang Dang , Chong Li , Yifang Wang , Imran Razzak , Jionglong Su

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction through spoken or written language. This spoken-language-centered interface can limit access…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Youngmin Kim , Kyobin Choo , Jiwoo Park , Minseo Kim , Chanyoung Kim , Junhyeok Kim , Seong Jae Hwang

Semantic image synthesis aims at generating photorealistic images from semantic layouts. Previous approaches with conditional generative adversarial networks (GAN) show state-of-the-art performance on this task, which either feed the…

Computer Vision and Pattern Recognition · Computer Science 2020-01-13 Xihui Liu , Guojun Yin , Jing Shao , Xiaogang Wang , Hongsheng Li

We present SignAvatars, the first large-scale, multi-prompt 3D sign language (SL) motion dataset designed to bridge the communication gap for Deaf and hard-of-hearing individuals. While there has been an exponentially growing number of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Zhengdi Yu , Shaoli Huang , Yongkang Cheng , Tolga Birdal

Sign language is commonly used by deaf or mute people to communicate but requires extensive effort to master. It is usually performed with the fast yet delicate movement of hand gestures, body posture, and even facial expressions. Current…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Songyao Jiang , Bin Sun , Lichen Wang , Yue Bai , Kunpeng Li , Yun Fu

We present a novel approach to generating photo-realistic images of a face with accurate lip sync, given an audio input. By using a recurrent neural network, we achieved mouth landmarks based on audio features. We exploited the power of…

Computer Vision and Pattern Recognition · Computer Science 2018-03-21 Seyed Ali Jalalifar , Hosein Hasani , Hamid Aghajan

Sign Language Translation (SLT) aims to convert sign language videos into spoken or written text. While early systems relied on gloss annotations as an intermediate supervision, such annotations are costly to obtain and often fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Ozge Mercanoglu Sincan , Richard Bowden

The absence of effective communication the deaf population represents the main social gap in this community. Furthermore, the sign language, main deaf communication tool, is unlettered, i.e., there is no formal written representation. In…

Computation and Language · Computer Science 2025-03-26 Fredy Alejandro Mendoza López , Jefferson Rodriguez , Fabio Martínez

Transformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding. While existing research has examined how well SLMs encode shallow acoustic and phonetic features, the extent to which…

Computation and Language · Computer Science 2025-09-22 Linyang He , Qiaolin Wang , Xilin Jiang , Nima Mesgarani

Transcribed datasets typically contain speaker identity for each instance in the data. We investigate two ways to incorporate this information during training: Multi-Task Learning and Adversarial Learning. In multi-task learning, the goal…

Machine Learning · Computer Science 2019-02-15 Yossi Adi , Neil Zeghidour , Ronan Collobert , Nicolas Usunier , Vitaliy Liptchinsky , Gabriel Synnaeve

This paper studies learning fair encoders in a self-supervised learning (SSL) setting, in which all data are unlabeled and only a small portion of them are annotated with sensitive attribute. Adversarial fair representation learning is well…

Machine Learning · Computer Science 2024-06-11 Qi Qi , Quanqi Hu , Qihang Lin , Tianbao Yang

State-of-the-art sign language generation frameworks lack expressivity and naturalness which is the result of only focusing manual signs, neglecting the affective, grammatical and semantic functions of facial expressions. The purpose of…

Computation and Language · Computer Science 2024-10-21 Carla Viegas , Mert İnan , Lorna Quandt , Malihe Alikhani

Besides the well-known classification task, these days neural networks are frequently being applied to generate or transform data, such as images and audio signals. In such tasks, the conventional loss functions like the mean squared error…

Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based…

Computation and Language · Computer Science 2025-08-08 Wenqian Cui , Dianzhi Yu , Xiaoqi Jiao , Ziqiao Meng , Guangyan Zhang , Qichao Wang , Yiwen Guo , Irwin King