English
Related papers

Related papers: Fine-Tuning Video Transformers for Word-Level Bang…

200 papers

This study presents TSLFormer, a light and robust word-level Turkish Sign Language (TSL) recognition model that treats sign gestures as ordered, string-like language. Instead of using raw RGB or depth videos, our method only works with 3D…

Computation and Language · Computer Science 2025-06-19 Kutay Ertürk , Furkan Altınışık , İrem Sarıaltın , Ömer Nezih Gerek

Sign Language Translation (SLT) aims to convert sign language (SL) videos into spoken language text, thereby bridging the communication gap between the sign and the spoken community. While most existing works focus on translating a single…

Computation and Language · Computer Science 2025-06-02 Sihan Tan , Taro Miyazaki , Kazuhiro Nakadai

The lack of fluency in sign language remains a barrier to seamless communication for hearing and speech-impaired communities. In this work, we propose a low-cost, real-time ASL-to-speech translation glove and an exhaustive training dataset…

Computation and Language · Computer Science 2024-07-22 Aditya Makkar , Divya Makkar , Aarav Patel , Liam Hebert

Sign Language Production (SLP) is the tough task of turning sign language into sign videos. The main goal of SLP is to create these videos using a sign gloss. In this research, we've developed a new method to make high-quality sign videos…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Pan Xie , Taiyi Peng , Yao Du , Qipeng Zhang

Continuous Sign Language Recognition (CSLR) focuses on the interpretation of a sequence of sign language gestures performed continually without pauses. In this study, we conduct an empirical evaluation of recent deep learning CSLR…

Computation and Language · Computer Science 2024-10-23 Sarah Alyami , Hamzah Luqman

Handwriting recognition remains challenging for some of the most spoken languages, like Bangla, due to the complexity of line and word segmentation brought by the curvilinear nature of writing and lack of quality datasets. This paper solves…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Sheikh Mohammad Jubaer , Nazifa Tabassum , Md. Ataur Rahman , Mohammad Khairul Islam

Self-supervised automatic speech recognition (SSL-ASR) is an ASR approach that uses speech encoders pretrained on large amounts of unlabeled audio (e.g., wav2vec2.0 or HuBERT) and then fine-tunes them with limited labeled data to perform…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-07 Eyal Cohen , Bhiksha Raj , Joseph Keshet

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic)…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Umberto Cappellazzo , Minsu Kim , Honglie Chen , Pingchuan Ma , Stavros Petridis , Daniele Falavigna , Alessio Brutti , Maja Pantic

Fine-tuning multilingual ASR models like Whisper for low-resource languages often improves read speech but degrades spontaneous audio performance, a phenomenon we term studio-bias. To diagnose this mismatch, we introduce Vividh-ASR, a…

Computation and Language · Computer Science 2026-05-14 Kush Juvekar , Kavya Manohar , Aditya Srinivas Menon , Arghya Bhattacharya , Kumarmanas Nethil

The recent advent of Large Language Models (LLMs) has ushered sophisticated reasoning capabilities into the realm of video through Video Large Language Models (VideoLLMs). However, VideoLLMs currently rely on a single vision encoder for all…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Jihoon Chung , Tyler Zhu , Max Gonzalez Saez-Diez , Juan Carlos Niebles , Honglu Zhou , Olga Russakovsky

Sign language recognition (SLR) is a challenging problem, involving complex manual features, i.e., hand gestures, and fine-grained non-manual features (NMFs), i.e., facial expression, mouth shapes, etc. Although manual features are…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Hezhen Hu , Wengang Zhou , Junfu Pu , Houqiang Li

Sign language recognition and translation technologies have the potential to increase access and inclusion of deaf signing communities, but research progress is bottlenecked by a lack of representative data. We introduce a new resource for…

Computation and Language · Computer Science 2023-10-03 Lee Kezar , Elana Pontecorvo , Adele Daniels , Connor Baer , Ruth Ferster , Lauren Berger , Jesse Thomason , Zed Sevcikova Sehyr , Naomi Caselli

This paper presents an Arabic Alphabet Sign Language recognition approach, using deep learning methods in conjunction with transfer learning and transformer-based models. We study the performance of the different variants on two publicly…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Mazen Balat , Rewaa Awaad , Hend Adel , Ahmed B. Zaky , Salah A. Aly

Surgical scene understanding is critical for surgical training and robotic decision-making in robot-assisted surgery. Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated great potential for advancing scene…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Guankun Wang , Junyi Wang , Wenjin Mo , Long Bai , Kun Yuan , Ming Hu , Jinlin Wu , Junjun He , Yiming Huang , Nicolas Padoy , Zhen Lei , Hongbin Liu , Nassir Navab , Hongliang Ren

This study introduces an integrated approach to recognizing Arabic Sign Language (ArSL) using state-of-the-art deep learning models such as MobileNetV3, ResNet50, and EfficientNet-B2. These models are further enhanced by explainable AI…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Mazen Balat , Rewaa Awaad , Ahmed B. Zaky , Salah A. Aly

Sign Language Translation (SLT) attempts to convert sign language videos into spoken sentences. However, many existing methods struggle with the disparity between visual and textual representations during end-to-end learning. Gloss-based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Sobhan Asasi , Mohamed Ilyes Lakhal , Richard Bowden

In recent years, deep learning techniques have been used to develop sign language recognition systems, potentially serving as a communication tool for millions of hearing-impaired individuals worldwide. However, there are inherent…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Alvaro Leandro Cavalcante Carneiro , Denis Henrique Pinheiro Salvadeo , Lucas de Brito Silva

Human action understanding is crucial for the advancement of multimodal systems. While recent developments, driven by powerful large language models (LLMs), aim to be general enough to cover a wide range of categories, they often overlook…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Yongle Huang , Haodong Chen , Zhenbang Xu , Zihan Jia , Haozhou Sun , Dian Shao

Financial fraud detection has emerged as a critical research challenge amid the rapid expansion of digital financial platforms. Although machine learning approaches have demonstrated strong performance in identifying fraudulent activities,…

Machine Learning · Computer Science 2026-03-13 Mohammad Shihab Uddin , Md Hasibul Amin , Nusrat Jahan Ema , Bushra Uddin , Tanvir Ahmed , Arif Hassan Zidan

Sign language translation (SLT) is challenging, as it involves converting sign language videos into natural language. Previous studies have prioritized accuracy over diversity. However, diversity is crucial for handling lexical and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 JiHwan Moon , Jihoon Park , Jungeun Kim , Jongseong Bae , Hyeongwoo Jeon , Ha Young Kim