中文
相关论文

相关论文: YouTube-SL-25: A Large-Scale, Open-Domain Multilin…

200 篇论文

The primary concern of this research is to take American Sign Language (ASL) data through real time camera footage and be able to convert the data and information into text. Adding to that, we are also putting focus on creating a framework…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Hasnat Jamil Bhuiyan , Mubtasim Fuad Mozumder , Md. Rabiul Islam Khan , Md. Sabbir Ahmed , Nabuat Zaman Nahim

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect…

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Thanadol Singkhornart , Olarik Surinta

Sign Language Recognition (SLR) is an essential yet challenging task since sign language is performed with the fast and complex movement of hand gestures, body posture, and even facial expressions. %Skeleton Aware Multi-modal Sign Language…

计算机视觉与模式识别 · 计算机科学 2023-02-17 Maxim Novopoltsev , Leonid Verkhovtsev , Ruslan Murtazin , Dmitriy Milevich , Iuliia Zemtsova

Sign language recognition (SLR) is a machine learning task aiming to identify signs in videos. Due to the scarcity of annotated data, unsupervised methods like contrastive learning have become promising in this field. They learn meaningful…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Ariel Basso Madjoukeng , Jérôme Fink , Pierre Poitier , Edith Belise Kenmogne , Benoit Frenay

Sign languages are essential for the Deaf and Hard-of-Hearing (DHH) community. Sign language generation systems have the potential to support communication by translating from written languages, such as English, into signed videos. However,…

Sign Language Recognition (SLR) involves the automatic identification and classification of sign gestures from images or video, converting them into text or speech to improve accessibility for the hearing-impaired community. In Bangladesh,…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Jubayer Ahmed Bhuiyan Shawon , Hasan Mahmud , Kamrul Hasan

The objective of this work is to annotate sign instances across a broad vocabulary in continuous sign language. We train a Transformer model to ingest a continuous signing stream and output a sequence of written tokens on a large-scale…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Gül Varol , Liliane Momeni , Samuel Albanie , Triantafyllos Afouras , Andrew Zisserman

In this paper, we propose SignLLM, a multilingual Sign Language Production (SLP) large language model, which includes two novel multilingual SLP modes MLSF and Prompt2LangGloss that allow sign language gestures generation from query texts…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Sen Fang , Chen Chen , Lei Wang , Ce Zheng , Chunyu Sui , Yapeng Tian

Sign Language (SL), as the mother tongue of the deaf community, is a special visual language that most hearing people cannot understand. In recent years, neural Sign Language Translation (SLT), as a possible way for bridging communication…

计算与语言 · 计算机科学 2022-11-02 Jiangbin Zheng , Siyuan Li , Cheng Tan , Chong Wu , Yidong Chen , Stan Z. Li

Self-supervised learning (SSL) has helped extend speech technologies to more languages by reducing the need for labeled data. However, models are still far from supporting the world's 7000+ languages. We propose XEUS, a Cross-lingual…

Sign language recognition (SLR) has recently achieved a breakthrough in performance thanks to deep neural networks trained on large annotated sign datasets. Of the many different sign languages, these annotated datasets are only available…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Ahmet Alp Kindiroglu , Ozgur Kara , Ogulcan Ozdemir , Lale Akarun

According to interviews with people who work with speech impaired persons, speech impaired people have difficulties in communicating with other people around them who do not know the sign language, and this situation may cause them to…

计算机视觉与模式识别 · 计算机科学 2021-02-24 Arda Mavi

Many technologies for human-computer interaction have been designed for hearing individuals and depend upon vocalized speech, precluding users of American Sign Language (ASL) in the Deaf community from benefiting from these advancements.…

Models of sign language have historically lagged behind those for spoken language (text and speech). Recent work has greatly improved their performance on tasks like sign language translation and isolated sign recognition. However, it…

计算与语言 · 计算机科学 2026-05-01 Serpil Karabüklü , Kanishka Misra , Shester Gueuwou , Diane Brentari , Greg Shakhnarovich , Karen Livescu

Automatic sign language recognition is a research area that encompasses human-computer interaction, computer vision and machine learning. Robust automatic recognition of sign language could assist in the translation process and the…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Franco Ronchetti , Facundo Manuel Quiroga , César Estrebou , Laura Lanzarini , Alejandro Rosete

Accurate sign language understanding serves as a crucial communication channel for individuals with disabilities. Current sign language translation algorithms predominantly rely on RGB frames, which may be limited by fixed frame rates,…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Xiao Wang , Yuehang Li , Fuling Wang , Bo Jiang , Yaowei Wang , Yonghong Tian , Jin Tang , Bin Luo

While Online Learning is growing and becoming widespread, the associated curricula often suffer from a lack of coverage and outdated content. In this regard, a key question is how to dynamically define the topics that must be covered to…

计算机与社会 · 计算机科学 2024-12-11 Mohammad Moein , Mohammadreza Molavi Hajiagha , Abdolali Faraji , Mohammadreza Tavakoli , Gàbor Kismihòk

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect this multi-scale…

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Zhenyu Yang , Yuhang Hu , Zemin Du , Dizhan Xue , Shengsheng Qian , Jiahong Wu , Fan Yang , Weiming Dong , Changsheng Xu