中文

Gloss2Text:基于 LLM 和语义感知标签平滑的手语gloss翻译

计算机视觉与模式识别 2024-07-15 v2 计算与语言 机器学习

摘要

从视频到 spoken text 的手语翻译 presents unique challenges due to the distinct grammar、expression nuances and high variation of visual appearance across different speakers and contexts。 intermediate gloss 注释旨在指导翻译过程。在我们的 work 中,我们 focus on Gloss2Text 翻译 stage,并通过利用 pre-trained large language models(LLMs)、data augmentation and novel label-smoothing loss function exploiting gloss translation ambiguities 来提出 several advances,显著提高 state-of-the-art approaches 的性能。通过在 PHOENIX Weather 2014T 数据集上进行 extensive experiments and ablation studies,我们的方法在 Gloss2Text 翻译方面超越了 state-of-the-art performance,表明其在 addressing sign language translation and suggesting promising avenues for future research and development 中具有效用。

关键词

引用

@article{arxiv.2407.01394,
  title  = {Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing},
  author = {Pooya Fayyazsanavi and Antonios Anastasopoulos and Jana Košecká},
  journal= {arXiv preprint arXiv:2407.01394},
  year   = {2024}
}