English
Related papers

Related papers: VK-G2T: Vision and Context Knowledge enhanced Glos…

200 papers

Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language translation (SLT) models might be inadequate, particularly in…

Multimedia · Computer Science 2024-10-28 Xin Shen , Lei Shen , Shaozu Yuan , Heming Du , Haiyang Sun , Xin Yu

Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps…

Computer Vision and Pattern Recognition · Computer Science 2021-10-14 Mingkang Tang , Zhanyu Wang , Zhenhua Liu , Fengyun Rao , Dian Li , Xiu Li

Simultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible. The translation thus has to start with an incomplete source text, which is…

Computation and Language · Computer Science 2020-10-14 Ozan Caglayan , Julia Ive , Veneta Haralampieva , Pranava Madhyastha , Loïc Barrault , Lucia Specia

Despite the promising progress of recent autoregressive models in text-to-image (T2I) generation, their ability to handle multi-attribute and ambiguous prompts remains limited. To address these limitations, existing works have applied…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Yaqi Li , Peng Chen , Mingyang Han , Pi Bu , Haoxiang Shi , Runzhou Zhao , Yang Yao , Xuan Zhang , Jun Song , Bo Zheng

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Haotian Zhang , Pengchuan Zhang , Xiaowei Hu , Yen-Chun Chen , Liunian Harold Li , Xiyang Dai , Lijuan Wang , Lu Yuan , Jenq-Neng Hwang , Jianfeng Gao

We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing phoneme recognition as…

Computation and Language · Computer Science 2025-09-30 Gerard I. Gállego , Oriol Pareras , Martí Cortada Garcia , Lucas Takanori , Javier Hernando

Recently, automatic image caption generation has been an important focus of the work on multimodal translation task. Existing approaches can be roughly categorized into two classes, i.e., top-down and bottom-up, the former transfers the…

Computer Vision and Pattern Recognition · Computer Science 2019-09-06 Wei Wei , Ling Cheng , Xianling Mao , Guangyou Zhou , Feida Zhu

Sign language translation systems are complex and require many components. As a result, it is very hard to compare methods across publications. We present an open-source implementation of a text-to-gloss-to-pose-to-video pipeline approach,…

Computation and Language · Computer Science 2023-05-30 Amit Moryossef , Mathias Müller , Anne Göhring , Zifan Jiang , Yoav Goldberg , Sarah Ebling

Video-grounded dialogue generation (VDG) requires the system to generate a fluent and accurate answer based on multimodal knowledge. However, the difficulty in multimodal knowledge utilization brings serious hallucinations to VDG models in…

Computation and Language · Computer Science 2024-02-20 Hongcheng Liu , Pingjie Wang , Yu Wang , Yanfeng Wang

The advances in automatic sign language translation (SLT) to spoken languages have been mostly benchmarked with datasets of limited size and restricted domains. Our work advances the state of the art by providing the first baseline results…

Computation and Language · Computer Science 2023-04-17 Laia Tarrés , Gerard I. Gállego , Amanda Duarte , Jordi Torres , Xavier Giró-i-Nieto

Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning process evolves. Existing methods that augment…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Martin Q. Ma , Yuxiao Qu , Aditya Agrawal , Willis Guo , Paul Pu Liang , Ruslan Salakhutdinov , Louis-Philippe Morency

Sign language translation (SLT) addresses the problem of translating information from a sign language in video to a spoken language in text. Existing studies, while showing progress, are often limited to narrow domains and/or few sign…

Computation and Language · Computer Science 2024-07-17 Biao Zhang , Garrett Tanzer , Orhan Firat

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

Computation and Language · Computer Science 2025-08-19 Shumin Que , Anton Ragni

A persistent challenge in sign language video processing, including the task of sign to written language translation, is how we learn representations of sign language in an effective and efficient way that preserves the important attributes…

Computation and Language · Computer Science 2025-06-04 Shester Gueuwou , Xiaodan Du , Greg Shakhnarovich , Karen Livescu

Sign language translation (SLT) is typically trained with text in a single spoken language, which limits scalability and cross-language generalization. Earlier approaches have replaced gloss supervision with text-based sentence embeddings,…

Computation and Language · Computer Science 2025-10-23 Yasser Hamidullah , Shakib Yazdani , Cennet Oguz , Josef van Genabith , Cristina España-Bonet

Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However, existing Vid-LLMs rely on uniform frame sampling to extract…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Rong Fan , Kaiyan Xiao , Minghao Zhu , Liuyi Wang , Kai Dai , Zhao Yang

Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision-language models and is particularly critical in sign language translation (SLT). In SLT, meaning depends on precise grounding in…

Sign language translation (SLT) is challenging, as it involves converting sign language videos into natural language. Previous studies have prioritized accuracy over diversity. However, diversity is crucial for handling lexical and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 JiHwan Moon , Jihoon Park , Jungeun Kim , Jongseong Bae , Hyeongwoo Jeon , Ha Young Kim

Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerable room for improving the generation performance, especially…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhefan Rao , Liya Ji , Yazhou Xing , Runtao Liu , Zhaoyang Liu , Jiaxin Xie , Ziqiao Peng , Yingqing He , Qifeng Chen

In this paper, we explore the visual representations produced from a pre-trained text-to-video (T2V) diffusion model for video understanding tasks. We hypothesize that the latent representation learned from a pretrained generative T2V model…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Zixin Zhu , Xuelu Feng , Dongdong Chen , Junsong Yuan , Chunming Qiao , Gang Hua
‹ Prev 1 4 5 6 7 8 10 Next ›