English
Related papers

Related papers: Prompting with Sign Parameters for Low-resource Si…

200 papers

Silent Speech Interfaces (SSIs) have gained attention for their ability to generate intelligible speech from non-acoustic signals. While significant progress has been made in advancing speech generation pipelines, limited work has addressed…

Computation and Language · Computer Science 2025-09-08 Nithyashree Sivasubramaniam

Speaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions. These models primarily extract embeddings for speaker identification but…

Computation and Language · Computer Science 2025-08-26 Massa Baali , Shuo Han , Syed Abdul Hannan , Purusottam Samal , Karanveer Singh , Soham Deshmukh , Rita Singh , Bhiksha Raj

Sign Language Translation (SLT) is a challenging task that aims to generate spoken language sentences from sign language videos, both of which have different grammar and word/gloss order. From a Neural Machine Translation (NMT) perspective,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Ozge Mercanoglu Sincan , Necati Cihan Camgoz , Richard Bowden

Open-vocabulary semantic segmentation aims to segment images into distinct semantic regions for both seen and unseen categories at the pixel level. Current methods utilize text embeddings from pre-trained vision-language models like CLIP…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Ziyu Zhao , Xiaoguang Li , Linjia Shi , Nasrin Imanpour , Song Wang

Benefiting from effective speech modeling, current Speech Large Language Models (SLLMs) have demonstrated exceptional capabilities in in-context speech generation and efficient generalization to unseen speakers. However, the prevailing…

Computation and Language · Computer Science 2024-01-26 Dong Zhang , Xin Zhang , Jun Zhan , Shimin Li , Yaqian Zhou , Xipeng Qiu

Systems that can efficiently search collections of sign language videos have been highlighted as a useful application of sign language technology. However, the problem of searching videos beyond individual keywords has received limited…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Amanda Duarte , Samuel Albanie , Xavier Giró-i-Nieto , Gül Varol

Sign language to spoken language audio translation is important to connect the hearing- and speech-challenged humans with others. We consider sign language videos with isolated sign sequences rather than continuous grammatical signing. Such…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Harsh Kavediya , Vighnesh Nayak , Bheeshm Sharma , Balamurugan Palaniappan

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Zirui Wang , Jiahui Yu , Adams Wei Yu , Zihang Dai , Yulia Tsvetkov , Yuan Cao

Audio-Language Models (ALMs) have recently achieved remarkable success in zero-shot audio recognition tasks, which match features of audio waveforms with class-specific text prompt features, inspired by advancements in Vision-Language…

Sound · Computer Science 2024-10-01 Asif Hanif , Maha Tufail Agro , Mohammad Areeb Qazi , Hanan Aldarmaki

The source-free cross-domain few-shot learning (CD-FSL) task aims to transfer pretrained models to target domains utilizing minimal samples, eliminating the need for source domain data. Addressing this issue requires models to have robust…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Linhai Zhuo , Zheng Wang , Yuqian Fu , Tianwen Qian

We present, to our knowledge, the first sign language-driven Vision-Language-Action (VLA) framework for intuitive and inclusive human-robot interaction. Unlike conventional approaches that rely on gloss annotations as intermediate…

Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos into hidden…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Yutong Chen , Ronglai Zuo , Fangyun Wei , Yu Wu , Shujie Liu , Brian Mak

A machine can understand human activities, and the meaning of signs can help overcome the communication barriers between the inaudible and ordinary people. Sign Language Recognition (SLR) is a fascinating research area and a crucial task…

Computer Vision and Pattern Recognition · Computer Science 2024-09-02 M. Madhiarasan , Partha Pratim Roy

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users…

Computation and Language · Computer Science 2024-03-27 Dominik Wagner , Alexander Churchill , Siddharth Sigtia , Panayiotis Georgiou , Matt Mirsamadi , Aarshee Mishra , Erik Marchi

In this work, our goals are two fold: large-vocabulary continuous sign language recognition (CSLR), and sign language retrieval. To this end, we introduce a multi-task Transformer model, CSLR2, that is able to ingest a signing sequence and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Charles Raude , K R Prajwal , Liliane Momeni , Hannah Bull , Samuel Albanie , Andrew Zisserman , Gül Varol

Existing driving style recognition systems largely depend on low-level sensor-derived features for training, neglecting the rich semantic reasoning capability inherent to human experts. This discrepancy results in a fundamental misalignment…

Robotics · Computer Science 2026-05-06 Zhaokun Chen , Chaopeng Zhang , Xiaohan Li , Wenshuo Wang , Gentiane Venture , Junqiang Xi

Session-based recommendation (SBR) methods often rely on user behavior data, which can struggle with the sparsity of session data, limiting performance. Researchers have identified that beyond behavioral signals, rich semantic information…

Information Retrieval · Computer Science 2025-04-15 Shutong Qiao , Wei Zhou , Junhao Wen , Chen Gao , Qun Luo , Peixuan Chen , Yong Li

Weakly Supervised Semantic Segmentation (WSSS) with image level labels aims to produce pixel level predictions without requiring dense annotations. While recent approaches have leveraged generative models to augment existing data, they…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Wangyu Wu , Zhenhong Chen , Xiaowei Huang , Fei Ma , Jimin Xiao

Automatic sign language recognition (SLR) has become a key enabler of inclusive human-computer interaction, fostering seamless communication between deaf individuals and hearing communities. Despite significant advances in multimodal…

Human-Computer Interaction · Computer Science 2026-05-08 Xiaofang Xiao , Guangchao Li , Guangrong Zhao , Qi Lin , Wen Ma , Hongkai Wen , Yanxiang Wang , Yiran Shen

Textless spoken language models (SLMs) are generative models of speech that do not rely on text supervision. Most textless SLMs learn to predict the next semantic token, a discrete representation of linguistic content, and rely on a…

Computation and Language · Computer Science 2025-10-23 Ju-Chieh Chou , Jiawei Zhou , Karen Livescu