English
Related papers

Related papers: Sign-IDD: Iconicity Disentangled Diffusion for Sig…

200 papers

Lifelong person re-identification (LReID) aims to train a generalizable model with sequentially collected data. However, such models often suffer from semantic drift, limited adaptability, and catastrophic forgetting as new domains emerge.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Wen Wen , Hao Chen , Shiliang Zhang

Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction through spoken or written language. This spoken-language-centered interface can limit access…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Youngmin Kim , Kyobin Choo , Jiwoo Park , Minseo Kim , Chanyoung Kim , Junhyeok Kim , Seong Jae Hwang

Zero-shot 3D anomaly detection aims to identify anomalies without access to training data from target categories. However, existing methods mainly rely on projecting 3D observations into multi-view representations that primarily capture…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Letian Bai , Xuanming Cao , Juan Du , Chengyu Tao

Accurate recognition and interpretation of sign language are crucial for enhancing communication accessibility for deaf and hard of hearing individuals. However, current approaches of Isolated Sign Language Recognition (ISLR) often face…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Karina Kvanchiani , Roman Kraynov , Elizaveta Petrova , Petr Surovcev , Aleksandr Nagaev , Alexander Kapitanov

Large-scale short-video search ranking models are typically trained on sparse co-occurrence signals over hashed item identifiers (HIDs). While effective at memorizing frequent interactions, such ID-based models struggle to generalize to…

Information Retrieval · Computer Science 2026-04-14 Guowen Li , Yuepeng Zhang , Shunyu Zhang , Yi Zhang , Xiaoze Jiang , Yi Wang , Jingwei Zhuo

Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rotate paradigms. To this end, we introduce Pose-Aware Diffusion (PAD), a novel end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Zihan Zhou , Luxi Chen , Jingzhi Zhou , Yuhao Wan , Min Zhao , Baoyu Fan , Chongxuan Li

Sign language translation (SLT) aims to translate natural language from sign language videos, serving as a vital bridge for inclusive communication. While recent advances leverage powerful visual backbones and large language models, most…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Wenfang Wu , Tingting Yuan , Yupeng Li , Daling Wang , Xiaoming Fu

Generating natural and linguistically accurate sign language avatars remains a formidable challenge. Current Sign Language Production (SLP) frameworks face a stark trade-off: direct text-to-pose models suffer from regression-to-the-mean…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Jianhe Low , Alexandre Symeonidis-Herzig , Maksym Ivashechkin , Ozge Mercanoglu Sincan , Richard Bowden

The recent surge in large language models has automated translations of spoken and written languages. However, these advances remain largely inaccessible to American Sign Language (ASL) users, whose language relies on complex visual cues.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Daniel Perkins , Davis Hunter , Dhrumil Patel , Galen Flanagan

In this paper, we show different fine-tuning methods for Stable Diffusion XL; this includes inference steps, and caption customization for each image to align with generating images in the style of a commercial 2D icon training set. We also…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Youssef Sultan , Jiangqin Ma , Yu-Ying Liao

A key challenge in lifelong imitation learning (LIL) is enabling agents to acquire new skills from expert demonstrations while retaining prior knowledge. This requires preserving the low-dimensional manifolds and geometric structures that…

Machine Learning · Computer Science 2026-03-11 Kaushik Roy , Giovanni D'urso , Nicholas Lawrance , Brendan Tidd , Peyman Moghadam

Focal cortical dysplasia (FCD) lesions in epilepsy FLAIR MRI are subtle and scarce, making joint image--mask generative modeling prone to instability and memorization. We propose SLIM-Diff, a compact joint diffusion model whose main…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Mario Pascual-González , Ariadna Jiménez-Partinen , R. M. Luque-Baena , Fátima Nagib-Raya , Ezequiel López-Rubio

Recent studies on facial expression editing have obtained very promising progress. On the other hand, existing methods face the constraint of requiring a large amount of expression labels which are often expensive and time-consuming to…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Rongliang Wu , Shijian Lu

Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fidelity but are computationally expensive, while…

Graphics · Computer Science 2025-10-10 Yongzhi Li , Saining Zhang , Yibing Chen , Boying Li , Yanxin Zhang , Xiaoyu Du

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent…

Computer Vision and Pattern Recognition · Computer Science 2020-10-14 Jianrong Wang , Tong Wu , Shanyu Wang , Mei Yu , Qiang Fang , Ju Zhang , Li Liu

Existing 3D anomaly detection methods are built on a rigid prior: normal geometry is pose-invariant and can be canonicalized through registration or alignment. This prior does not hold for articulated objects with hinge or sliding joints,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Jinye Gan , Bozhong Zheng , Xiaohao Xu , Junye Ren , Zixuan Zhang , Na Ni , Yingna Wu

Most of the vision-based sign language research to date has focused on Isolated Sign Language Recognition (ISLR), where the objective is to predict a single sign class given a short video clip. Although there has been significant progress…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Ryan Wong , Necati Cihan Camgöz , Richard Bowden

Diffusion-based approaches have recently achieved strong results in face swapping, offering improved visual quality over traditional GAN-based methods. However, even state-of-the-art models often suffer from fine-grained artifacts and poor…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Weston Bondurant , Arkaprava Sinha , Hieu Le , Srijan Das , Stephanie Schuckers

Recent advancements in deep generative models, particularly with the application of CLIP (Contrastive Language Image Pretraining) to Denoising Diffusion Probabilistic Models (DDPMs), have demonstrated remarkable effectiveness in text to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Cristian Sbrolli , Paolo Cudrano , Matteo Matteucci

Reconstruction-based approaches have achieved remarkable outcomes in anomaly detection. The exceptional image reconstruction capabilities of recently popular diffusion models have sparked research efforts to utilize them for enhanced…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Haoyang He , Jiangning Zhang , Hongxu Chen , Xuhai Chen , Zhishan Li , Xu Chen , Yabiao Wang , Chengjie Wang , Lei Xie