English
Related papers

Related papers: DA-Font: Few-Shot Font Generation via Dual-Attenti…

200 papers

The increasing difficulty in accurately detecting forged images generated by AIGC(Artificial Intelligence Generative Content) poses many risks, necessitating the development of effective methods to identify and further locate forged areas.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Yang Liu , Xiaofei Li , Jun Zhang , Shengze Hu , Jun Lei

Transformer-based object detectors often struggle with occlusions, fine-grained localization, and computational inefficiency caused by fixed queries and dense attention. We propose DAMM, Dual-stream Attention with Multi-Modal queries, a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Noreen Anwar , Guillaume-Alexandre Bilodeau , Wassim Bouachir

Can a pre-trained generator be adapted to the hybrid of multiple target domains and generate images with integrated attributes of them? In this work, we introduce a new task -- Few-shot Hybrid Domain Adaptation (HDA). Given a source…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Hengjia Li , Yang Liu , Linxuan Xia , Yuqi Lin , Tu Zheng , Zheng Yang , Wenxiao Wang , Xiaohui Zhong , Xiaobo Ren , Xiaofei He

Generative AI for automated glaucoma diagnostic report generation faces two predominant challenges: content redundancy in narrative outputs and inadequate highlighting of pathologically significant features including optic disc cupping,…

Computational Engineering, Finance, and Science · Computer Science 2025-10-14 Cheng Huang , Weizheng Xie , Zeyu Han , Tsengdar Lee , Karanjit Kooner , Jui-Ka Wang , Ning Zhang , Jia Zhang

Fine-grained image classification is a challenging problem, since the difficulty of finding discriminative features. To handle this circumstance, basically, there are two ways to go. One is use attention based method to focus on informative…

Computer Vision and Pattern Recognition · Computer Science 2020-01-08 ZiChao Dong , JiLong Wu , TingTing Ren , Yue Wang , MengYing Ge

Traditional sentiment analysis has long been a unimodal task, relying solely on text. This approach overlooks non-verbal cues such as vocal tone and prosody that are essential for capturing true emotional intent. We introduce Dynamic…

Computation and Language · Computer Science 2025-09-30 Sadia Abdulhalim , Muaz Albaghdadi , Moshiur Farazi

Understanding not only where drivers look but also why their attention shifts is essential for interpretable human-AI collaboration in autonomous driving. Driver attention is not purely perceptual but semantically structured. Thus,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Kaiser Hamid , Can Cui , Khandakar Ashrafi Akbar , Ziran Wang , Nade Liang

Font generation is a difficult and time-consuming task, especially in those languages using ideograms that have complicated structures with a large number of characters, such as Chinese. To solve this problem, few-shot font generation and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Haibin He , Xinyuan Chen , Chaoyue Wang , Juhua Liu , Bo Du , Dacheng Tao , Yu Qiao

Most existing text-to-image generation methods adopt a multi-stage modular architecture which has three significant problems: 1) Training multiple networks increases the run time and affects the convergence and stability of the generative…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Zhenxing Zhang , Lambert Schomaker

Optical Character Recognition (OCR) is essential in applications such as document processing, license plate recognition, and intelligent surveillance. However, existing OCR models often underperform in real-world scenarios due to irregular…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Inho Jake Park , Jaehoon Jay Jeong , Ho-Sang Jo

We propose the Multi-Head Density Adaptive Attention Mechanism (DAAM), a novel probabilistic attention framework that can be used for Parameter-Efficient Fine-tuning (PEFT), and the Density Adaptive Transformer (DAT), designed to enhance…

Machine Learning · Computer Science 2024-10-01 Georgios Ioannides , Aman Chadha , Aaron Elkins

While recent progress has significantly boosted few-shot classification (FSC) performance, few-shot object detection (FSOD) remains challenging for modern learning systems. Existing FSOD systems follow FSC approaches, ignoring critical…

Computer Vision and Pattern Recognition · Computer Science 2021-09-17 Tung-I Chen , Yueh-Cheng Liu , Hung-Ting Su , Yu-Cheng Chang , Yu-Hsiang Lin , Jia-Fong Yeh , Wen-Chin Chen , Winston H. Hsu

To accomplish the punctuation restoration task, most existing approaches focused on leveraging extra information (e.g., part-of-speech tags) or addressing the class imbalance problem. Recent works have widely applied the transformer-based…

Computation and Language · Computer Science 2022-04-12 Yangjun Wu , Kebin Fang , Yao Zhao

Zero-shot learning (ZSL) aims to recognize novel classes through transferring shared semantic knowledge (e.g., attributes) from seen classes to unseen classes. Recently, attention-based methods have exhibited significant progress which…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Jinwei Han , Yingguo Gao , Zhiwen Lin , Ke Yan , Shouhong Ding , Yuan Gao , Gui-Song Xia

Effective deep feature extraction via feature-level fusion is crucial for multimodal object detection. However, previous studies often involve complex training processes that integrate modality-specific features by stacking multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Lei Hao , Lina Xu , Chang Liu , Yanni Dong

Domain adaptive object detection aims to leverage the knowledge learned from a labeled source domain to improve the performance on an unlabeled target domain. Prior works typically require the access to the source domain data for…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Han Sun , Rui Gong , Konrad Schindler , Luc Van Gool

One of the important research topics in image generative models is to disentangle the spatial contents and styles for their separate control. Although StyleGAN can generate content feature vectors from random noises, the resulting spatial…

Computer Vision and Pattern Recognition · Computer Science 2021-07-26 Gihyun Kwon , Jong Chul Ye

Due to the advancement of Generative Adversarial Networks (GAN), Autoencoders, and other AI technologies, it has been much easier to create fake images such as "Deepfakes". More recent research has introduced few-shot learning, which uses a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Young Oh Bang , Simon S. Woo

In this paper, we focus on the semantic image synthesis task that aims at transferring semantic label maps to photo-realistic images. Existing methods lack effective semantic constraints to preserve the semantic information and ignore the…

Computer Vision and Pattern Recognition · Computer Science 2020-09-01 Hao Tang , Song Bai , Nicu Sebe

Generating new fonts is a time-consuming and labor-intensive task, especially in a language with a huge amount of characters like Chinese. Various deep learning models have demonstrated the ability to efficiently generate new fonts with a…

Computer Vision and Pattern Recognition · Computer Science 2022-12-16 Haoyang He , Xin Jin , Angela Chen