English

VDLF-Net: Variational Feature Fusion for Adaptive and Few-Shot Visual Learning

Computer Vision and Pattern Recognition 2026-04-28 v1

Abstract

This paper introduces VDLF-Net, which attaches a compact VAE to a multi-scale CNN backbone. Latent vectors and softmax-gate support the backbone feature maps, while 2\ell_2-normalized embeddings from the gated maps contribute toward supervised classification or episodic few-shot prediction. Under standard CIFAR-100 and Mini-ImageNet protocols, VDLF-Net demonstrates an improved performance over ResNet-50 Enhanced, VGG-16, Prototypical Networks, and Matching Networks. Extensive ablations show that removing the fine-resolution scale has the greatest impact on VDLF-Net's performance. At the same time, KL and reconstruction at the chosen α\alpha pose a minor performance reduction, demonstrating that performance gains over classical episodic baselines mainly originate from the full VDLF-Net architecture and training strategy.

Keywords

Cite

@article{arxiv.2604.23641,
  title  = {VDLF-Net: Variational Feature Fusion for Adaptive and Few-Shot Visual Learning},
  author = {Jiawei Yan},
  journal= {arXiv preprint arXiv:2604.23641},
  year   = {2026}
}
R2 v1 2026-07-01T12:35:40.427Z