English

Vision without Images: End-to-End Computer Vision from Single Compressive Measurements

Computer Vision and Pattern Recognition 2025-09-03 v3 Artificial Intelligence

Abstract

Snapshot Compressed Imaging (SCI) offers high-speed, low-bandwidth, and energy-efficient image acquisition, but remains challenged by low-light and low signal-to-noise ratio (SNR) conditions. Moreover, practical hardware constraints in high-resolution sensors limit the use of large frame-sized masks, necessitating smaller, hardware-friendly designs. In this work, we present a novel SCI-based computer vision framework using pseudo-random binary masks of only 8×\times8 in size for physically feasible implementations. At its core is CompDAE, a Compressive Denoising Autoencoder built on the STFormer architecture, designed to perform downstream tasks--such as edge detection and depth estimation--directly from noisy compressive raw pixel measurements without image reconstruction. CompDAE incorporates a rate-constrained training strategy inspired by BackSlash to promote compact, compressible models. A shared encoder paired with lightweight task-specific decoders enables a unified multi-task platform. Extensive experiments across multiple datasets demonstrate that CompDAE achieves state-of-the-art performance with significantly lower complexity, especially under ultra-low-light conditions where traditional CMOS and SCI pipelines fail.

Keywords

Cite

@article{arxiv.2501.15122,
  title  = {Vision without Images: End-to-End Computer Vision from Single Compressive Measurements},
  author = {Fengpu Pan and Heting Gao and Jiangtao Wen and Yuxing Han},
  journal= {arXiv preprint arXiv:2501.15122},
  year   = {2025}
}
R2 v1 2026-06-28T21:17:22.904Z