Light-Bound Transformers: Hardware-Anchored Robustness for Silicon-Photonic Computer Vision Systems
Abstract
Deploying Vision Transformers (ViTs) on near-sensor analog accelerators demands training pipelines that are explicitly aligned with device-level noise and energy constraints. We introduce a compact framework for silicon-photonic execution of ViTs that integrates measured hardware noise, robust attention training, and an energy-aware processing flow. We first characterize bank-level noise in microring-resonator (MR) arrays, including fabrication variation, thermal drift, and amplitude noise, and convert these measurements into closed-form, activation-dependent variance proxies for attention logits and feed-forward activations. Using these proxies, we develop Chance-Constrained Training (CCT), which enforces variance-normalized logit margins to bound attention rank flips, and a noise-aware LayerNorm that stabilizes feature statistics without changing the optical schedule. These components yield a practical ``measure model train run'' pipeline that optimizes accuracy under noise while respecting system energy limits. Hardware-in-the-loop experiments with MR photonic banks show that our approach restores near-clean accuracy under realistic noise budgets, with no in-situ learning or additional optical MACs.
Cite
@article{arxiv.2604.04330,
title = {Light-Bound Transformers: Hardware-Anchored Robustness for Silicon-Photonic Computer Vision Systems},
author = {Xuming Chen and Deniz Najafi and Chengwei Zhou and Pietro Mercati and Arman Roohi and Mohsen Imani and Mahdi Nikdast and Shaahin Angizi and Gourav Datta},
journal= {arXiv preprint arXiv:2604.04330},
year = {2026}
}
Comments
Accepted at Design Automation Conference (DAC) 2026