From the perspective of control theory, convolutional layers (of neural networks) are 2-D (or N-D) linear time-invariant dynamical systems. The usual representation of convolutional layers by the convolution kernel corresponds to the representation of a dynamical system by its impulse response. However, many analysis tools from control theory, e.g., involving linear matrix inequalities, require a state space representation. For this reason, we explicitly provide a state space representation of the Roesser type for 2-D convolutional layers with cinr1+coutr2 states, where cin/cout is the number of input/output channels of the layer and r1/r2 characterizes the width/length of the convolution kernel. This representation is shown to be minimal for cin=cout. We further construct state space representations for dilated, strided, and N-D convolutions.
Cite
@article{arxiv.2403.11938,
title = {State space representations of the Roesser type for convolutional layers},
author = {Patricia Pauli and Dennis Gramlich and Frank Allgöwer},
journal= {arXiv preprint arXiv:2403.11938},
year = {2024}
}