Related papers: Finite size scaling in neural networks
We numerically investigate the heterogeneity in cluster sizes in the two-dimensional Ising model and verify its scaling form recently proposed in the context of percolation problems [Phys. Rev. E 84, 010101(R) (2011)]. The scaling exponents…
{}From a finite-size scaling (FSS) theory of cumulants of the order parameter at phase coexistence points, we reconstruct the scaling of the moments. Assuming that the cumulants allow a reconstruction of the free energy density no better…
Finite-size criteria have emerged as an effective tool for deriving spectral gaps in higher-dimensional frustration-free quantum spin systems. We quantitatively improve the existing finite-size criteria by introducing a novel subsystem…
We survey the application of a relatively new branch of statistical physics--"community detection"-- to data mining. In particular, we focus on the diagnosis of materials and automated image segmentation. Community detection describes the…
We study the scaling properties of critical particle systems confined by a potential. Using renormalization-group arguments, we show that their critical behavior can be cast in the form of a trap-size scaling, resembling finite-size scaling…
Recent studies point to the potential storage of a large number of patterns in the celebrated Hopfield associative memory model, well beyond the limits obtained previously. We investigate the properties of new fixed points to discover that…
We present a new model for time series classification, called the hidden-unit logistic model, that uses binary stochastic hidden units to model latent structure in the data. The hidden units are connected in a chain structure that models…
We reformulate the problem of encoding a multi-scale representation of a sequence in a language model by casting it in a continuous learning framework. We propose a hierarchical multi-scale language model in which short time-scale…
The universal behaviour of the directed percolation universality class is well understood, both the critical scaling as well as finite size scaling. This article focuses on the block (finite size) scaling of the order parameter and its…
We present a framework to define a large class of neural networks for which, by construction, training by gradient flow provably reaches arbitrarily low loss when the number of parameters grows. Distinct from the fixed-space global…
Scaling laws have shaped recent advances in machine learning by enabling predictable scaling of model performance based on model size, computation, and data volume. Concurrently, the rise in computational cost for AI has motivated model…
We propose an end-to-end learned image data hiding framework that embeds and extracts secrets in the latent representations of a generic neural compressor. By leveraging a perceptual loss function in conjunction with our proposed message…
In statistical setting of the pattern recognition problem the number of examples required to approximate an unknown labelling function is linear in the VC dimension of the target learning class. In this work we consider the question whether…
We investigate the problem of maintaining an encoded distributed storage system when some nodes contain adversarial errors. Using the error-correction capabilities that are built into the existing redundancy of the system, we propose a…
The push to train ever larger neural networks has motivated the study of initialization and training at large network width. A key challenge is to scale training so that a network's internal representations evolve nontrivially at all…
Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study…
The aim of this thesis is to compare the capacity of different models of neural networks. We start by analysing the problem solving capacity of a single perceptron using a simple combinatorial argument. After some observations on the…
After training complex deep learning models, a common task is to compress the model to reduce compute and storage demands. When compressing, it is desirable to preserve the original model's per-example decisions (e.g., to go beyond top-1…
Critical phenomena on scale-free networks with a degree distribution $p_k \sim k^{-\lambda}$ exhibit rich finite-size effects due to its structural heterogeneity. We systematically study the finite-size scaling of percolation and identify…
Besides its original spin representation, the Ising model is known to have the Fortuin-Kasteleyn (FK) bond and loop representations, of which the former was recently shown to exhibit two upper critical dimensions $(d_c=4,d_p=6)$. Using a…