Related papers: Distributional Spectral Diagnostics for Localizing…
In lean premixed combustors, flame stabilization is an important operational concern that can affect efficiency, robustness and pollutant formation. The focus of this paper is on flame lift-off and re-attachment to the nozzle of a swirl…
We propose a new viewpoint on the study of localization transitions in disordered quantum systems, showing how critical properties can be seen also as a geometric transition in the data space generated by the classically encoded…
We explore statistical fluctuations over the ensemble of quantum trajectories in a model of two-dimensional free fermions subject to projective monitoring of local charge across the measurement-induced phase transition. Our observables are…
Ensuring generalization to unseen environments remains a challenge. Domain shift can lead to substantially degraded performance unless shifts are well-exercised within the available training environments. We introduce a simple robust…
We demonstrate the existence of a complexity phase transition in neural networks by studying the grokking phenomenon, where networks suddenly transition from memorization to generalization long after overfitting their training data. To…
Switchback experiments--alternating treatment and control over time--are widely used when unit-level randomization is infeasible, outcomes are aggregated, or user interference is unavoidable. In practice, experimentation must support fast…
Using the level--spacing distribution and the total probability function of the numbers of levels in a given energy interval we analyze the crossover of the level statistics between the delocalized and the localized regimes. By numerically…
This paper proposes deception as a mechanism for out-of-distribution (OOD) generalization: by learning data representations that make training data appear independent and identically distributed (iid) to an observer, we can identify stable…
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accuracy in the training process. It is often taken as an example…
Pre-trained model assessment for transfer learning aims to identify the optimal candidate for the downstream tasks from a model hub, without the need of time-consuming fine-tuning. Existing advanced works mainly focus on analyzing the…
Geographic distribution shift arises when the distribution of locations on Earth in a training dataset is different from what is seen at inference time. Using standard empirical risk minimization (ERM) in this setting can lead to uneven…
A key challenge for the machine learning community is to understand and accelerate the training dynamics of deep networks that lead to delayed generalisation and emergent robustness to input perturbations, also known as grokking. Prior work…
Many problems in computational science and engineering become one-to-many after coarse graining, partial observation, or inverse reconstruction: a resolved state may not determine a unique subgrid forcing, a structural descriptor may not…
When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining…
Estimating uncertainty in deep learning models is critical for reliable decision-making in high-stakes applications such as medical imaging. Prior research has established that the difference between an input sample and its reconstructed…
We discuss here in detail a new analytical random walk approach to calculating the phase-diagram for spatially extended systems with multiplicative noise. We use the Anderson localization problem as an example. The transition from…
While the phenomenon of grokking, i.e., delayed generalization, has been studied extensively, it remains an open problem whether there is a mathematical framework that characterizes what kind of features will emerge, how and in which…
Recent work has shown that the performance of machine learning models can vary substantially when models are evaluated on data drawn from a distribution that is close to but different from the training distribution. As a result, predicting…
In many real applications of statistical learning, collecting sufficiently many training data is often expensive, time-consuming, or even unrealistic. In this case, a transfer learning approach, which aims to leverage knowledge from a…
We consider the quickest change-point detection problem where the aim is to detect the onset of a pre-specified drift in "live"-monitored standard Brownian motion; the change-point is assumed unknown (nonrandom). The topic of interest is…