Related papers: Transfer entropy and O-information to detect grokk…
In grokking, a model first fits the training data while test accuracy remains low, and only later begins to generalize. We ask whether this transition can be localized from observed training trajectories before the test accuracy rises, and…
The data representation in a machine-learning model strongly influences its performance. This becomes even more important for quantum machine learning models implemented on noisy intermediate scale quantum (NISQ) devices. Encoding high…
Transformers trained on modular arithmetic exhibit sharp transitions between memorization, generalization, and collapse. We show that weight decay acts as a scalar empirical control parameter for these regimes, and introduce two cheap…
Many important quantities in quantum information science, such as entropy and entanglement, are non-linear functions of the density matrix and cannot be expressed as operator observables. Standard open-system approaches evolve only a single…
The maximum entropy encoding framework provides a unified perspective for many non-contrastive learning methods like SimSiam, Barlow Twins, and MEC. Inspired by this framework, we introduce Matrix-SSL, a novel approach that leverages matrix…
Multiscale systems are ubiquitous in science and technology, but are notoriously challenging to simulate as short spatiotemporal scales must be appropriately linked to emergent bulk physics. When expensive high-dimensional dynamical systems…
We study holographic aspects of mixed state entanglement measures in various large $N$ top-down as well as bottom-up confining models. For the top-down models, we consider wrapped $D3$ and $D4$ branes gravity solutions whereas, for the…
The cross-entropy loss commonly used in deep learning is closely related to the defining properties of optimal representations, but does not enforce some of the key properties. We show that this can be solved by adding a regularization…
This paper introduces a high-order Markov chain task to investigate how transformers learn to integrate information from multiple past positions with varying statistical significance. We demonstrate that transformers learn this task…
Incremental learning suffers from two challenging problems; forgetting of old knowledge and intransigence on learning new knowledge. Prediction by the model incrementally learned with a subset of the dataset are thus uncertain and the…
Grokking, the sudden generalization that occurs after prolonged overfitting, is a surprising phenomenon challenging our understanding of deep learning. Although significant progress has been made in understanding grokking, the reasons…
Over-parameterized deep neural networks are able to achieve excellent training accuracy while maintaining a small generalization error. It has also been found that they are able to fit arbitrary labels, and this behaviour is referred to as…
We study the macroscopic entanglement properties of a low dimensional quantum spin system by investigating its magnetic properties at low temperatures and high magnetic fields. The tempera- ture and magnetic field dependence of entanglement…
Learning-based approaches to robotic manipulation are limited by the scalability of data collection and accessibility of labels. In this paper, we present a multi-task domain adaptation framework for instance grasping in cluttered scenes by…
We demonstrate the use of tensor networks for image classification with the TensorNetwork open source library. We explain in detail the encoding of image data into a matrix product state form, and describe how to contract the network in a…
This paper reveals the intrinsic structure of Matrix Product States (MPS) by establishing their deep connection to entangled hidden Markov models (EHMMs). It is demonstrated that a significant class of MPS can be derived as the outcomes of…
Transfer entropy (TE) is an information theoretic measure that reveals the directional flow of information between processes, providing valuable insights for a wide range of real-world applications. This work proposes Transfer Entropy…
We investigate fundamental connections between thermodynamics and quantum information theory. First, we show that the operational framework of thermal operations is nonequivalent to the framework of Gibbs-preserving maps, and we comment on…
In this work, we demonstrate that a simple two-layer neural network with standard activation functions can learn an arbitrary word operation in any finite group, provided sufficient width is available and exhibits grokking while doing so.…
Sampling problems have emerged as a central avenue for demonstrating quantum advantage on noisy intermediate-scale quantum devices. However, physical noise can fundamentally alter their computational complexity, often making them…