Related papers: Overlap Gap and Computational Thresholds in the Sq…
Artificial neural networks are functions depending on a finite number of parameters typically encoded as weights and biases. The identification of the parameters of the network from finite samples of input-output pairs is often referred to…
The UHF wave function may be written as a spin-contaminated \textit{pair} wave function of the APSG form, and the overlap of the alpha and beta corresponding orbitals of the UHF solution can be taken as a proxy for the strength of the…
The permutation symmetry of neurons in each layer of a deep neural network gives rise not only to multiple equivalent global minima of the loss function, but also to first-order saddle points located on the path between the global minima.…
Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory does not faithfully explain the behavior of practical…
We consider a spin-orbit coupled system of particles in an external trap that is represented by a deformed harmonic oscillator potential. The spin-orbit interaction is a Rashba interaction that does not commute with the trapping potential…
Quadratically constrained quadratic programs (QCQPs) are a fundamental class of optimization problems well-known to be NP-hard in general. In this paper we study conditions under which the standard semidefinite program (SDP) relaxation of a…
We consider neural networks with a single hidden layer and non-decreasing homogeneous activa-tion functions like the rectified linear units. By letting the number of hidden units grow unbounded and using classical non-Euclidean…
We describe a model for s-wave collisions between ground state atoms in optical lattices, considering especially the limits of quasi-one and two dimensional axisymmetric harmonic confinement. When the atomic interactions are modelled by an…
We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher-student setup. We focus on the extensive-width regime, where the teacher and student network widths scale proportionally with…
We consider the scenario where a 4-lattice constant, rotationally symmetric charge density wave (CDW) is present in the underdoped cuprates. We prove a theorem that puts strong constraint on the possible form factor of such a CDW. We…
We study numerically the geometric entanglement in the Laughlin wave function, which is of great importance in condensed matter physics. The Slater determinant having the largest overlap with the Laughlin wave function is constructed by an…
The most important hint of physics beyond the Standard Model (SM) from the 1995 precision electroweak data is that the most precisely measured quantities, the total, leptonic and hadronic decay widths of the $Z$ and the effective weak…
We discuss the isospin-breaking effects on threshold cusp structures in multichannel scattering near two-body thresholds. In hadronic systems with isospin symmetry, two or more nearly degenerate thresholds can appear, and their small…
We provide a closed form upper bound formulation for the average pairwise-error probability (PEP) of selective decode and forward (SDF) cooperation protocol for a keyhole (pinhole) channel condition. We have employed orthogonal space-time…
Motivated by a wide-spread use of convex optimization techniques, convexity properties of bit error rate of the maximum likelihood detector operating in the AWGN channel are studied for arbitrary constellations and bit mappings, which also…
We study support recovery for a $k \times k$ principal submatrix with elevated mean $\lambda/N$, hidden in an $N\times N$ symmetric mean zero Gaussian matrix. Here $\lambda>0$ is a universal constant, and we assume $k = N \rho$ for some…
Hard-label classification is usually trained with smooth surrogate losses, most prominently softmax cross-entropy. We isolate an asymptotic mechanism by which this mismatch between smooth surrogate and discrete labels produces power-law…
We introduce WARP (Weight-space Adaptive Recurrent Prediction), a simple yet powerful model that unifies weight-space learning with linear recurrence to redefine sequence modeling. Unlike conventional recurrent neural networks (RNNs) which…
This paper provides a theoretical justification of the superior classification performance of deep rectifier networks over shallow rectifier networks from the geometrical perspective of piecewise linear (PWL) classifier boundaries. We show…
The secrecy rate region of wiretap interference channels with a multi-antenna passive eavesdropper is studied under receiver energy harvesting constraints. To stay operational in the network, the legitimate receivers demand energy alongside…