Related papers: HMC and gradient flow with machine-learned classic…
Continual learning (CL) presents a fundamental challenge in training neural networks on sequential tasks without experiencing catastrophic forgetting. Traditionally, the dominant approach in CL has been gradient-based optimization, where…
We generalize the Hamiltonian Monte Carlo algorithm with a stack of neural network layers and evaluate its ability to sample from different topologies in a two dimensional lattice gauge theory. We demonstrate that our model is able to…
Sampling a target probability distribution with an unknown normalization constant is a fundamental challenge in computational science and engineering. Recent work shows that algorithms derived by considering gradient flows in the space of…
Recent studies on diffusion-based sampling methods have shown that Langevin Monte Carlo (LMC) algorithms can be beneficial for non-convex optimization, and rigorous theoretical guarantees have been proven for both asymptotic and finite-time…
We propose an optimization algorithm called Frictionless Hamiltonian Descent, which is a direct counterpart of classical Hamiltonian Monte Carlo in sampling. We analyze Frictionless Hamiltonian Descent for strongly convex quadratic…
This paper presents a new gradient flow dissipation geometry over non-negative and probability measures. This is motivated by a principled construction that combines the unbalanced optimal transport and interaction forces modeled by…
We present a novel method for guaranteeing linear momentum in learned physics simulations. Unlike existing methods, we enforce conservation of momentum with a hard constraint, which we realize via antisymmetrical continuous convolutional…
Electronic structure methods offer in principle accurate predictions of molecular properties, however, their applicability is limited by computational costs. Empirical methods are cheaper, but come with inherent approximations and are…
Non-zero topological charge is prohibited in the chiral limit of gauge-fermion systems because any instanton would create a zero mode of the Dirac operator. On the lattice, however, the geometric $Q_\text{geom}=\langle F{\tilde F}\rangle…
We recently demonstrated that standard fixed-time lattice random-walk models cannot be modified to properly represent biased diffusion processes in more than two dimensions. The origin of this fundamental limitation appears to be the fact…
As the continuum limit is approached, lattice QCD simulations tend to get trapped in the topological charge sectors of field space and may consequently give biased results in practice. We propose to bypass this problem by imposing open…
We report on the determination of the gradient flow scales in $N_f=2+1$ QCD using highly improved staggered quark (HISQ) ensembles generated by the HotQCD Collaboration for bare gauge couplings ranging from $\beta = 6.423$ to $8.400$. Using…
It has become customary to use a smoothing algorithm called "gradient flow" to fix the lattice spacing in a simulation, through a parameter called $t_0$. It is shown that in order to keep the length $t_0$ fixed with respect to mesonic or…
At fine lattice spacings, Markov chain Monte Carlo simulations of QCD and other gauge theories with or without fermions are plagued by slow modes that give rise to large autocorrelation times. This can lead to simulation runs that are…
Classifier-free guidance (CFG) is a widely used technique for controllable generation in diffusion and flow-based models. Despite its empirical success, CFG relies on a heuristic linear extrapolation that is often sensitive to the guidance…
Scaling deep reinforcement learning networks is challenging and often results in degraded performance, yet the root causes of this failure mode remain poorly understood. Several recent works have proposed mechanisms to address this, but…
Federated learning (FL) is an emerging machine learning method that can be applied in mobile edge systems, in which a server and a host of clients collaboratively train a statistical model utilizing the data and computation resources of the…
A variety of lattice discretisations of continuum actions has been considered, usually requiring the correct classical continuum limit. Here we discuss "weird" lattice formulations without that property, namely lattice actions that are…
The paper surveys recent progresses in understanding the dynamics and loss landscape of the gradient flow equations associated to deep linear neural networks, i.e., the gradient descent training dynamics (in the limit when the step size…
It is well understood that neural networks with carefully hand-picked weights provide powerful function approximation and that they can be successfully trained in over-parametrized regimes. Since over-parametrization ensures zero training…