Related papers: Improving Multigrid and Conventional Relaxation Al…
Stochastic-gradient-based optimization has been a core enabling methodology in applications to large-scale problems in machine learning and related areas. Despite the progress, the gap between theory and practice remains significant, with…
This paper proposes a novel analysis for the Scaffold algorithm, a popular method for dealing with data heterogeneity in federated learning. While its convergence in deterministic settings--where local control variates mitigate client…
We propose and study an improved method to calculate the fermionic determinant of dynamical configurations. The evaluation or at least stochastic estimation of ratios of fermionic determinants is essential for a recently proposed updating…
Federated Learning (FL) incurs high communication overhead, which can be greatly alleviated by compression for model updates. Yet the tradeoff between compression and model accuracy in the networked environment remains unclear and, for…
We study deterministic and stochastic primal-dual sub-gradient algorithms for distributed optimization of a separable objective function with global inequality constraints. In both algorithms, the norm of the Lagrangian multipliers are…
Of all the vector fields surrounding the minima of recurrent learning setups, the gradient field with its exploding and vanishing updates appears a poor choice for optimization, offering little beyond efficient computability. We seek to…
This paper presents distributed conjugate gradient algorithms for distributed parameter estimation and spectrum estimation over wireless sensor networks. In particular, distributed conventional conjugate gradient (CCG) and modified…
We present a bosonization procedure which replaces fermions with generalized spin variables subject to local constraints. It requires that the number of Majorana modes per lattice site matches the coordination number modulo two. If this…
We provide tight finite-time convergence bounds for gradient descent and stochastic gradient descent on quadratic functions, when the gradients are delayed and reflect iterates from $\tau$ rounds ago. First, we show that without stochastic…
This note addresses the problem of computing fermion propagators in a broad variety of strongly correlated systems that can be mapped onto the theory of fermions coupled to an (over)damped bosonic mode. A number of the previously applied…
By using Symanzik's improvement program, we study on-shell improved lattice QCD with staggered fermions. We find that there are as many as 15 independent lattice operators of dimension of six~(including both gauge and fermion operators)…
This paper presents an adaptive combination strategy for distributed learning over diffusion networks. Since learning relies on the collaborative processing of the stochastic information at the dispersed agents, the overall performance can…
This paper is focused on the convergence analysis of an adaptive stochastic collocation algorithm for the stationary diffusion equation with parametric coefficient. The algorithm employs sparse grid collocation in the parameter domain…
Algebraic Multigrid (AMG) methods are often robust and effective solvers for solving the large and sparse linear systems that arise from discretized PDEs and other problems, relying on heuristic graph algorithms to achieve their…
This paper studies the co-design of actuators, sensors, and communication in the distributed setting, where a networked plant is partitioned into subsystems each equipped with a sub-controller interacting with other sub-controllers. The…
In this paper, we investigate the impact of compression on stochastic gradient algorithms for machine learning, a technique widely used in distributed and federated learning. We underline differences in terms of convergence rates between…
Applying domain decomposition to the lattice Dirac operator and the associated quark propagator, we arrive at expressions which, with the proper insertion of random sources therein, can provide improvement to the estimation of the…
Tensor network techniques have proved to be powerful tools that can be employed to explore the large scale dynamics of lattice systems. Nonetheless, the redundancy of degrees of freedom in lattice gauge theories (and related models) poses a…
The Gumbel-Max trick is the basis of many relaxed gradient estimators. These estimators are easy to implement and low variance, but the goal of scaling them comprehensively to large combinatorial distributions is still outstanding. Working…
Lagrangian decomposition (LD) is a relaxation method that provides a dual bound for constrained optimization problems by decomposing them into more manageable sub-problems. This bound can be used in branch-and-bound algorithms to prune the…