Related papers: Switching-Geometry Analysis of Deflated Q-Value It…
This paper investigates an iterative rank-one decomposition scheme for positive operators on a Hilbert space based on a residual-weighted congruence update. At each step the operator is compressed along a chosen unit vector while remaining…
This paper focuses on solving a stochastic variational inequality (SVI) problem under relaxed smoothness assumption for a class of structured non-monotone operators. The SVI problem has attracted significant interest in the machine learning…
Passive intelligent reconfigurable surfaces (IRS) are becoming an attractive component of cellular networks due to their ability of shaping the propagation environment and thereby improving the coverage. While passive IRS nodes incorporate…
Recently discovered polyhedral structures of the value function for finite state-action discounted Markov decision processes (MDP) shed light on understanding the success of reinforcement learning. We investigate the value function polytope…
Soft Q-learning is a variation of Q-learning designed to solve entropy regularized Markov decision problems where an agent aims to maximize the entropy regularized value function. Despite its empirical success, there have been limited…
Value decomposition (VD) methods have achieved remarkable success in cooperative multi-agent reinforcement learning (MARL). However, their reliance on the max operator for temporal-difference (TD) target calculation leads to systematic…
The recognition network in deep latent variable models such as variational autoencoders (VAEs) relies on amortized inference for efficient posterior approximation that can scale up to large datasets. However, this technique has also been…
The objective of this paper is to investigate the finite-time analysis of a Q-learning algorithm applied to two-player zero-sum Markov games. Specifically, we establish a finite-time analysis of both the minimax Q-learning algorithm and the…
This paper deals with the parameter estimation problem of the Single-Input-Single-Output (SISO) switched Hammerstein system. Suppose that the switching law is arbitrary but can be observed online. All subsystems are parameterized and the…
This paper presents novel adaptive reduced-rank filtering algorithms based on joint iterative optimization of adaptive filters. The novel scheme consists of a joint iterative optimization of a bank of full-rank adaptive filters that…
We consider a rank-one symmetric matrix corrupted by additive noise. The rank-one matrix is formed by an $n$-component unknown vector on the sphere of radius $\sqrt{n}$, and we consider the problem of estimating this vector from the…
We study the partially asymmetric exclusion process with open boundaries. We generalise the matrix approach previously used to solve the special case of total asymmetry and derive exact expressions for the partition sum and currents valid…
Evaluating the entanglement spectrum is essential for characterizing exotic quantum phases such as quantum criticality and topological order. However, for large quantum many-body systems, this task is hindered by the exponential measurement…
Recovering a low-rank tensor from incomplete information is a recurring problem in signal processing and machine learning. The most popular convex relaxation of this problem minimizes the sum of the nuclear norms of the unfoldings of the…
This contribution reviews the parallel dynamics of Q-Ising neural networks for various architectures: extremely diluted asymmetric, layered feedforward, extremely diluted symmetric, and fully connected. Using a probabilistic signal-to-noise…
In this paper, we propose a general scale invariant approach for sparse signal recovery via the minimization of the $q$-ratio sparsity. When $1 < q \leq \infty$, both the theoretical analysis based on $q$-ratio constrained minimal singular…
We make a complete variational treatment of rank-one Proper Generalised Decomposition for separable fractional partial differential equations with conformable derivatives. The setting is Hilbertian, the energy is induced by a symmetric…
In a discounted reward Markov Decision Process (MDP), the objective is to find the optimal value function, i.e., the value function corresponding to an optimal policy. This problem reduces to solving a functional equation known as the…
This paper proposes an elegant optimization framework consisting of a mix of linear-matrix-inequality and second-order-cone constraints. The proposed framework generalizes the semidefinite relaxation (SDR) enabled solution to the typical…
Deep learning is an increasingly popular approach for inverting surface wave dispersion curves to obtain Vs profiles. However, its generalizability is constrained by the depth and velocity scales of training data. We propose a unified deep…