Related papers: A Note on Equivalent Conditions for Majorization
We propose new concentration inequalities for self-normalized martingales. The main idea is to introduce a suitable weighted sum of the predictable quadratic variation and the total quadratic variation of the martingale. It offers much more…
This letter reports two moment extensions of the entropy of a distribution. By understanding the traditional entropy as the average of the original distribution up to a random variable transformation, the traditional moments equation become…
Bayesian hierarchical models are frequently used in practical data analysis contexts. One interpretation of these models is that they provide an indirect way of assigning a prior for unknown parameters, through the introduction of…
Learning representations that capture the underlying data generating process is a key problem for data efficient and robust use of neural networks. One key property for robustness which the learned representation should capture and which…
We show that the naive application of the maximum entropy principle can yield answers which depend on the level of description, i.e. the result is not invariant under coarse-graining. We demonstrate that the correct approach, even for…
Chinese text recognition is more challenging than Latin text due to the large amount of fine-grained Chinese characters and the great imbalance over classes, which causes a serious overfitting problem. We propose to apply Maximum Entropy…
Iterative majorize-minimize (MM) (also called optimization transfer) algorithms solve challenging numerical optimization problems by solving a series of "easier" optimization problems that are constructed to guarantee monotonic descent of…
This paper clarifies the main research methods and ideas of the thesis [1,2,4]. The special calculation process is also realized by corresponding computer algorithm. Finally, we introduce zero rows sum case and give the corresponding…
The problem of membrane topology in the matrix model of M-theory is considered. The matrix regularization procedure, which makes a correspondence between finite-sized matrices and functions defined on a two-dimensional base space, is…
We develop several efficient algorithms for the classical \emph{Matrix Scaling} problem, which is used in many diverse areas, from preconditioning linear systems to approximation of the permanent. On an input $n\times n$ matrix $A$, this…
Although neural networks can solve very complex machine-learning problems, the theoretical reason for their generalizability is still not fully understood. Here we use Wang-Landau Mote Carlo algorithm to calculate the entropy (logarithm of…
In this paper we describe the alternative approach to the sample boundedness and continuity of stochastic processes. We show that the regularity of paths can be understood in terms of a distribution of the argument maximum. For a centered…
Regularization is one of the crucial ingredients of deep learning, yet the term regularization has various definitions, and regularization methods are often studied separately from each other. In our work we present a systematic, unifying…
We develop a family of reformulations of an arbitrary consistent linear system into a stochastic problem. The reformulations are governed by two user-defined parameters: a positive definite matrix defining a norm, and an arbitrary discrete…
The muon optimizer has picked up much attention as of late as a possible replacement to the seemingly omnipresent Adam optimizer. Recently, care has been taken to document the scaling laws of hyper-parameters under muon such as weight decay…
Estimating large covariance and precision matrices are fundamental in modern multivariate analysis. The problems arise from statistical analysis of large panel economics and finance data. The covariance matrix reveals marginal correlations…
We analyze statistical properties of the complex system with conditions which manifests through specific constraints on the column/row sum of the matrix elements. The presence of additional constraints besides symmetry leads to new…
This note discusses an interesting matrix factorization called the CUR Decomposition. We illustrate various viewpoints of this method by comparing and contrasting them in different situations. Additionally, we offer a new characterization…
We introduce a notion of complexity of diagrams (and in particular of objects and morphisms) in an arbitrary category, as well as a notion of complexity of functors between categories equipped with complexity functions. We discuss several…
The classification of complex data usually requires the composition of processing steps. Here, a major challenge is the selection of optimal algorithms for preprocessing and classification (including parameterizations). Nowadays, parts of…