Related papers: Designing a Linearized Potential Function in Neura…
Optimization methods play a crucial role in modern machine learning, powering the remarkable empirical achievements of deep learning models. These successes are even more remarkable given the complex non-convex nature of the loss landscape…
Recent works have shown that on sufficiently over-parametrized neural nets, gradient descent with relatively large initialization optimizes a prediction function in the RKHS of the Neural Tangent Kernel (NTK). This analysis leads to global…
In this paper, we consider the problem of estimating Tsallis entropy from a given data set. We propose four different estimators for Tsallis entropy measure based on higher-order sample spacings, and then discuss estimation of Tsallis…
Tsallis relative operator entropy is defined and then its properties are given. Shannon inequality and its reverse one in Hilbert space operators derived by T.Furuta \cite{Fu:par} are extended in terms of the parameter of the Tsallis…
The Lagrangian technique of Niven (2004, Physica A, 334(3-4): 444) is used to determine the constrained forms of the Tsallis entropy function - i.e. Lagrangian functions in which the probabilities of each state are independent - for each…
Evolutionary computation can be used to optimize several different aspects of neural network architectures. For instance, the TaylorGLO method discovers novel, customized loss functions, resulting in improved performance, faster training,…
Compressed Counting (CC)} was recently proposed for approximating the $\alpha$th frequency moments of data streams, for $0<\alpha \leq 2$. Under the relaxed strict-Turnstile model, CC dramatically improves the standard algorithm based on…
A way to pose the entropic uncertainty principle for trace-preserving super-operators is presented. It is based on the notion of extremal unraveling of a super-operator. For given input state, different effects of each unraveling result in…
In this paper, a shallow Ritz-type neural network for solving elliptic equations with delta function singular sources on an interface is developed. There are three novel features in the present work; namely, (i) the delta function…
Temporal Graph Learning (TGL) has become a prevalent technique across diverse real-world applications, especially in domains where data can be represented as a graph and evolves over time. Although TGL has recently seen notable progress in…
The canonical probability distribution function (pdf) obtained by optimizing the Tsallis entropy under the linear mean energy constraint (first formalism) or the escort mean energy constraint (third formalism) suffer self-referentiality. In…
We propose and analyze a regularization approach for structured prediction problems. We characterize a large class of loss functions that allows to naturally embed structured outputs in a linear space. We exploit this fact to design…
Following the work on Shannon entropy together with the principle of maximum entropy, Luo & Singh (J. Hydrol. Eng., 2011, 16(4): 303-315) and Singh & Luo (J. Hydrol. Eng., 2011, 16(9): 725-735) explored the concept of non-extensive Tsallis…
Maximum Tsallis entropy (MTE) framework in reinforcement learning has gained popularity recently by virtue of its flexible modeling choices including the widely used Shannon entropy and sparse entropy. However, non-Shannon entropies suffer…
We present a probabilistic framework for nonlinearities, based on doubly truncated Gaussian distributions. By setting the truncation points appropriately, we are able to generate various types of nonlinearities within a unified framework,…
The primary objective of learning methods is generalization. Classic uniform generalization bounds, which rely on VC-dimension or Rademacher complexity, fail to explain the significant attribute that over-parameterized models in deep…
Recent years have seen the emergence of nonlinear methods for solving partial differential equations (PDEs), such as physics-informed neural networks (PINNs). While these approaches often perform well in practice, their theoretical analysis…
Fusing and balancing multi-modal inputs from novel sensors for dense prediction tasks, particularly semantic segmentation, is critically important yet remains a significant challenge. One major limitation is the tendency of multi-modal…
We propose a new way of thinking about deep neural networks, in which the linear and non-linear components of the network are naturally derived and justified in terms of principles in probability theory. In particular, the models…
Numerical experiments support the interesting conjecture that statistical methods be applicable not only to fully-chaotic systems, but also at the edge of chaos by using Tsallis' generalizations of the standard exponential and entropy. In…