English
Related papers

Related papers: Local Softmax and Global Weights in Non-Boolean Ev…

200 papers

By using fluid-kinetic simulations of confined and concentrated emulsion droplets, we investigate the nature of space non-homogeneity in soft-glassy dynamics and provide quantitative measurements of the statistical features of plastic…

Soft Condensed Matter · Physics 2014-02-19 R. Benzi , M. Sbragaglia , P. Perlekar , M. Bernaschi , S. Succi , F. Toschi

Localization of plastic strain induced by softening can be objectively described by a regularized plasticity model that postulates a dependence of the current yield stress on a nonlocal softening variable defined by a differential…

Materials Science · Physics 2012-02-02 Milan Jirasek , Jan Zeman , Jaroslav Vondrejc

Optimization analyses for cross-entropy training rely on local Taylor models of the loss to predict whether a proposed step will decrease the objective. These surrogates are reliable only inside the Taylor convergence radius of the true…

Machine Learning · Computer Science 2026-03-17 Piyush Sao

Three extensions and reinterpretations of nonclassical probabilities are reviewed. (i) We propose to generalize the probability axiom of quantum mechanics to self-adjoint positive operators of trace one. Furthermore, we discuss the…

Quantum Physics · Physics 2007-05-23 Karl Svozil

The biases present in training datasets have been shown to affect models for sentence pair classification tasks such as natural language inference (NLI) and fact verification. While fine-tuning models on additional data has been used to…

Computation and Language · Computer Science 2021-02-05 James Thorne , Andreas Vlachos

Neural collapse (NC) describes the structured geometry that emerges in the features and weights of trained classifiers. Recent theory suggests NC can be suboptimal in deep architectures, attributing this to an explicit low-rank bias from L2…

Machine Learning · Computer Science 2026-05-25 Connall Garrod , Jonathan P. Keating , Christos Thrampoulidis

The paper establishes generalization bounds for multitask deep neural networks using operator-theoretic techniques. The authors propose a tighter bound than those derived from conventional norm based methods by leveraging small condition…

Machine Learning · Computer Science 2026-05-29 Mahdi Mohammadigohari , Giuseppe Di Fatta , Giuseppe Nicosia , Panos M. Pardalos

The state of many physical, biological and socio-technical systems evolves by combining smooth local transitions and abrupt resetting events to a set of reference values. The inclusion of the resetting mechanism not only provides the…

Statistical Mechanics · Physics 2022-12-21 Oriol Artime

Self-normalized processes are basic to many probabilistic and statistical studies. They arise naturally in the the study of stochastic integrals, martingale inequalities and limit theorems, likelihood-based methods in hypothesis testing and…

Probability · Mathematics 2009-09-29 Victor H. de la Peña , Michael J. Klass , Tze Leung Lai

We investigate $\beta$-functions of quantum gravity using dimensional regularisation. In contrast to minimal subtraction, a non-minimal renormalisation scheme is employed which is sensitive to power-law divergences from mass terms or…

High Energy Physics - Theory · Physics 2024-09-17 Yannick Kluth

A number of prototypical optimization problems in multi-agent systems (e.g., task allocation and network load-sharing) exhibit a highly local structure: that is, each agent's decision variables are only directly coupled to few other agent's…

Multiagent Systems · Computer Science 2020-03-04 Robin Brown , Federico Rossi , Kiril Solovey , Michael T. Wolf , Marco Pavone

Spatial heterogeneity in the elastic properties of soft random solids is examined via vulcanization theory. The spatial heterogeneity in the \emph{structure} of soft random solids is a result of the fluctuations locked-in at their…

Disordered Systems and Neural Networks · Physics 2011-12-08 Xiaoming Mao , Paul M. Goldbart , Xiangjun Xing , Annette Zippelius

Weight averaging has become a standard technique for enhancing model performance. However, methods such as Stochastic Weight Averaging (SWA) and Latest Weight Averaging (LAWA) often require manually designed procedures to sample from the…

Machine Learning · Computer Science 2025-02-17 Peng Wang , Shengchao Hu , Zerui Tao , Guoxia Wang , Dianhai Yu , Li Shen , Quan Zheng , Dacheng Tao

Expectation Maximization (EM) is among the most popular algorithms for maximum likelihood estimation, but it is generally only guaranteed to find its stationary points of the log-likelihood objective. The goal of this article is to present…

Machine Learning · Computer Science 2018-10-29 Ji Xu , Daniel Hsu , Arian Maleki

The results of the renormalization group are commonly advertised as the existence of power law singularities near critical points. The classic predictions are often violated and logarithmic and exponential corrections are treated on a…

We provide the first proof of convergence for normalized error feedback algorithms across a wide range of machine learning problems. Despite their popularity and efficiency in training deep neural networks, traditional analyses of error…

Machine Learning · Computer Science 2024-10-23 Sarit Khirirat , Abdurakhmon Sadiev , Artem Riabinin , Eduard Gorbunov , Peter Richtárik

We derive the N=1 supersymmetric extension for a class of weakly nonlocal four dimensional gravitational theories.The construction is explicitly done in the superspace and the tree-level perturbative unitarity is explicitly proved both in…

High Energy Physics - Theory · Physics 2016-05-13 Stefano Giaccari , Leonardo Modesto

In this paper, we introduce a threshold-based framework for multiclass classification that generalizes the standard argmax rule. This is done by replacing the probabilistic interpretation of softmax outputs with a geometric one on the…

Machine Learning · Computer Science 2025-12-02 Francesco Marchetti , Edoardo Legnaro , Sabrina Guastavino

Softmax function is widely used in artificial neural networks for multiclass classification, multilabel classification, attention mechanisms, etc. However, its efficacy is often questioned in literature. The log-softmax loss has been shown…

Machine Learning · Computer Science 2020-11-24 Kunal Banerjee , Vishak Prasad C , Rishi Raj Gupta , Karthik Vyas , Anushree H , Biswajit Mishra

We prove that, for finite-arm bandits with linear function approximation, the global convergence of policy gradient (PG) methods depends on inter-related properties between the policy update and the representation. textcolor{blue}{First},…

Machine Learning · Computer Science 2025-04-04 Jincheng Mei , Bo Dai , Alekh Agarwal , Mohammad Ghavamzadeh , Csaba Szepesvari , Dale Schuurmans