English
Related papers

Related papers: Grammatical Evolution with Restarts for Fast Fract…

200 papers

In this work, we study the convergence \emph{in high probability} of clipped gradient methods when the noise distribution has heavy tails, ie., with bounded $p$th moments, for some $1<p\le2$. Prior works in this setting follow the same…

Optimization and Control · Mathematics 2023-04-05 Ta Duy Nguyen , Alina Ene , Huy L. Nguyen

While large-scale neural language models, such as GPT2 and BART, have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms…

Computation and Language · Computer Science 2022-10-11 Jin Xu , Xiaojiang Liu , Jianhao Yan , Deng Cai , Huayang Li , Jian Li

It is argued that there is a need for fat-tailed distributions that become thin in the extreme tail. A 3-parameter distribution is introduced that visually resembles the t-distribution and interpolates between the normal distribution and…

Statistics Theory · Mathematics 2022-02-08 Rose D Baker

The distribution of data in the world (eg, internet, etc.) significantly differs from the well-curated datasets and is often over-populated with samples from common categories. The algorithms designed for well-curated datasets perform…

Machine Learning · Computer Science 2025-07-30 Harsh Rangwani

A particular direction of recent advance about stochastic deep-learning algorithms has been about uncovering a rather mysterious heavy-tailed nature of the stationary distribution of these algorithms, even when the data distribution is not…

Machine Learning · Computer Science 2022-04-28 Sayar Karmakar , Anirbit Mukherjee

We propose a construction procedure which generates a wide class of random evolving networks with fat-tailed degree distributions and an arbitrary clustering. This procedure applies the stochastic transformations of edges, which can be used…

Statistical Mechanics · Physics 2007-05-23 S. N. Dorogovtsev , J. F. F. Mendes , A. N. Samukhin

In this paper we consider a transformer with an $n$-gram structure, such as the one underlying ChatGPT. The transformer provides next word probabilities, which can be used to generate word sequences. We consider methods for computing word…

Machine Learning · Computer Science 2024-03-26 Yuchao Li , Dimitri Bertsekas

Using graphical methods based on a `lookdown' and pruned version of the {\em ancestral selection graph}, we obtain a representation of the type distribution of the ancestor in a two-type Wright-Fisher population with mutation and selection,…

Probability · Mathematics 2020-03-17 Ellen Baake , Ute Lenz , Anton Wakolbinger

We propose an analytical approach to the computation of tail probabilities of compound distributions whose individual components have heavy tails. Our approach is based on the contour integration method, and gives rise to a representation…

Computational Finance · Quantitative Finance 2017-10-04 Igor Halperin

In real-world scenarios, where knowledge distributions exhibit long-tail. Humans manage to master knowledge uniformly across imbalanced distributions, a feat attributed to their diligent practices of reviewing, summarizing, and correcting…

Computer Vision and Pattern Recognition · Computer Science 2024-09-16 Qihao Zhao , Yalun Dai , Shen Lin , Wei Hu , Fan Zhang , Jun Liu

It is difficult to anticipate the myriad challenges that a predictive model will encounter once deployed. Common practice entails a reactive, cyclical approach: model deployment, data mining, and retraining. We instead develop a proactive…

The restarted primal-dual hybrid gradient method (rPDHG) is a first-order method that has recently received significant attention for its computational effectiveness in solving linear program (LP) problems. Despite its impressive practical…

Optimization and Control · Mathematics 2026-02-17 Zikai Xiong

We establish sharp large deviation asymptotics for the maximum order statistic of independent and identically distributed heavy-tailed random variables, valid for all Borel subsets of the right tail. This result yields exact decay rates for…

Probability · Mathematics 2026-01-09 José M. Zapata

Reinforcement Learning (RL) for Large Language Models (LLMs) faces a fundamental tension: the numerical divergence between high-throughput inference engines and numerically precise training engines. Although these systems share the same…

Machine Learning · Computer Science 2026-02-09 Yingru Li , Jiawei Xu , Jiacai Liu , Yuxuan Tong , Ziniu Li , Tianle Cai , Ge Zhang , Qian Liu , Baoxiang Wang

Big data can easily be contaminated by outliers or contain variables with heavy-tailed distributions, which makes many conventional methods inadequate. To address this challenge, we propose the adaptive Huber regression for robust…

Statistics Theory · Mathematics 2018-10-11 Qiang Sun , Wenxin Zhou , Jianqing Fan

The purpose of this paper is to introduce a new Markov chain Monte Carlo method and exhibit its efficiency by simulation and high-dimensional asymptotic theory. Key fact is that our algorithm has a reversible proposal transition kernel,…

Methodology · Statistics 2014-12-22 Kengo Kamatani

In this paper we discuss the problem of the estimation of extreme event occurrence probability for data drawn from some multifractal process. We also study the heavy (power-law) tail behavior of probability density function associated with…

Statistical Mechanics · Physics 2009-11-11 Jean-Francois Muzy , Emmanuel Bacry , Alexey Kozhemyak

Evolutionary algorithms have been widely used for a range of stochastic optimization problems in order to address complex real-world optimization problems. We consider the knapsack problem where the profits involve uncertainties. Such a…

Neural and Evolutionary Computing · Computer Science 2022-04-13 Aneta Neumann , Yue Xie , Frank Neumann

Recent works have proposed incorporating heavy-tailed (HT) noise into diffusion- and flow-based generative models, with the goals of better recovering the tails of target distributions and improving generative diversity. This motivation is…

Machine Learning · Computer Science 2026-05-14 Hamza Cherkaoui , Hélène Halconruy , Antonio Ocello

For genetic algorithms using a bit-string representation of length~$n$, the general recommendation is to take $1/n$ as mutation rate. In this work, we discuss whether this is really justified for multimodal functions. Taking jump functions…

Neural and Evolutionary Computing · Computer Science 2017-03-23 Benjamin Doerr , Huu Phuoc Le , Régis Makhmara , Ta Duy Nguyen