English
Related papers

Related papers: Large data limits and scaling laws for tSNE

200 papers

We establish convergence of the training dynamics of residual neural networks (ResNets) to their joint infinite depth L, hidden width M, and embedding dimension D limit. Specifically, we consider ResNets with two-layer perceptron blocks in…

Machine Learning · Statistics 2026-03-23 Louis-Pierre Chaintron , Lénaïc Chizat , Javier Maass

In contrast to classical techniques for exploratory analysis of high-dimensional data sets, such as principal component analysis (PCA), neighbor embedding (NE) techniques tend to better preserve the local structure/topology of…

Machine Learning · Statistics 2022-09-07 Roman Josef Rainer , Michael Mayr , Johannes Himmelbauer , Ramin Nikzad-Langerodi

Information networks are ubiquitous and are ideal for modeling relational data. Networks being sparse and irregular, network embedding algorithms have caught the attention of many researchers, who came up with numerous embeddings algorithms…

Machine Learning · Computer Science 2020-09-25 Junshan Wang , Yilun Jin , Guojie Song , Xiaojun Ma

We propose a linear-complexity method for sampling from truncated multivariate normal (TMVN) distributions with high fidelity by applying nearest-neighbor approximations to a product-of-conditionals decomposition of the TMVN density. To…

Computation · Statistics 2024-06-26 Jian Cao , Matthias Katzfuss

In this paper, we address the generalization of deep neural network (DNN) based speech enhancement to unseen noise conditions for the case that training data is limited in size and diversity. To gain more insights, we analyze the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-18 Robert Rehr , Timo Gerkmann

High-dimensional limit theorems have been shown useful to derive tuning rules for finding the optimal scaling in random-walk Metropolis algorithms. The assumptions under which weak convergence results are proved are however restrictive: the…

Methodology · Statistics 2022-02-16 Sebastian M Schmon , Philippe Gagnon

We study the problem of exact support recovery for high-dimensional sparse linear regression under independent Gaussian design when the signals are weak, rare, and possibly heterogeneous. Under a suitable scaling of the sample size and…

Statistics Theory · Mathematics 2023-07-19 Saptarshi Roy , Ambuj Tewari , Ziwei Zhu

Unsupervised dimensionality reduction is one of the commonly used techniques in the field of high dimensional data recognition problems. The deep autoencoder network which constrains the weights to be non-negative, can learn a low…

Computer Vision and Pattern Recognition · Computer Science 2020-09-18 Anyong Qin , Zhaowei Shang , Zhuolin Tan , Taiping Zhang , Yuan Yan Tang

Neural networks are becoming an increasingly important tool in applications. However, neural networks are not widely used in statistical genetics. In this paper, we propose a new neural networks method called expectile neural networks. When…

Statistics Theory · Mathematics 2020-11-04 Jinghang Lin , Xiaoxi Shen , Qing Lu

Stochastic iterative algorithms, including stochastic gradient descent (SGD) and stochastic gradient Langevin dynamics (SGLD), are widely utilized for optimization and sampling in large-scale and high-dimensional problems in machine…

Machine Learning · Statistics 2025-01-22 Xiaoyu Wang , Mikolaj J. Kasprzak , Jeffrey Negrea , Solesne Bourguin , Jonathan H. Huggins

In this paper, we address the issue on non-asymptotic convergence bounds of Euler-type schemes associated with non-dissipative SDEs. On the one hand, for non-degenerate SDEs with super-linear drifts, we propose a novel modified Euler scheme…

Probability · Mathematics 2025-12-09 Jianhai Bao , Jiaqing Hao , Panpan Ren

Many approaches in machine learning rely on a weighted graph to encode the similarities between samples in a dataset. Entropic affinities (EAs), which are notably used in the popular Dimensionality Reduction (DR) algorithm t-SNE, are…

Machine Learning · Computer Science 2023-10-31 Hugues Van Assel , Titouan Vayer , Rémi Flamary , Nicolas Courty

The self-consistent expansion (SCE) is a powerful technique for obtaining perturbative solutions to problems in statistical physics but it suffers from a subtle problem - too much freedom! The SCE can be used to generate an enormous number…

Statistical Mechanics · Physics 2024-07-12 Chanania Steinbock , Eytan Katzav

We study the scaling limits of stochastic gradient descent (SGD) with constant step-size in the high-dimensional regime. We prove limit theorems for the trajectories of summary statistics (i.e., finite-dimensional functions) of SGD as the…

Machine Learning · Statistics 2023-08-21 Gerard Ben Arous , Reza Gheissari , Aukosh Jagannath

Neural networks often require large amounts of data to generalize and can be ill-suited for modeling small and noisy experimental datasets. Standard network architectures trained on scarce and noisy data will return predictions that violate…

Machine Learning · Computer Science 2021-05-20 Gregory Barber , Mulugeta A. Haile , Tzikang Chen

It is generally recognized that finite learning rate (LR), in contrast to infinitesimal LR, is important for good generalization in real-life deep nets. Most attempted explanations propose approximating finite-LR SGD with Ito Stochastic…

Machine Learning · Computer Science 2021-06-18 Zhiyuan Li , Sadhika Malladi , Sanjeev Arora

We examine the Bayes-consistency of a recently proposed 1-nearest-neighbor-based multiclass learning algorithm. This algorithm is derived from sample compression bounds and enjoys the statistical advantages of tight, fully empirical…

Machine Learning · Computer Science 2019-06-27 Aryeh Kontorovich , Sivan Sabato , Roi Weiss

Modern supervised learning techniques, particularly those using deep nets, involve fitting high dimensional labelled data sets with functions containing very large numbers of parameters. Much of this work is empirical. Interesting phenomena…

Machine Learning · Statistics 2018-05-30 Partha P Mitra

Sparse signal recovery is one of the most fundamental problems in various applications, including medical imaging and remote sensing. Many greedy algorithms based on the family of hard thresholding operators have been developed to solve the…

Signal Processing · Electrical Eng. & Systems 2023-06-09 Rachel Grotheer , Shuang Li , Anna Ma , Deanna Needell , Jing Qin

In this paper, we propose a Tensor Train Neighborhood Preserving Embedding (TTNPE) to embed multi-dimensional tensor data into low dimensional tensor subspace. Novel approaches to solve the optimization problem in TTNPE are proposed. For…

Machine Learning · Computer Science 2018-05-09 Wenqi Wang , Vaneet Aggarwal , Shuchin Aeron