English
Related papers

Related papers: Minimax Rates for Learning Pairwise Interactions i…

200 papers

When and how can an attention mechanism learn to selectively attend to informative tokens, thereby enabling detection of weak, rare, and sparsely located features? We address these questions theoretically in a sparse-token classification…

Machine Learning · Computer Science 2025-09-30 Nicholas Barnfield , Hugo Cui , Yue M. Lu

Attention layers -- which map a sequence of inputs to a sequence of outputs -- are core building blocks of the Transformer architecture which has achieved significant breakthroughs in modern artificial intelligence. This paper presents a…

Machine Learning · Computer Science 2023-07-24 Hengyu Fu , Tianyu Guo , Yu Bai , Song Mei

We present a novel approach for nonparametric regression using wavelet basis functions. Our proposal, $\texttt{waveMesh}$, can be applied to non-equispaced data with sample size not necessarily a power of 2. We develop an efficient proximal…

Machine Learning · Statistics 2019-03-13 Asad Haris , Noah Simon , Ali Shojaie

Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly understood, and many theoretical works have focused on the…

Machine Learning · Computer Science 2026-02-27 O. Duranthon , P. Marion , C. Boyer , B. Loureiro , L. Zdeborová

Pairwise learning is receiving increasing attention since it covers many important machine learning tasks, e.g., metric learning, AUC maximization, and ranking. Investigating the generalization behavior of pairwise learning is thus of…

Machine Learning · Computer Science 2021-11-10 Shaojie Li , Yong Liu

Systems of interacting particles or agents have wide applications in many disciplines such as Physics, Chemistry, Biology and Economics. These systems are governed by interaction laws, which are often unknown: estimating them from…

Machine Learning · Statistics 2020-07-16 Fei Lu , Mauro Maggioni , Sui Tang

Supervised learning is often affected by a covariate shift in which the marginal distributions of instances (covariates $x$) of training and testing samples $\mathrm{p}_\text{tr}(x)$ and $\mathrm{p}_\text{te}(x)$ are different but the label…

Machine Learning · Statistics 2023-06-12 José I. Segovia-Martín , Santiago Mazuelas , Anqi Liu

The success of self-attention lies in its ability to capture long-range dependencies and enhance context understanding, but it is limited by its computational complexity and challenges in handling sequential data with inherent…

Computation and Language · Computer Science 2025-05-05 Md Kowsher , Nusrat Jahan Prottasha , Chun-Nam Yu , Ozlem Ozmen Garibay , Niloofar Yousefi

Transfer learning for nonparametric regression is considered. We first study the non-asymptotic minimax risk for this problem and develop a novel estimator called the confidence thresholding estimator, which is shown to achieve the minimax…

Machine Learning · Statistics 2024-01-24 T. Tony Cai , Hongming Pu

We investigate the learning rate of multiple kernel leaning (MKL) with elastic-net regularization, which consists of an $\ell_1$-regularizer for inducing the sparsity and an $\ell_2$-regularizer for controlling the smoothness. We focus on a…

Machine Learning · Statistics 2011-07-14 Taiji Suzuki , Ryota Tomioka , Masashi Sugiyama

We propose a framework for the joint inference of network topology, multi-type interaction kernels, and latent type assignments in heterogeneous interacting particle systems from multi-trajectory data. This learning task is a challenging…

Machine Learning · Statistics 2026-02-05 Quanjun Lang , Xiong Wang , Fei Lu , Mauro Maggioni

We consider the nonparametric regression with a random design model, and we are interested in the adaptive estimation of the regression at a point $x\_0$ where the design is degenerate. When the design density is $\beta$-regularly varying…

Statistics Theory · Mathematics 2016-08-16 Stéphane Gaiffas

Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their implications for…

Enforcing orthonormal or isometric property for the weight matrices has been shown to enhance the training of deep neural networks by mitigating gradient exploding/vanishing and increasing the robustness of the learned networks. However,…

Machine Learning · Computer Science 2024-03-01 Zhen Qin , Xuwei Tan , Zhihui Zhu

Sparse additive models are families of $d$-variate functions that have the additive decomposition $f^* = \sum_{j \in S} f^*_j$, where $S$ is an unknown subset of cardinality $s \ll d$. In this paper, we consider the case where each…

Statistics Theory · Mathematics 2011-12-20 Garvesh Raskutti , Martin J. Wainwright , Bin Yu

In the brain, fine-scale correlations combine to produce macroscopic patterns of activity. However, as experiments record from larger and larger populations, we approach a fundamental bottleneck: the number of correlations one would like to…

Biological Physics · Physics 2024-02-02 Christopher W. Lynn , Qiwei Yu , Rich Pang , Stephanie E. Palmer , William Bialek

Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby…

Machine Learning · Computer Science 2026-05-11 Zhengkai Sun , Dibyakanti Kumar , Alejandro F Frangi , Anirbit Mukherjee , Mingfei Sun

Human learners have the natural ability to use knowledge gained in one setting for learning in a different but related setting. This ability to transfer knowledge from one task to another is essential for effective learning. In this paper,…

Statistics Theory · Mathematics 2019-06-10 T. Tony Cai , Hongji Wei

Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of the model capabilities. Yet, a theoretical characterization…

Machine Learning · Computer Science 2025-02-10 Hugo Cui , Freya Behrens , Florent Krzakala , Lenka Zdeborová

In this paper, we study last-iterate convergence of learning algorithms in bilinear saddle-point problems, a preferable notion of convergence that captures the day-to-day behavior of learning dynamics. We focus on the challenging setting…