English
Related papers

Related papers: Compositional regression using principal nested sp…

200 papers

Representing token embeddings as probability distributions over learned manifolds allows for more flexible contextual inference, reducing representational rigidity while enhancing semantic granularity. Comparative evaluations demonstrate…

Computation and Language · Computer Science 2025-04-25 Christopher Nightingale , Dominic Lavington , Jonathan Thistlethwaite , Sebastian Penhaligon , Thomas Belinski , David Boldo

Confidence sets from i.i.d. data are constructed for the extrinsic mean of a probabilty measure P on spheres, real projective spaces, and complex projective spaces, as well as Grassmann manifolds, with the latter three embedded by the…

Statistics Theory · Mathematics 2016-02-15 Thomas Hotz , Florian Kelma

Regression models for compositional data are common in several areas of knowledge. As in other classes of regression models, it is desirable to perform diagnostic analysis in these models using residuals that are approximately standard…

Methodology · Statistics 2024-03-21 Gustavo H. A. Pereira , Jianwen Cai

We study harmonic map regression, a nonparametric estimator for manifold-valued responses, that penalizes the empirical Fr\'echet risk by the Dirichlet energy. By connecting penalized regression to the theory of harmonic maps, the estimator…

Statistics Theory · Mathematics 2026-04-13 Xiaoyu Chen

Principal component analysis (PCA) aims at estimating the direction of maximal variability of a high-dimensional dataset. A natural question is: does this task become easier, and estimation more accurate, when we exploit additional…

Information Theory · Computer Science 2014-06-19 Andrea Montanari , Emile Richard

The nested error regression model is a useful tool for analyzing clustered (grouped) data, and is especially used in small area estimation. The classical nested error regression model assumes normality of random effects and error terms, and…

Methodology · Statistics 2016-05-16 Shonosuke Sugasawa , Tatsuya Kubokawa

In many applications, particularly in the natural sciences, the available high-dimensional set of features may contain variables that are not correlated with the response under consideration. Such irrelevant features can, in certain cases,…

Statistics Theory · Mathematics 2025-07-28 Gianluca Finocchio , Tatyana Krivobokova

Gaussian processes (GPs) are very widely used for modeling of unknown functions or surfaces in applications ranging from regression to classification to spatial processes. Although there is an increasingly vast literature on applications,…

Methodology · Statistics 2017-06-28 Lizhen Lin , Mu Niu , Pokman Cheung , David Dunson

Unsupervised text embedding has shown great power in a wide range of NLP tasks. While text embeddings are typically learned in the Euclidean space, directional similarity is often more effective in tasks such as word similarity and document…

Computation and Language · Computer Science 2019-11-05 Yu Meng , Jiaxin Huang , Guangyuan Wang , Chao Zhang , Honglei Zhuang , Lance Kaplan , Jiawei Han

Leveraging the compositional nature of our world to expedite learning and facilitate generalization is a hallmark of human perception. In machine learning, on the other hand, achieving compositional generalization has proven to be an…

Machine Learning · Computer Science 2023-07-13 Thaddäus Wiedemer , Prasanna Mayilvahanan , Matthias Bethge , Wieland Brendel

Gaussian Process (GP) regression is a powerful nonparametric Bayesian framework, but its performance depends critically on the choice of covariance kernel. Selecting an appropriate kernel is therefore central to model quality, yet remains…

Machine Learning · Computer Science 2026-01-14 Md Shafiqul Islam , Shakti Prasad Padhy , Douglas Allaire , Raymundo Arróyave

In many applications, data can be heterogeneous in the sense of spanning latent groups with different underlying distributions. When predictive models are applied to such data the heterogeneity can affect both predictive performance and…

Machine Learning · Statistics 2022-05-04 Thomas Lartigue , Sach Mukherjee

The parameter space of nonnegative trigonometric sums (NNTS) models for circular data is the surface of a hypersphere; thus, constructing regression models for a circular-dependent variable using NNTS models can comprise fitting great…

Methodology · Statistics 2025-01-10 J. J. Fernández-Durán , M. M. Gregorio-Domínguez

Driven by a wide range of applications, many principal subspace estimation problems have been studied individually under different structural constraints. This paper presents a unified framework for the statistical analysis of a general…

Statistics Theory · Mathematics 2020-11-17 T. Tony Cai , Hongzhe Li , Rong Ma

A new method is proposed for variable screening, variable selection and prediction in linear regression problems where the number of predictors can be much larger than the number of observations. The method involves minimizing a penalized…

Statistics Theory · Mathematics 2017-09-14 D. Vasiliu , T. Dey , I. L. Dryden

The paper introduces a new estimation method for the standard linear regression model. The procedure is not driven by the optimisation of any objective function rather, it is a simple weighted average of slopes from observation pairs. The…

Econometrics · Economics 2024-02-27 Felix Chan , Laszlo Matyas

Principal Components Regression (PCR) is a traditional tool for dimension reduction in linear regression that has been both criticized and defended. One concern about PCR is that obtaining the leading principal components tends to be…

Statistics Theory · Mathematics 2017-10-10 Martin Slawski

Representation learning plays a central role in structuring internal embeddings to capture the statistical properties of language, influencing the coherence and contextual consistency of generated text. Statistical Coherence Alignment is…

Computation and Language · Computer Science 2025-08-11 Jonathan Gale , Godfrey Aldington , Harriet Thistlewood , Thomas Tattershall , Basil Wentworth , Vincent Enoasmo

Graph-structured data is ubiquitous in scientific domains, where models often face imbalanced learning settings. In imbalanced regression, domain preferences focus on specific target value ranges that represent the most scientifically…

Machine Learning · Computer Science 2025-07-15 Brenda Nogueira , Gabe Gomes , Meng Jiang , Nitesh V. Chawla , Nuno Moniz