English
Related papers

Related papers: Predicting with Proxies: Transfer Learning in High…

200 papers

Explainable artificially intelligent (XAI) systems form part of sociotechnical systems, e.g., human+AI teams tasked with making decisions. Yet, current XAI systems are rarely evaluated by measuring the performance of human+AI teams on…

Artificial Intelligence · Computer Science 2020-01-24 Zana Buçinca , Phoebe Lin , Krzysztof Z. Gajos , Elena L. Glassman

Numerous variable selection methods rely on a two-stage procedure, where a sparsity-inducing penalty is used in the first stage to predict the support, which is then conveyed to the second stage for estimation or inference purposes. In this…

Applications · Statistics 2015-05-28 Jean-Michel Bécu , Yves Grandvalet , Christophe Ambroise , Cyril Dalmasso

Sampling from the posterior is a key technical problem in Bayesian statistics. Rigorous guarantees are difficult to obtain for Markov Chain Monte Carlo algorithms of common use. In this paper, we study an alternative class of algorithms…

Statistics Theory · Mathematics 2024-08-26 Andrea Montanari , Yuchen Wu

A machine learning model may exhibit discrimination when used to make decisions involving people. One potential cause for such outcomes is that the model uses a statistical proxy for a protected demographic attribute. In this paper we…

Machine Learning · Computer Science 2018-11-29 Samuel Yeom , Anupam Datta , Matt Fredrikson

We introduce sparse random projection, an important dimension-reduction tool from machine learning, for the estimation of discrete-choice models with high-dimensional choice sets. Initially, high-dimensional data are compressed into a…

Machine Learning · Statistics 2016-04-21 Khai X. Chiong , Matthew Shum

Statistical prediction models are often trained on data from different probability distributions than their eventual use cases. One approach to proactively prepare for these shifts harnesses the intuition that causal mechanisms should…

Machine Learning · Computer Science 2023-08-02 Bijan Mazaheri , Atalanti Mastakouri , Dominik Janzing , Michaela Hardt

High-dimensional vector autoregression with measurement error is frequently encountered in a large variety of scientific and business applications. In this article, we study statistical inference of the transition matrix under this model.…

Methodology · Statistics 2020-09-18 Xiang Lyu , Jian Kang , Lexin Li

Adaptive collection of data is commonplace in applications throughout science and engineering. From the point of view of statistical inference however, adaptive data collection induces memory and correlation in the samples, and poses…

Methodology · Statistics 2020-05-07 Yash Deshpande , Adel Javanmard , Mohammad Mehrabi

Aggregate outcome variables collected through surveys and administrative records are often subject to systematic measurement error. For instance, in disaster loss databases, county-level losses reported may differ from the true damages due…

Machine Learning · Computer Science 2026-03-17 Saketh Vishnubhatla , Shu Wan , Andre Harrison , Adrienne Raglin , Huan Liu

This paper proposes a novel two-step strategy for testing the goodness-of-fit of parametric regression models in ultra-high dimensional sparse settings, where the predictor dimension far exceeds the sample size. This regime usually renders…

Methodology · Statistics 2025-12-30 Falong Tan , Jie Liu , Heng Peng , Lixing Zhu

We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small…

Machine Learning · Statistics 2013-04-11 Hamed Firouzi , Bala Rajaratnam , Alfred Hero

Among the most popular variable selection procedures in high-dimensional regression, Lasso provides a solution path to rank the variables and determines a cut-off position on the path to select variables and estimate coefficients. In this…

Methodology · Statistics 2018-06-19 X. Jessie Jeng , Huimin Peng , Wenbin Lu

Predictive biomarkers are playing an essential role in precision medicine. Identifying an optimal cutoff to select patient subsets with greater benefit from treatment is critical and more challenging for predictive biomarkers measured with…

Applications · Statistics 2025-04-29 Chi Zhang , Wei Shi , Spencer Woody , Qing Liu

Methods for Projection Pursuit aim to facilitate the visual exploration of high-dimensional data by identifying interesting low-dimensional projections. A major challenge is the design of a suitable quality metric of projections, commonly…

Machine Learning · Computer Science 2015-11-30 Tijl De Bie , Jefrey Lijffijt , Raul Santos-Rodriguez , Bo Kang

We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as they are often fine-tuned on human preferences that may not…

Bayesian variable selection methods are powerful techniques for fitting and inferring on sparse high-dimensional linear regression models. However, many are computationally intensive or require restrictive prior distributions on model…

Methodology · Statistics 2023-10-10 Alexander C. McLain , Anja Zgodic , Howard Bondell

The explosion of large-scale data in fields such as finance, e-commerce, and social media has outstripped the processing capabilities of single-machine systems, driving the need for distributed statistical inference methods. Traditional…

Machine Learning · Statistics 2024-09-02 Jingguo Lan , Hongmei Lin , Xueqin Wang

Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which training recipe to apply. Making these decisions well requires reliable performance…

Computation and Language · Computer Science 2026-05-19 Arkil Patel , Siva Reddy , Marius Mosbach , Dzmitry Bahdanau

Latent factor model estimation typically relies on either using domain knowledge to manually pick several observed covariates as factor proxies, or purely conducting multivariate analysis such as principal component analysis. However, the…

Methodology · Statistics 2023-01-04 Runzhe Wan , Yingying Li , Wenbin Lu , Rui Song

Proxy-based metric learning losses are superior to pair-based losses due to their fast convergence and low training complexity. However, existing proxy-based losses focus on learning class-discriminative features while overlooking the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Zhibo Yang , Muhammet Bastan , Xinliang Zhu , Doug Gray , Dimitris Samaras