English
Related papers

Related papers: Predicting Batting Averages in Specific Matchups U…

200 papers

An important challenge in statistical analysis lies in controlling the bias of estimators due to the ever-increasing data size and model complexity. Approximate numerical methods and data features like censoring and misclassification often…

Statistics Theory · Mathematics 2020-11-17 Stéphane Guerrier , Mucyo Karemera , Samuel Orso , Maria-Pia Victoria-Feser , Yuming Zhang

We present an algorithm for learning the intrinsic value of a batted ball in baseball. This work addresses the fundamental problem of separating the value of a batted ball at contact from factors such as the defense, weather, and ballpark…

Applications · Statistics 2016-03-02 Glenn Healey

When multitudes of features can plausibly be associated with a response, both privacy considerations and model parsimony suggest grouping them to increase the predictive power of a regression model. Specifically, the identification of…

Methodology · Statistics 2024-05-07 Brandon Woosuk Park , Anand N. Vidyashankar , Tucker S. McElroy

Sufficient dimension reduction methods often require stringent conditions on the joint distribution of the predictor, or, when such conditions are not satisfied, rely on marginal transformation or reweighting to fulfill them approximately.…

Statistics Theory · Mathematics 2009-04-27 Bing Li , Yuexiao Dong

Simultaneous variable selection and statistical inference is challenging in high-dimensional data analysis. Most existing post-selection inference methods require explicitly specified regression models, which are often linear, as well as…

Methodology · Statistics 2026-03-19 Shangyuan Ye , Shauna Rakshe , Ye Liang

Estimating ballpark effects and team defense in baseball is challenging because batted-ball outcomes are influenced by multiple factors, including contact quality, ballpark environment, defensive performance, and random variation. In this…

Applications · Statistics 2026-03-24 Jhe-Jia Wu , Tian-Li Yan , Ting-Li Chen

In this study, we propose a projection estimation method for large-dimensional matrix factor models with cross-sectionally spiked eigenvalues. By projecting the observation matrix onto the row or column factor space, we simplify factor…

Methodology · Statistics 2020-12-04 Long Yu , Yong He , Xin-bing Kong , Xinsheng Zhang

We investigate different methods for regularizing quantile regression when predicting either a subset of quantiles or the full inverse CDF. We show that minimizing an expected pinball loss over a continuous distribution of quantiles is a…

Machine Learning · Statistics 2021-02-11 Taman Narayan , Serena Wang , Kevin Canini , Maya Gupta

Latent variable models represent a useful tool for the analysis of complex data when the constructs of interest are not observable. A problem related to these models is that the integrals involved in the likelihood function cannot be solved…

Methodology · Statistics 2015-03-05 Silvia Bianconcini , Silvia Cagnone , Dimitris Rizopoulos

We consider forecasting a single time series using a large number of predictors in the presence of a possible nonlinear forecast function. Assuming that the predictors affect the response through the latent factors, we propose to first…

Statistics Theory · Mathematics 2021-04-22 Wei Luo , Lingzhou Xue , Jiawei Yao , Xiufan Yu

Two new approaches for checking the dimension of the basis functions when using penalized regression smoothers are presented. The first approach is a test for adequacy of the basis dimension based on an estimate of the residual variance…

Methodology · Statistics 2016-02-23 Natalya Pya , Simon N Wood

Single-cell RNA sequencing allows the quantification of gene expression at the individual cell level, enabling the study of cellular heterogeneity and gene expression dynamics. Dimensionality reduction is a common preprocessing step…

Computation · Statistics 2025-10-14 Cristian Castiglione , Alexandre Segers , Lieven Clement , Davide Risso

We propose the nuclear norm penalty as an alternative to the ridge penalty for regularized multinomial regression. This convex relaxation of reduced-rank multinomial regression has the advantage of leveraging underlying structure among the…

Machine Learning · Statistics 2025-06-09 Scott Powers , Trevor Hastie , Robert Tibshirani

Dimensionality reduction techniques play an essential role in data analytics, signal processing and machine learning. Dimensionality reduction is usually performed in a preprocessing stage that is separate from subsequent data analysis,…

Machine Learning · Computer Science 2016-12-21 Bo Yang , Xiao Fu , Nicholas D. Sidiropoulos

A new methodological framework suitable for era-adjusting baseball statistics is developed in this article. Within this methodological framework specific models are motivated. We call these models Full House Models. Full House Models work…

Applications · Statistics 2024-04-25 Shen Yan , Adrian Burgos , Christopher Kinson , Daniel J. Eck

A discrete-time stochastic process derived from a model of basketball is used to generalize any discrete distribution. The generalized distributions can have one or two more parameters than the parent distribution. Those derived from…

Applications · Statistics 2020-06-25 Rose Baker

Forecast combination and model averaging have become popular tools in forecasting and prediction, both of which combine a set of candidate estimates with certain weights and are often shown to outperform single estimates. A data-driven…

Statistics Theory · Mathematics 2025-10-31 Jiahui Zou , Andrey Vasnev , Wendun Wang , Xinyu Zhang

In this paper, we address the problem of how a network of agents can collaboratively fit a linear model when each agent only ever has an arbitrary summand of the regression data. This problem generalizes previously studied…

Optimization and Control · Mathematics 2014-08-06 François D. Côté , Ioannis N. Psaromiligkos , Warren J. Gross

Approximate Bayesian computation (ABC) and other likelihood-free inference methods have gained popularity in the last decade, as they allow rigorous statistical inference for complex models without analytically tractable likelihood…

Computation · Statistics 2019-06-21 Jukka Sirén , Samuel Kaski

This paper introduces an approach to reference class selection in distributional forecasting with an application to corporate sales growth rates using several co-variates as reference variables, that are implicit predictors. The method can…

Statistical Finance · Quantitative Finance 2024-05-07 Etienne Theising
‹ Prev 1 3 4 5 6 7 10 Next ›