English
Related papers

Related papers: Assessing inter-rater reliability with heterogeneo…

200 papers

In dealing with high-dimensional data sets, factor models are often useful for dimension reduction. The estimation of factor models has been actively studied in various fields. In the first part of this paper, we present a new approach to…

Statistical Finance · Quantitative Finance 2017-11-27 Joongyeub Yeo , George Papanicolaou

Over the past years, topics ranging from climate change to human rights have seen increasing importance for investment decisions. Hence, investors (asset managers and asset owners) who wanted to incorporate these issues started to assess…

Artificial Intelligence · Computer Science 2021-09-22 Tim Krappel , Alex Bogun , Damian Borth

We consider inference in linear regression models that is robust to heteroskedasticity and the presence of many control variables. When the number of control variables increases at the same rate as the sample size the usual…

Statistics Theory · Mathematics 2020-09-29 Koen Jochmans

Backtests on historical data are the basis for practical evaluations of portfolio selection rules, but their reliability is often limited by reliance on a single sample path. This can lead to high estimation variance. Resampling techniques…

Portfolio Management · Quantitative Finance 2025-10-14 Andrew Paskaramoorthy , Terence van Zyl , Tim Gebbie

Multi-objective evaluation is a necessary aspect when managing complex systems, as the intrinsic complexity of a system is generally closely linked to the potential number of optimization objectives. However, an evaluation makes no sense…

Physics and Society · Physics 2016-08-03 Juste Raimbault

Random forest (RF) stands out as a highly favored machine learning approach for classification problems. The effectiveness of RF hinges on two key factors: the accuracy of individual trees and the diversity among them. In this study, we…

Machine Learning · Computer Science 2024-10-28 Ye-eun Kim , Seoung Yun Kim , Hyunjoong Kim

Estimating varying treatment effects in randomized trials with noncompliance is inherently challenging since variation comes from two separate sources: variation in the impact itself and variation in the compliance rate. In this setting,…

Applications · Statistics 2024-08-28 Jared D. Fisher , David W. Puelz , Sameer K. Deshpande

Recent causal inference literature has introduced causal effect decompositions to quantify sources of observed inequalities or disparities in outcomes, but these approaches are typically limited to pairwise comparisons. In healthcare…

Methodology · Statistics 2026-04-27 Lin Yu , Zhihui Liu , Kathy Han , Olli Saarela

Estimating the dependences between random variables, and ranking them accordingly, is a prevalent problem in machine learning. Pursuing frequentist and information-theoretic approaches, we first show that the p-value and the mutual…

Machine Learning · Computer Science 2012-07-02 Harald Steck

Adjusting for covariates is a well established method to estimate the total causal effect of an exposure variable on an outcome of interest. Depending on the causal structure of the mechanism under study there may be different adjustment…

Statistics Theory · Mathematics 2021-04-27 Jack Kuipers , Giusi Moffa

A new empirical Bayes approach to variable selection in the context of generalized linear models is developed. The proposed algorithm scales to situations in which the number of putative explanatory variables is very large, possibly much…

Methodology · Statistics 2021-06-29 Haim Bar , James Booth , Martin T. Wells

Traditional credibility analysis of risks in insurance is based on the random effects model, where the heterogeneity across the policyholders is assumed to be time-invariant. One popular extension is the dynamic random effects (or…

Applications · Statistics 2021-09-17 Jae Youn Ahn , Himchan Jeong , Yang Lu

Individuals often make different decisions when faced with the same context, due to personal preferences and background. For instance, judges may vary in their leniency towards certain drug-related offenses, and doctors may vary in their…

Machine Learning · Computer Science 2021-10-28 Justin Lim , Christina X Ji , Michael Oberst , Saul Blecker , Leora Horwitz , David Sontag

Modeling the time-varying covariance structures of high-dimensional variables is critical across diverse scientific and industrial applications; however, existing approaches exhibit notable limitations in either modeling flexibility or…

Methodology · Statistics 2026-01-21 Taehee Lee , Jun S. Liu

We reconcile the two worlds of dense and sparse modeling by exploiting the positive aspects of both. We employ a factor model and assume {the dynamic of the factors is non-pervasive while} the idiosyncratic term follows a sparse vector…

Methodology · Statistics 2022-05-25 Jonas Krampe , Luca Margaritella

This paper studies model selection consistency for high dimensional sparse regression when data exhibits both cross-sectional and serial dependency. Most commonly-used model selection methods fail to consistently recover the true model when…

Methodology · Statistics 2018-09-12 Jianqing Fan , Yuan Ke , Kaizheng Wang

This paper shows that the degree of approximate multicollinearity in a linear regression model increases simply by including independent variables, even if these are not highly linearly related. In the current situation where it is…

Methodology · Statistics 2025-03-07 Román Salmerón Gómez , Catalina García García

In latent variable models the parameter estimation can be implemented by using the joint or the marginal likelihood, based on independence or conditional independence assumptions. The same dilemma occurs within the Bayesian framework with…

Computation · Statistics 2014-09-18 Silia Vitoratou , Ioannis Ntzoufras , Irini Moustaki

Statistical learning evolves quickly with more and more sophisticated models proposed to incorporate the complicated data structure from modern scientific and business problems. Varying index coefficient models extend varying coefficient…

Statistics Theory · Mathematics 2019-03-05 Li Jialiang , Lv Jing

Modern regression applications can involve hundreds or thousands of variables which motivates the use of variable selection methods. Bayesian variable selection defines a posterior distribution on the possible subsets of the variables…

Methodology · Statistics 2024-10-16 J. E. Griffin