English
Related papers

Related papers: Understanding the effects of dichotomization of co…

200 papers

The focus of modern biomedical studies has gradually shifted to explanation and estimation of joint effects of high dimensional predictors on disease risks. Quantifying uncertainty in these estimates may provide valuable insight into…

Methodology · Statistics 2021-03-09 Zhe Fei , Yi Li

A common approach in forecasting problems is to estimate a least-squares regression (or other statistical learning models) from past data, which is then applied to predict future outcomes. An underlying assumption is that the same…

Methodology · Statistics 2022-03-22 Malte Schierholz

Many scientific applications involve mixed spatially indexed outcomes of heterogeneous types that are driven by shared latent mechanisms. Modeling such data is challenging due to complex, nonlinear, and potentially nonstationary spatial…

Methodology · Statistics 2026-03-10 Yeseul Jeon , Kyeong Eun Lee , Joon Jin Song

Observational studies of causal effects require adjustment for confounding factors. In the tabular setting, where these factors are well-defined, separate random variables, the effect of confounding is well understood. However, in public…

Machine Learning · Computer Science 2023-02-16 Connor T. Jerzak , Fredrik Johansson , Adel Daoud

Cosmological large-scale structure analyses based on two-point correlation functions often assume a Gaussian likelihood function with a fixed covariance matrix. We study the impact on cosmological parameter estimation of ignoring the…

Cosmology and Nongalactic Astrophysics · Physics 2019-03-21 Darsh Kodwani , David Alonso , Pedro Ferreira

The prediction of phenotypic traits using high-density genomic data has many applications such as the selection of plants and animals of commercial interest; and it is expected to play an increasing role in medical diagnostics. Statistical…

Methodology · Statistics 2016-09-29 Marco Scutari , Ian Mackay , David Balding

Heteroskedasticity is a statistical anomaly that describes differing variances of error terms in a time series dataset. The presence of heteroskedasticity in data imposes serious challenges for forecasting models and many statistical tests…

Statistics Theory · Mathematics 2016-09-21 Marwa Hassan , Mo Hossny , Douglas Creighton , Saeid Nahavandi

We examine the errors on counts in cells extracted from galaxy surveys. The measurement error, related to the finite number of sampling cells, is disentangled from the ``cosmic error'', due to the finiteness of the survey. Using the…

Astrophysics · Physics 2009-10-28 István Szapudi , Stéphane Colombi

Difference-in-differences (DID) is a widely used quasi-experimental design for causal inference, traditionally applied to scalar or Euclidean outcomes, while extensions to outcomes residing in non-Euclidean spaces remain limited. Existing…

Methodology · Statistics 2025-01-30 Yidong Zhou , Daisuke Kurisu , Taisuke Otsu , Hans-Georg Müller

Propensity score (PS) matching to estimate causal effects of exposure is biased when unmeasured spatial confounding exists. Some exposures are continuous yet dependent on a binary variable (e.g., level of a contaminant (continuous) within a…

Methodology · Statistics 2026-05-04 Honghyok Kim , Michelle Bell

We develop Bayesian nonparametric models for spatially indexed data of mixed type. Our work is motivated by challenges that occur in environmental epidemiology, where the usual presence of several confounding variables that exhibit complex…

Methodology · Statistics 2014-10-17 Georgios Papageorgiou , Sylvia Richardson , Nicky Best

The problem of dynamic prediction with time-dependent covariates, given by biomarkers, repeatedly measured over time, has received much attention over the last decades. Two contrasting approaches have become in widespread use. The first is…

Methodology · Statistics 2021-03-31 Hein Putter , Hans C. van Houwelingen

Model-based geostatistical design involves the selection of locations to collect data to minimise an expected loss function over a set of all possible locations. The loss function is specified to reflect the aim of data collection, which,…

Computation · Statistics 2021-12-03 S. G. Jagath Senarathne , Werner G. Müller , James M. McGree

Air pollution remains a major environmental risk factor that is often associated with adverse health outcomes. However, quantifying and evaluating its effects on human health is challenging due to the complex nature of exposure data. Recent…

Methodology · Statistics 2025-06-02 Soumyakanti Pan , Sudipto Banerjee

In statistical genetics an important task involves building predictive models for the genotype-phenotype relationships and thus attribute a proportion of the total phenotypic variance to the variation in genotypes. Numerous models have been…

Applications · Statistics 2016-03-30 Deniz Akdemir , Jean-Luc Jannink

Intensive longitudinal biomarker data are increasingly common in scientific studies that seek temporally granular understanding of the role of behavioral and physiological factors in relation to outcomes of interest. Intensive longitudinal…

Methodology · Statistics 2024-01-17 Mingyan Yu , Zhenke Wu , Margaret Hicken , Michael R. Elliott

In this paper we study the asymptotics of linear regression in settings with non-Gaussian covariates where the covariates exhibit a linear dependency structure, departing from the standard assumption of independence. We model the covariates…

Machine Learning · Statistics 2024-12-10 Behrad Moniri , Hamed Hassani

This paper studies algorithmic fairness when the protected attribute is location. To handle protected attributes that are continuous, such as age or income, the standard approach is to discretize the domain into predefined groups, and…

Machine Learning · Computer Science 2023-02-27 Dimitris Sacharidis , Giorgos Giannopoulos , George Papastefanatos , Kostas Stefanidis

We consider continuous-time models with a large panel of moment conditions, where the structural parameter depends on a set of characteristics, whose effects are of interest. The leading example is the linear factor model in financial…

Econometrics · Economics 2018-12-04 Yuan Liao , Xiye Yang

Preferential sampling is a common feature in geostatistics and occurs when the locations to be sampled are chosen based on information about the phenomena under study. In this case, point pattern models are commonly used as the probability…

Methodology · Statistics 2022-10-27 Douglas Mateus da Silva , Dani Gamerman