Related papers: Statistical Analysis of Galaxy Surveys - I. Robust…
Data clustering reduces the effective sample size from the number of observations towards the number of clusters. For instrumental variable models this reduced effective sample size makes the instruments more likely to be weak, in the sense…
The error or variability of machine learning algorithms is often assessed by repeatedly re-fitting a model with different weighted versions of the observed data. The ubiquitous tools of cross-validation (CV) and the bootstrap are examples…
Propensity score (PS) methods are widely used to estimate treatment effects in non-randomized studies. Variance is typically estimated using sandwich or bootstrap methods, which can either treat the PS as estimated or fixed. The latter is…
Statistical resampling methods have become feasible for parametric estimation, hypothesis testing, and model validation now that the computer is a ubiquitous tool for statisticians. This essay focuses on the resampling technique for…
Measuring the angular clustering of galaxies as a function of redshift is a powerful method for extracting information from the three-dimensional galaxy distribution. The precision of such measurements will dramatically increase with…
The combination of two- and three-point clustering statistics of galaxies and the underlying matter distribution has the potential to break degeneracies between cosmological parameters and nuisance parameters and can lead to significantly…
We present a novel method to compress galaxy clustering three-point statistics and apply it to redshift space galaxy bispectrum monopole measurements from BOSS DR12 CMASS data considering a $k$-space range of $0.03-0.12\,h/\mathrm{Mpc}$.…
We measure the monopole moment of the three-point correlation function on scales $1\mpc-70\mpc$ in the Two degree Field Galaxy Redshift Survey (2dFGRS). Volume limited samples are constructed using a series of integral magnitudes bins…
Sequential trial emulation (STE) is an approach to estimating causal treatment effects by emulating a sequence of target trials from observational data. In STE, inverse probability weighting is commonly utilised to address time-varying…
Conventional cluster-robust inference can be invalid when data contain clusters of unignorably large size. We formalize this issue by deriving a necessary and sufficient condition for its validity, and show that this condition is frequently…
In cluster-randomized trials, generalized linear mixed models and generalized estimating equations have conventionally been the default analytic methods for estimating the average treatment effect as routine practice. However, recent…
Measurements of clustering in large-scale imaging surveys that make use of photometric redshifts depend on the uncertainties in the redshift determination. We have used light-cone simulations to show how the deprojection method successfully…
The caustic technique uses galaxy redshifts alone to measure the escape velocity and mass profiles of galaxy clusters to clustrocentric distances well beyond the virial radius, where dynamical equilibrium does not necessarily hold. We…
Introductory texts on statistics typically only cover the classical "two sigma" confidence interval for the mean value and do not describe methods to obtain confidence intervals for other estimators. The present technical report fills this…
In this paper we propose a flexible nested error regression small area model with high dimensional parameter that incorporates heterogeneity in regression coefficients and variance components. We develop a new robust small area specific…
The manuscript discusses how to incorporate random effects for quantile regression models for clustered data with focus on settings with many but small clusters. The paper has three contributions: (i) documenting that existing methods may…
Nested-error regression models are widely used for analyzing clustered data. For example, they are often applied to two-stage sample surveys, and in biology and econometrics. Prediction is usually the main goal of such analyses, and…
We present a mitigation strategy to reduce the impact of non-linear galaxy bias on the joint `$3 \times 2 $pt' cosmological analysis of weak lensing and galaxy surveys. The $\Psi$-statistics that we adopt are based on Complete Orthogonal…
The existence of galaxy intrinsic clustering severely hampers the weak lensing reconstruction from cosmic magnification. In paper I \citep{Yang2011}, we proposed a minimal variance estimator to overcome this problem. By utilizing the…
We critically investigate current statistical tests applied to high redshift clusters of galaxies in order to test the standard cosmological model and describe their range of validity. We carefully compare a sample of high-redshift,…