English
Related papers

Related papers: Inference for Large Panel Data with Many Covariate…

200 papers

Corrupted data sets containing noisy or missing observations are prevalent in various contemporary applications such as economics, finance and bioinformatics. Despite the recent methodological and algorithmic advances in high-dimensional…

Methodology · Statistics 2020-05-12 J. Wu , Z. Zheng , Y. Li , Y. Zhang

We consider the problem of large-scale inference on the row or column variables of data in the form of a matrix. Often this data is transposable, meaning that both the row variables and column variables are of potential interest. An example…

Methodology · Statistics 2015-03-13 Genevera I. Allen , Robert Tibshirani

The use of weights provides an effective strategy to incorporate prior domain knowledge in large-scale inference. This paper studies weighted multiple testing in a decision-theoretic framework. We develop oracle and data-driven procedures…

Methodology · Statistics 2017-05-10 Pallavi Basu , T. Tony Cai , Kiranmoy Das , Wenguang Sun

Important objectives in cancer research are the prediction of a patient's risk based on molecular measurements such as gene expression data and the identification of new prognostic biomarkers (e.g. genes). In clinical practice, this is…

Applications · Statistics 2020-04-17 Katrin Madjar , Manuela Zucknick , Katja Ickstadt , Jörg Rahnenführer

Identifying co-varying causal elements in very high dimensional feature space with internal structures, e.g., a space with as many as millions of linearly ordered features, as one typically encounters in problems such as whole genome…

Methodology · Statistics 2012-06-18 Seyoung Kim , Eric P. Xing

Unmeasured confounding presents a significant challenge in causal inference from observational studies. Classical approaches often rely on collecting proxy variables, such as instrumental variables. However, in applications where the…

Methodology · Statistics 2025-01-16 Xiaochuan Shi , Dehan Kong , Linbo Wang

Instrumental variable methods provide useful tools for inferring causal effects in the presence of unmeasured confounding. To apply these methods with large-scale data sets, a major challenge is to find valid instruments from a possibly…

Methodology · Statistics 2024-09-24 Xinyi Zhang , Linbo Wang , Stanislav Volgushev , Dehan Kong

In this paper, we propose a new approach to causal inference with panel data. Instead of using panel data to adjust for differences in the distribution of unobserved heterogeneity between the treated and comparison groups, we instead use…

Econometrics · Economics 2025-12-01 Brantly Callaway , Derek Dyal , Pedro H. C. Sant'Anna , Emmanuel S. Tsyawo

Big data applications, such as medical imaging and genetics, typically generate datasets that consist of few observations n on many more variables p, a scenario that we denote as p>>n. Traditional data processing methods are often…

Data Analysis, Statistics and Probability · Physics 2016-05-18 Magnus O. Ulfarsson , Frosti Palsson , Jakob Sigurdsson , Johannes R. Sveinsson

Fitting high-dimensional data involves a delicate tradeoff between faithful representation and the use of sparse models. Too often, sparsity assumptions on the fitted model are too restrictive to provide a faithful representation of the…

Machine Learning · Statistics 2013-12-17 Majid Janzamin , Animashree Anandkumar

This paper presents a selective survey of recent developments in statistical inference and multiple testing for high-dimensional regression models, including linear and logistic regression. We examine the construction of confidence…

Methodology · Statistics 2023-01-26 T. Tony Cai , Zijian Guo , Yin Xia

We propose a novel distributional regression model for a multivariate response vector based on a copula process over the covariate space. It uses the implicit copula of a Gaussian multivariate regression, which we call a ``regression…

Methodology · Statistics 2024-03-06 Nadja Klein , Michael Stanley Smith , David Nott , Ryan Chisholm

While reliable data-driven decision-making hinges on high-quality labeled data, the acquisition of quality labels often involves laborious human annotations or slow and expensive scientific measurements. Machine learning is becoming an…

Machine Learning · Statistics 2024-03-01 Tijana Zrnic , Emmanuel J. Candès

Covariance estimation for high-dimensional datasets is a fundamental problem in modern day statistics with numerous applications. In these high dimensional datasets, the number of variables p is typically larger than the sample size n. A…

Methodology · Statistics 2016-10-11 Kshitij Khare , Sang Oh , Syed Rahman , Bala Rajaratnam

Sparse Gaussian processes and various extensions thereof are enabled through inducing points, that simultaneously bottleneck the predictive capacity and act as the main contributor towards model complexity. However, the number of inducing…

Machine Learning · Computer Science 2021-07-27 Anders Kirk Uhrenholt , Valentin Charvet , Bjørn Sand Jensen

The estimation of causal treatment effects from observational data is a fundamental problem in causal inference. To avoid bias, the effect estimator must control for all confounders. Hence practitioners often collect data for as many…

Machine Learning · Statistics 2020-11-05 Kristjan Greenewald , Dmitriy Katz-Rogozhnikov , Karthik Shanmugam

The paper proposes a new covariance estimator for large covariance matrices when the variables have a natural ordering. Using the Cholesky decomposition of the inverse, we impose a banded structure on the Cholesky factor, and select the…

Applications · Statistics 2008-12-18 Elizaveta Levina , Adam Rothman , Ji Zhu

In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in…

Methodology · Statistics 2019-07-09 Yinchu Zhu , Jelena Bradic

This paper studies model selection consistency for high dimensional sparse regression when data exhibits both cross-sectional and serial dependency. Most commonly-used model selection methods fail to consistently recover the true model when…

Methodology · Statistics 2018-09-12 Jianqing Fan , Yuan Ke , Kaizheng Wang

Motivation: The high dimensionality of genomic data calls for the development of specific classification methodologies, especially to prevent over-optimistic predictions. This challenge can be tackled by compression and variable selection,…

Methodology · Statistics 2021-04-10 G. Durif , L. Modolo , J. Michaelsson , J. E. Mold , S. Lambert-Lacroix , F. Picard
‹ Prev 1 3 4 5 6 7 10 Next ›