中文
相关论文

相关论文: Inference in generalized linear models with robust…

200 篇论文

In modern scientific applications, large volumes of covariate data are readily available, while outcome labels are costly, sparse, and often subject to distribution shift. This asymmetry has spurred interest in semi-supervised (SS)…

统计理论 · 数学 2026-05-12 Lorenzo Testa , Qi Xu , Jing Lei , Kathryn Roeder

The issue addressed in this paper is that of testing for common breaks across or within equations of a multivariate system. Our framework is very general and allows integrated regressors and trends as well as stationary regressors. The null…

统计理论 · 数学 2018-01-12 Tatsushi Oka , Pierre Perron

Qualifying gene and isoform expression is one of the primary tasks for RNA-Seq experiments. Given a sequence of counts representing numbers of reads mapped to different positions (exons and junctions) of isoforms, methods based on Poisson…

应用统计 · 统计学 2014-10-27 Jun Li , Hui Jiang

In this paper we study the problem of statistical inference on the parameters of the semiparametric variance-mean mixtures. This class of mixtures has recently become rather popular in statistical and financial modelling. We design a…

其他统计学 · 统计学 2017-05-23 Denis Belomestny , Vladimir Panov

Assessing model generalization under distribution shift is essential for real-world deployment, particularly when labeled test data is unavailable. This paper presents a unified and practical framework for unsupervised model evaluation and…

机器学习 · 计算机科学 2025-10-06 Weijian Deng , Weijie Tu , Ibrahim Radwan , Mohammad Abu Alsheikh , Stephen Gould , Liang Zheng

We consider high-dimensional generalized linear models when the covariates are contaminated by measurement error. Estimates from errors-in-variables regression models are well-known to be biased in traditional low-dimensional settings if…

统计计算 · 统计学 2020-01-06 Michael Byrd , Monnie McGee

We consider a family of problems that are concerned about making predictions for the majority of unlabeled, graph-structured data samples based on a small proportion of labeled samples. Relational information among the data samples, often…

机器学习 · 计算机科学 2019-11-05 Jiaqi Ma , Weijing Tang , Ji Zhu , Qiaozhu Mei

Quantitative studies in many fields involve the analysis of multivariate data of diverse types, including measurements that we may consider binary, ordinal and continuous. One approach to the analysis of such mixed data is to use a copula…

统计理论 · 数学 2007-06-13 Peter D. Hoff

In semi-supervised learning, the prevailing understanding suggests that observing additional unlabeled samples improves estimation accuracy for linear parameters only in the case of model misspecification. In this work, we challenge such a…

统计方法学 · 统计学 2025-09-03 Kai Chen , Yuqian Zhang

We introduce inference methods for score decompositions, which partition scoring functions for predictive assessment into three interpretable components: miscalibration, discrimination, and uncertainty. Our estimation and inference relies…

计量经济学 · 经济学 2026-03-05 Timo Dimitriadis , Marius Puke

Recently-developed genotype imputation methods are a powerful tool for detecting untyped genetic variants that affect disease susceptibility in genetic association studies. However, existing imputation methods require individual-level…

应用统计 · 统计学 2010-11-15 Xiaoquan Wen , Matthew Stephens

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

统计方法学 · 统计学 2024-09-05 F. Richard Guo , Rajen D. Shah

We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of…

统计方法学 · 统计学 2018-08-15 Anru Zhang , Lawrence D. Brown , T. Tony Cai

A general random effects model is proposed that allows for continuous as well as discrete distributions of the responses. Responses can be unrestricted continuous, bounded continuous, binary, ordered categorical or given in the form of…

统计方法学 · 统计学 2024-04-30 Gerhard Tutz

We study a general factor analysis framework where the $n$-by-$p$ data matrix is assumed to follow a general exponential family distribution entry-wise. While this model framework has been proposed before, we here further relax its…

统计方法学 · 统计学 2025-12-02 Liang Wang , Luis Carvalho

We consider semiparametric location-scatter models for which the $p$-variate observation is obtained as $X=\Lambda Z+\mu$, where $\mu$ is a $p$-vector, $\Lambda$ is a full-rank $p\times p$ matrix and the (unobserved) random $p$-vector $Z$…

统计理论 · 数学 2012-02-24 Pauliina Ilmonen , Davy Paindaveine

For many applications, an ensemble of base classifiers is an effective solution. The tuning of its parameters(number of classes, amount of data on which each classifier is to be trained on, etc.) requires G, the generalization error of a…

We consider a class of semiparametric regression models which are one-parameter extensions of the Cox [J. Roy. Statist. Soc. Ser. B 34 (1972) 187-220] model for right-censored univariate failure times. These models assume that the hazard…

统计理论 · 数学 2007-06-13 Michael R. Kosorok , Bee Leng Lee , Jason P. Fine

Data sets obtained from linking multiple files are frequently affected by mismatch error, as a result of non-unique or noisy identifiers used during record linkage. Accounting for such mismatch error in downstream analysis performed on the…

统计方法学 · 统计学 2023-06-02 Martin Slawski , Brady T. West , Priyanjali Bukke , Guoqing Diao , Zhenbang Wang , Emanuel Ben-David

With the evolution of single-cell RNA sequencing techniques into a standard approach in genomics, it has become possible to conduct cohort-level causal inferences based on single-cell-level measurements. However, the individual gene…

统计方法学 · 统计学 2025-04-23 Jin-Hong Du , Zhenghao Zeng , Edward H. Kennedy , Larry Wasserman , Kathryn Roeder