English
Related papers

Related papers: Semiparametric count data regression for self-repo…

200 papers

Continuous-time multi-state survival models can be used to describe health-related processes over time. In the presence of interval-censored times for transitions between the living states, the likelihood is constructed using transition…

Methodology · Statistics 2017-03-24 Robson J. M. Machado , Ardo van den Hout

Event-related potentials (ERPs) extracted from electroencephalography (EEG) data in response to stimuli are widely used in psychological and neuroscience experiments. A major goal is to link ERP characteristic components to subject-level…

Methodology · Statistics 2024-06-11 Cheng-Han Yu , Meng Li , Marina Vannucci

Score-based generative models (SGMs) have gained prominence in sparse-view CT reconstruction for their precise sampling of complex distributions. In SGM-based reconstruction, data consistency in the score-based diffusion model ensures close…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Weiwen Wu , Yanyang Wang

The cumulative incidence is the probability of failure from the cause of interest over a certain time period in the presence of other risks. A semiparametric regression model proposed by Fine and Gray (1999) has become the method of choice…

Methodology · Statistics 2016-03-02 Lu Mao , D. Y. Lin

A state-space model is a statistical framework for inferring latent states from observed time-series data. However, inference with nonlinear and high-dimensional state-space models remains challenging. To this end, an approach based on…

Computational Engineering, Finance, and Science · Computer Science 2026-04-02 Yuma Yamaoka , Seiichi Uchida , Shoji Toyota

The increasing integration of distributed energy resources (DERs) is transforming power systems into complex, decentralized networks, particularly at the distribution level, where active distribution networks (ADNs) introduce new challenges…

Optimization and Control · Mathematics 2025-07-14 J. G. De la Varga , J. M. Morales , S. Pineda

We propose a novel framework for incorporating unlabeled data into semi-supervised classification problems, where scenarios involving the minimization of either i) adversarially robust or ii) non-robust loss functions have been considered.…

Subsampling is a general statistical method developed in the 1990s aimed at estimating the sampling distribution of a statistic $\hat \theta _n$ in order to conduct nonparametric inference such as the construction of confidence intervals…

Statistics Theory · Mathematics 2021-12-14 Dimitris N. Politis

Unsupervised learning on imbalanced data is challenging because, when given imbalanced data, current model is often dominated by the major category and ignores the categories with small amount of data. We develop a latent variable model…

Machine Learning · Computer Science 2016-07-04 Fariba Yousefi , Zhenwen Dai , Carl Henrik Ek , Neil Lawrence

When the response mechanism is believed to be not missing at random (NMAR), a valid analysis requires stronger assumptions on the response mechanism than standard statistical methods would otherwise require. Semiparametric estimators have…

Methodology · Statistics 2020-05-08 Kosuke Morikawa , Jae Kwang Kim

Researchers are often interested in understanding the relationship between a set of covariates and a set of response variables. To achieve this goal, the use of regression analysis, either linear or generalized linear models, is largely…

We consider the problem of estimating the average treatment effect (ATE) in a semi-supervised learning setting, where a very small proportion of the entire set of observations are labeled with the true outcome but features predictive of the…

Methodology · Statistics 2020-10-27 David Cheng , Ashwin Ananthakrishnan , Tianxi Cai

We consider generalized linear regression analysis with left-censored covariate due to the lower limit of detection. Complete case analysis by eliminating observations with values below limit of detection yields valid estimates for…

Methodology · Statistics 2014-12-09 Shengchun Kong , Bin Nan

Bayesian Additive Regression Trees (BART) is a flexible machine learning algorithm capable of capturing nonlinearities between an outcome and covariates and interaction among covariates. We extend BART to a semiparametric regression…

Applications · Statistics 2018-06-13 Bret Zeldow , Vincent Lo Re , Jason Roy

Gaussian processes (GPs) are a powerful tool for probabilistic inference over functions. They have been applied to both regression and non-linear dimensionality reduction, and offer desirable properties such as uncertainty estimates,…

Machine Learning · Statistics 2014-10-01 Yarin Gal , Mark van der Wilk , Carl E. Rasmussen

Clinical dietary assessment can generate detailed but high-dimensional nutrient and food-group information that is difficult to translate quickly into counselling priorities. This paper proposes an explainable unsupervised-to-supervised…

Quantitative Methods · Quantitative Biology 2026-05-12 Wing Yi Yu , Chun Yin Chiu

We study semiparametric factor models in high-dimensional panels where the factor loadings consist of a nonparametric component explained by observed covariates and an idiosyncratic component capturing unobserved heterogeneity. A key…

Methodology · Statistics 2025-12-09 Sijie Zheng

Assume that we observe a large number of curves, all of them with identical, although unknown, shape, but with a different random shift. The objective is to estimate the individual time shifts and their distribution. Such an objective…

Applications · Statistics 2015-03-13 T. Trigano , U. Isserles , Y. Ritov

Count data modeling has been extensively applied in medical sciences to analyze various healthcare datasets. Numerous probability models have been developed to address diverse aspects of healthcare data. In this study, we propose a novel…

Methodology · Statistics 2025-09-04 Peer Bilal Ahmad , Na Elah

This paper investigates the problem of making inference about a parametric model for the regression of an outcome variable $Y$ on covariates $(V,L)$ when data are fused from two separate sources, one which contains information only on $(V,…

Methodology · Statistics 2020-12-15 Katherine Evans , BaoLuo Sun , James Robins , Eric J. Tchetgen Tchetgen