English
Related papers

Related papers: Semiparametric count data regression for self-repo…

200 papers

We propose a simple yet powerful framework for modeling integer-valued data, such as counts, scores, and rounded data. The data-generating process is defined by Simultaneously Transforming and Rounding (STAR) a continuous-valued process,…

Methodology · Statistics 2019-09-04 Daniel R. Kowal , Antonio Canale

Clinical and epidemiological studies encode participant information in multivariate vectors with mixed type variables on continuous, truncated, ordinal, and binary scales. Semiparametric Gaussian Copula (SGC) assumes that observed data is…

Methodology · Statistics 2026-03-19 Debangan Dey , Vadim Zipunnikov

Amid growing global mental health concerns, particularly among vulnerable groups, natural language processing offers a tremendous potential for early detection and intervention of people's mental disorders via analyzing their postings and…

Machine Learning · Computer Science 2023-11-10 Haijian Shao , Ming Zhu , Shengjie Zhai

We investigate the parameter estimation of regression models with fixed group effects, when the group variable is missing while group related variables are available. This problem involves clustering to infer the missing group variable…

Methodology · Statistics 2020-12-29 Matthieu Marbac , Mohammed Sedki , Christophe Biernacki , Vincent Vandewalle

Electronic health record (EHR) data has emerged as a valuable resource for analyzing patient health status. However, the prevalence of missing data in EHR poses significant challenges to existing methods, leading to spurious correlations…

Machine Learning · Computer Science 2024-05-16 Zhihao Yu , Xu Chu , Yujie Jin , Yasha Wang , Junfeng Zhao

Big spatio-temporal datasets, available through both open and administrative data sources, offer significant potential for social science research. The magnitude of the data allows for increased resolution and analysis at individual level.…

Applications · Statistics 2017-11-27 Anastasia Ushakova , Slava J. Mikhaylov

Data harmonization is the process by which an equivalence is developed between two variables measuring a common trait. Our problem is motivated by dementia research in which multiple tests are used in practice to measure the same underlying…

Methodology · Statistics 2021-10-13 Steven Wilkins-Reeves , Yen-Chi Chen , Kwun Chuen Gary Chan

Tensors are becoming prevalent in modern applications such as medical imaging and digital marketing. In this paper, we propose a sparse tensor additive regression (STAR) that models a scalar response as a flexible nonparametric function of…

Machine Learning · Statistics 2021-03-08 Botao Hao , Boxiang Wang , Pengyuan Wang , Jingfei Zhang , Jian Yang , Will Wei Sun

The premise of independence among subjects in the same cluster/group often fails in practice, and models that rely on such untenable assumption can produce misleading results. To overcome this severe deficiency, we introduce a new…

Methodology · Statistics 2022-02-22 Jussiane Nader Gonçalves , Wagner Barreto-Souza , Hernando Ombao

It is possible to approach regression analysis with random covariates from a semiparametric perspective where information is combined from multiple multivariate sources. The approach assumes a semiparametric density ratio model where…

Methodology · Statistics 2012-10-02 Anastasia Voulgaraki , Benjamin Kedem , Barry I. Graubard

Interval-censoring frequently occurs in studies of chronic diseases where disease status is inferred from intermittently collected biomarkers. Although many methods have been developed to analyze such data, they typically assume perfect…

Methodology · Statistics 2026-05-26 Yuhao Deng , Donglin Zeng , Yuanjia Wang

The onset of several silent, chronic diseases such as diabetes can be detected only through diagnostic tests. Due to cost considerations, self-reported outcomes are routinely collected in lieu of expensive diagnostic tests in large-scale…

Applications · Statistics 2015-09-15 Xiangdong Gu , Yunsheng Ma , Raji Balasubramanian

In multi-center clinical trials, due to various reasons, the individual-level data are strictly restricted to be assessed publicly. Instead, the summarized information is widely available from published results. With the advance of…

Methodology · Statistics 2021-01-05 Jing Qin , Yukun Liu , Pengfei Li

Two key challenges in modern statistical applications are the large amount of information recorded per individual, and that such data are often not collected all at once but in batches. These batch effects can be complex, causing…

Applications · Statistics 2019-05-21 Alejandra Avalos-Pacheco , David Rossell , Richard S. Savage

Mobile health studies often collect multiple within-day self-reported assessments of participants' behavior and well-being on different scales such as physical activity (continuous), pain levels (truncated), mood states (ordinal), and life…

Methodology · Statistics 2023-09-22 Debangan Dey , Rahul Ghosal , Kathleen Merikangas , Vadim Zipunnikov

With the rapid advances of data acquisition techniques, spatio-temporal data are becoming increasingly abundant in a diverse array of disciplines. Here we develop spatio-temporal regression methodology for analyzing large amounts of…

Methodology · Statistics 2021-12-01 Ting Fung Ma , Fangfang Wang , Jun Zhu , Anthony R. Ives , Katarzyna E. Lewińska

Regarding the rising number of people suffering from mental health illnesses in today's society, the importance of mental health cannot be overstated. Wearable sensors, which are increasingly widely available, provide a potential way to…

Machine Learning · Computer Science 2023-10-16 Anket Patil , Dhairya Shah , Abhishek Shah , Mokshit Gala

Conducting valid statistical analyses is challenging in the presence of missing-not-at-random (MNAR) data, where the missingness mechanism is dependent on the missing values themselves even conditioned on the observed data. Here, we…

Methodology · Statistics 2023-06-13 Anna Guo , Jiwei Zhao , Razieh Nabi

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

Methodology · Statistics 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

The widespread adoption of social media has heightened interest in its psychological effects, particularly on mental health indicators such as anxiety, depression, loneliness, and sleep quality, as these platforms increasingly influence…

Machine Learning · Computer Science 2026-04-28 Md All Shahria , Sanjeda Dewan Mithila , Touhid Alam , Mohammad Sakib Mahmood , Mahfuza Khatun
‹ Prev 1 2 3 10 Next ›