English
Related papers

Related papers: Regression analysis of doubly truncated data

200 papers

The spatial scan statistic is widely used to detect disease clusters in epidemiological surveillance. Since the seminal work by~\cite{kulldorff1997}, numerous extensions have emerged, including methods for defining scan regions, detecting…

Methodology · Statistics 2025-02-11 Takayuki Kawashima , Daisuke Yoneoka , Yuta Tanoue , Akifumi Eguchi , Shuhei Nomura

High-dimensional mixed data as a combination of both continuous and ordinal variables are widely seen in many research areas such as genomic studies and survey data analysis. Estimating the underlying correlation among mixed data is hence…

Methodology · Statistics 2018-09-18 Xiaoyun Quan , James G. Booth , Martin T. Wells

Optimization problems with an auxiliary latent variable structure in addition to the main model parameters occur frequently in computer vision and machine learning. The additional latent variables make the underlying optimization task…

Machine Learning · Computer Science 2020-03-13 Christopher Zach , Huu Le

Truncation is a statistical phenomenon that occurs in many time to event studies. For example, autopsy-confirmed studies of neurodegenerative diseases are subject to an inherent left and right truncation, also known as double truncation.…

Methodology · Statistics 2018-03-28 Lior Rennert , Sharon X. Xie

All data are digitized, and hence are essentially integers rather than true real numbers. Ordinarily this causes no difficulties since the truncation or rounding usually occurs below the noise level. However, in some instances, when the…

Data Analysis, Statistics and Probability · Physics 2016-02-16 Kevin H. Knuth , J. Patrick Castle , Kevin R. Wheeler

We compare two theoretically distinct approaches to generating artificial (or ``surrogate'') data for testing hypotheses about a given data set. The first and more straightforward approach is to fit a single ``best'' model to the original…

comp-gas · Physics 2015-06-24 James Theiler , Dean Prichard

Astronomers often deal with data where the covariates and the dependent variable are measured with heteroscedastic non-Gaussian error. For instance, while TESS and Kepler datasets provide a wealth of information, addressing the challenges…

Instrumentation and Methods for Astrophysics · Physics 2024-12-17 Naomi Giertych , Jonathan P Williams , Sujit Ghosh

This review outlines concepts of mathematical statistics, elements of probability theory, hypothesis tests and point estimation for use in the analysis of modern astronomical data. Least squares, maximum likelihood, and Bayesian approaches…

Instrumentation and Methods for Astrophysics · Physics 2012-05-10 Eric D. Feigelson , G. Jogesh Babu

In two-way contingency tables under an asymmetric situation, where the row and column variables are defined as explanatory and response variables, respectively, quantifying the extent to which the explanatory variable contributes to…

Methodology · Statistics 2026-04-17 Wataru Urasaki , Kouji Tahata , Sadao Tomizawa

Two-sample inference for the difference of population means typically relies upon a Central Limit Theorem approximation. When data are drawn from a Negative Binomial distribution, previous work of Shilane et al. (2010) showed that a Normal…

Methodology · Statistics 2012-03-06 David Shilane , Derek Bean

We propose robust sparse reduced rank regression for analyzing large and complex high-dimensional data with heavy-tailed random noise. The proposed method is based on a convex relaxation of a rank- and sparsity-constrained non-convex…

Machine Learning · Statistics 2019-04-16 Kean Ming Tan , Qiang Sun , Daniela Witten

Robustness and counterfactual bias are usually evaluated on a test dataset. However, are these evaluations robust? If the test dataset is perturbed slightly, will the evaluation results keep the same? In this paper, we propose a "double…

Computation and Language · Computer Science 2021-04-13 Chong Zhang , Jieyu Zhao , Huan Zhang , Kai-Wei Chang , Cho-Jui Hsieh

An effective two-stage method for an estimation of parameters of the linear regression is considered. For this purpose we introduce a certain quasi-estimator that, in contrast to usual estimator, produces two alternative estimates. It is…

Statistics Theory · Mathematics 2010-10-06 Anatoly Gordinsky

This review article considers some of the most common methods used in astronomy for regressing one quantity against another in order to estimate the model parameters or to predict an observationally expensive quantity using trends between…

Instrumentation and Methods for Astrophysics · Physics 2012-10-24 S. Andreon , M. A. Hurn

In regression analysis of multivariate data, it is tacitly assumed that response and predictor variables in each observed response-predictor pair correspond to the same entity or unit. In this paper, we consider the situation of "permuted…

Statistics Theory · Mathematics 2017-11-17 Martin Slawski , Emanuel Ben-David

A number of models for generating statistical data in various fields of insurance, including life insurance, pensions, and general insurance have been considered. It is shown that the insurance statistics data, as a rule, are truncated and…

Methodology · Statistics 2019-04-16 Valery Baskakov , Anna Bartunova

In this paper we present an enhancement of the regression-based variance reduction approaches recently proposed in Belomestny et al. This enhancement is based on a truncation of the control variate and allows for a significant reduction of…

Probability · Mathematics 2017-11-10 Denis Belomestny , Stefan Häfner , Mikhail Urusov

The Mann-Whitney-Wilcoxon rank sum test (MWWRST) is a widely used method for comparing two treatment groups in randomized control trials, particularly when dealing with highly skewed data. However, when applied to observational study data,…

High-dimensional penalized rank regression is a powerful tool for modeling high-dimensional data due to its robustness and estimation efficiency. However, the non-smoothness of the rank loss brings great challenges to the computation. To…

Methodology · Statistics 2025-02-20 Leheng Cai , Xu Guo , Heng Lian , Liping Zhu

Boosting has garnered significant interest across both machine learning and statistical communities. Traditional boosting algorithms, designed for fully observed random samples, often struggle with real-world problems, particularly with…

Machine Learning · Statistics 2026-02-19 Yuan Bian , Grace Y. Yi , Wenqing He