English
Related papers

Related papers: Asymmetric Tobit analysis for correlation estimati…

200 papers

Importance sampling algorithms are discussed in detail, with an emphasis on implicit sampling, and applied to data assimilation via particle filters. Implicit sampling makes it possible to use the data to find high-probability samples at…

Computation · Statistics 2015-06-02 Alexandre J. Chorin , Fei Lu , Robert N. Miller , Matthias Morzfeld , Xuemin Tu

Multi-category data arise in diverse fields including marketing, chemistry, public policy, genomics, political science, and ecology. We consider the problem of estimating ratios of category-specific means in a fully nonparametric setting,…

Methodology · Statistics 2025-10-29 Grant Hopkins , Sarah Teichman , Ellen Graham , Amy D Willis

This paper addresses the problem of correlation estimation in sets of compressed images. We consider a framework where images are represented under the form of linear measurements due to low complexity sensing or security requirements. We…

Computer Vision and Pattern Recognition · Computer Science 2011-12-20 Vijayaraghavan Thirumalai , Pascal Frossard

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

Methodology · Statistics 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

Censoring occurs when an outcome is unobserved beyond some threshold value. Methods that do not account for censoring produce biased predictions of the unobserved outcome. This paper introduces Type I Tobit Bayesian Additive Regression Tree…

Econometrics · Economics 2024-02-21 Eoghan O'Neill

Data contamination undermines the validity of Large Language Model evaluation by enabling models to rely on memorized benchmark content rather than true generalization. While prior work has proposed contamination detection methods, these…

Computation and Language · Computer Science 2026-01-22 Chaymaa Abbas , Nour Shamaa , Mariette Awad

Data collection in economically constrained countries often necessitates using approximate and biased measurements due to the low-cost of the sensors used. This leads to potentially invalid predictions and poor policies or decision making.…

Machine Learning · Computer Science 2019-12-02 Michael T. Smith , Joel Ssematimba , Mauricio A. Alvarez , Engineer Bainomugisha

Accurate estimates of microbial species abundances are needed to advance our understanding of the role that microbiomes play in human and environmental health. However, artificially constructed microbiomes demonstrate that intuitive…

Methodology · Statistics 2025-03-17 David S Clausen , Amy D Willis

The US Census Bureau will deliberately corrupt data sets derived from the 2020 US Census, enhancing the privacy of respondents while potentially reducing the precision of economic analysis. To investigate whether this trade-off is…

Econometrics · Economics 2024-02-13 Anish Agarwal , Rahul Singh

Network intrusion detection sensors are usually built around low level models of network traffic. This means that their output is of a similarly low level and as a consequence, is difficult to analyze. Intrusion alert correlation is the…

Cryptography and Security · Computer Science 2010-07-05 Gianni Tedesco , Uwe Aickelin

The analysis of environmental mixtures is of growing importance in environmental epidemiology, and one of the key goals in such analyses is to identify exposures and their interactions that are associated with adverse health outcomes.…

Methodology · Statistics 2021-03-22 Srijata Samanta , Joseph Antonelli

Measurement error arises through a variety of mechanisms. A rich literature exists on the bias introduced by covariate measurement error and on methods of analysis to address this bias. By comparison, less attention has been given to errors…

Methodology · Statistics 2018-11-27 Pamela Shaw , Jiwei He , Bryan Shepherd

The setting of a right-censored random sample subject to contamination is considered. In various fields, expert information is often available and used to overcome the contamination. This paper integrates expert knowledge into the…

Methodology · Statistics 2023-03-28 Martin Bladt , Christian Furrer

In microbiome studies, it is often of great interest to identify clusters or partitions of microbiome profiles within a study population and to characterize the distinctive attributes of each resulting microbial community. While raw counts…

Methodology · Statistics 2025-08-18 Zhongmao Liu , Xiaohui Yin , Yanjiao Zhou , Gen Li , Kun Chen

Metagenomics sequencing is routinely applied to quantify bacterial abundances in microbiome studies, where the bacterial composition is estimated based on the sequencing read counts. Due to limited sequencing depth and DNA dropouts, many…

Methodology · Statistics 2019-04-26 Yuanpei Cao , Anru Zhang , Hongzhe Li

We study robust regression under a contamination model in which covariates are clean while the responses may be corrupted in an adaptive manner. Unlike the classical Huber's contamination model, where both covariates and responses may be…

Statistics Theory · Mathematics 2026-04-07 Ilias Diakonikolas , Chao Gao , Daniel M. Kane , Ankit Pensia , Dong Xie

Air pollution constitutes the highest environmental risk factor in relation to heath. In order to provide the evidence required for health impact analyses, to inform policy and to develop potential mitigation strategies comprehensive…

Applications · Statistics 2021-08-23 Matthew L. Thomas , Gavin Shaddick , Daniel Simpson , Kees de Hoogh , James V. Zidek

Several recent imaging experiments access the equilibrium density profiles of interacting particles confined to a two-dimensional substrate. When these particles are in a fluid phase, we show that such data yields precise information…

Disordered Systems and Neural Networks · Physics 2016-08-31 Ankush Sengupta , Surajit Sengupta , Gautam I. Menon

Content moderation typically combines the efforts of human moderators and machine learning models. However, these systems often rely on data where significant disagreement occurs during moderation, reflecting the subjective nature of…

Computation and Language · Computer Science 2025-09-01 Guillermo Villate-Castillo , Javier Del Ser , Borja Sanz

In many fields of study, we only observe lower bounds on the true response value of some experiments. When fitting a regression model to predict the distribution of the outcomes, we cannot simply drop these right-censored observations, but…

Artificial Intelligence · Computer Science 2020-09-30 Katharina Eggensperger , Kai Haase , Philipp Müller , Marius Lindauer , Frank Hutter