English
Related papers

Related papers: Asymmetric Tobit analysis for correlation estimati…

200 papers

Accurately quantifying and removing submerged underwater waste plays a crucial role in safeguarding marine life and preserving the environment. While detecting floating and surface debris is relatively straightforward, quantifying submerged…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Jaskaran Singh Walia , Karthik Seemakurthy

This study investigates the application of an artificial neural network framework for analysing water pollution caused by solids. Water pollution by suspended solids poses significant environmental and health risks. Traditional methods for…

Machine Learning · Computer Science 2026-01-12 I. Luviano Soto , Y. Concha Sánchez , A. Raya

Overlapping asymmetric datasets are common in data science and pose questions of how they can be incorporated together into a predictive analysis. In healthcare datasets there is often a small amount of information that is available for a…

Methodology · Statistics 2023-11-21 Matthew McTeer , Robin Henderson , Quentin M Anstee , Paolo Missier

Lossy compression plays a growing role in scientific simulations where the cost of storing their output data can span terabytes. Using error bounded lossy compression reduces the amount of storage for each simulation; however, there is no…

Applications · Statistics 2021-11-30 David Krasowska , Julie Bessac , Robert Underwood , Jon C. Calhoun , Sheng Di , Franck Cappello

In the face of growing needs for water and energy, a fundamental understanding of the environmental impacts of human activities becomes critical for managing water and energy resources, remedying water pollution, and making regulatory…

Signal Processing · Electrical Eng. & Systems 2019-08-30 Guanjie Zheng , Mengqi Liu , Tao Wen , Hongjian Wang , Huaxiu Yao , Susan L. Brantley , Zhenhui Li

This paper studies density estimation under pointwise loss in the setting of contamination model. The goal is to estimate $f(x_0)$ at some $x_0\in\mathbb{R}$ with i.i.d. observations, $$ X_1,\dots,X_n\sim (1-\epsilon)f+\epsilon g, $$ where…

Statistics Theory · Mathematics 2018-07-30 Haoyang Liu , Chao Gao

Public health researchers often estimate health effects of exposures (e.g., pollution, diet, lifestyle) that cannot be directly measured for study subjects. A common strategy in environmental epidemiology is to use a first-stage (exposure)…

Methodology · Statistics 2014-06-03 Adam A. Szpiro , Christopher J. Paciorek

Correlation networks are commonly used to infer associations between microbes and metabolites. The resulting p-values are then corrected for multiple comparisons using existing methods such as the Benjamini and Hochberg procedure to control…

Methodology · Statistics 2025-06-17 Jing Ma

The coexistence of multiple defect categories as well as the substantial class imbalance problem significantly impair the detection of sewer pipeline defects. To solve this problem, a multi-label pipe defect recognition method is proposed…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Xin Zuo , Yu Sheng , Jifeng Shen , Yongwei Shan

Conformal prediction is a non-parametric technique for constructing prediction intervals or sets from arbitrary predictive models under the assumption that the data is exchangeable. It is popular as it comes with theoretical guarantees on…

Machine Learning · Statistics 2025-12-01 Jase Clarkson , Wenkai Xu , Mihai Cucuringu , Yvik Swan , Gesine Reinert

Unmeasured confounding is a major challenge for identifying causal relationships from non-experimental data. Here, we propose a method that can accommodate unmeasured discrete confounding. Extending recent identifiability results in deep…

Machine Learning · Computer Science 2024-08-13 Patrick Burauel , Frederick Eberhardt , Michel Besserve

Classical semiparametric inference with missing outcome data is not robust to contamination of the observed data and a single observation can have arbitrarily large influence on estimation of a parameter of interest. This sensitivity is…

Methodology · Statistics 2021-03-02 Eva Cantoni , Xavier de Luna

One of the significant problems associated with imbalanced data classification is the lack of reliable metrics. This runs primarily from the fact that for most real-life (as well as commonly used benchmark) problems, we do not have…

Machine Learning · Computer Science 2024-04-16 Szymon Wojciechowski , Michał Woźniak

The amount of quality data in many machine learning tasks is limited to what is available locally to data owners. The set of quality data can be expanded through trading or sharing with external data agents. However, data buyers need…

Machine Learning · Statistics 2025-07-21 Martin V. Vejling , Shashi Raj Pandey , Christophe A. N. Biscio , Petar Popovski

One way to quantify exposure to air pollution and its constituents in epidemiologic studies is to use an individual's nearest monitor. This strategy results in potential inaccuracy in the actual personal exposure, introducing bias in…

Accurate pest population monitoring and tracking their dynamic changes are crucial for precision agriculture decision-making. A common limitation in existing vision-based automatic pest counting research is that models are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Xumin Gao , Mark Stevens , Grzegorz Cielniak

Inferring the true demand for a product or a service from aggregate data is often challenging due to the limited available supply, thus resulting in observations that are censored and correspond to the realized demand, thereby not…

Machine Learning · Computer Science 2025-01-22 Filipe Rodrigues

Background: Trace quantities of contaminating DNA are widespread in the laboratory environment, but their presence has received little attention in the context of high throughput sequencing. This issue is highlighted by recent works that…

Genomics · Quantitative Biology 2015-06-18 Richard W Lusk

When data are collected subject to a detection limit, observations below the detection limit may be considered censored. In addition, the domain of such observations may be restricted; for example, values may be required to be non-negative.…

Applications · Statistics 2020-06-30 Justin R. Williams , Hyung-Woo Kim , Catherine M. Crespi

When data are stored across multiple locations, directly pooling all the data together for statistical analysis may be impossible due to communication costs and privacy concerns. Distributed computing systems allow the analysis of such…

Methodology · Statistics 2025-02-27 Xian Li , Xuan Liang , A. H. Welsh , Tao Zou