中文
相关论文

相关论文: Multi-Target Tobit Models for Completing Water Qua…

200 篇论文

In human microbiome studies, sequencing reads data are often summarized as counts of bacterial taxa at various taxonomic levels specified by a taxonomic tree. This paper considers the problem of analyzing two repeated measurements of…

应用统计 · 统计学 2017-02-17 Pixu Shi , Hongzhe Li

Detection limits are common in biomedical and environmental studies, where key covariates or outcomes are censored below an assay-specific threshold. Standard approaches such as complete-case analysis, single-value substitution, and…

统计方法学 · 统计学 2025-12-12 Y. Xu , S. Tu L. Shao , T. Lin , X. M. Tu

Data quality is a key element for building and optimizing good learning models. Despite many attempts to characterize data quality, there is still a need for rigorous formalization and an efficient measure of the quality from available…

机器学习 · 计算机科学 2023-12-14 Jouseau Roxane , Salva Sébastien , Samir Chafik

Water quality has a direct impact on industry, agriculture, and public health. Algae species are common indicators of water quality. It is because algal communities are sensitive to changes in their habitats, giving valuable knowledge on…

计算机视觉与模式识别 · 计算机科学 2022-03-24 Peisheng Qian , Ziyuan Zhao , Haobing Liu , Yingcai Wang , Yu Peng , Sheng Hu , Jing Zhang , Yue Deng , Zeng Zeng

Our objective is to construct well-calibrated prediction sets for a time-to-event outcome subject to right-censoring with guaranteed coverage. Inspired by modern conformal inference, our approach avoids the need for a well-specified…

统计方法学 · 统计学 2026-01-27 Rebecca Farina , Eric J. Tchetgen Tchetgen , Arun Kumar Kuchibhotla

As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-training, commonly known as the data contamination problem. To…

计算与语言 · 计算机科学 2024-06-27 Kun Qian , Shunji Wan , Claudia Tang , Youzhi Wang , Xuanming Zhang , Maximillian Chen , Zhou Yu

Continuous high frequency water quality monitoring is becoming a critical task to support water management. Despite the advancements in sensor technologies, certain variables cannot be easily and/or economically monitored in-situ and in…

机器学习 · 计算机科学 2020-01-28 María Castrillo , Álvaro López García

Existing toxic detection models face significant limitations, such as lack of transparency, customization, and reproducibility. These challenges stem from the closed-source nature of their training data and the paucity of explanations for…

计算与语言 · 计算机科学 2025-01-24 Tinh Son Luong , Thanh-Thien Le , Thang Viet Doan , Linh Ngo Van , Thien Huu Nguyen , Diep Thi-Ngoc Nguyen

Motivated by the challenges in analyzing gut microbiome and metagenomic data, this work aims to tackle the issue of measurement errors in high-dimensional regression models that involve compositional covariates. This paper marks a…

统计方法学 · 统计学 2024-09-13 Huali Zhao , Tianying Wang

Tracking multiple time-varying states based on heterogeneous observations is a key problem in many applications. Here, we develop a statistical model and algorithm for tracking an unknown number of targets based on the probabilistic fusion…

信号处理 · 电气工程与系统科学 2022-01-10 Domenico Gaglione , Paolo Braca , Giovanni Soldi , Florian Meyer , Franz Hlawatsch , Moe Z. Win

We consider linear regression model estimation where the covariate of interest is randomly censored. Under a non-informative censoring mechanism, one may obtain valid estimates by deleting censored observations. However, this comes at a…

应用统计 · 统计学 2017-10-24 Folefac Atem , Roland A. Matsouaka

Censored response variables--where outcomes are only partially observed due to known bounds--arise in numerous scientific domains and present serious challenges for regression analysis. The Tobit model, a classical solution for handling…

统计方法学 · 统计学 2025-05-14 The Tien Mai

Annotating data for sensitive labels (e.g., disease, smoking) poses a potential threats to individual privacy in many real-world scenarios. To cope with this problem, we propose a novel setting to protect privacy of each instance, namely…

机器学习 · 计算机科学 2024-12-04 Zhongnian Li , Meng Wei , Peng Ying , Tongfeng Sun , Xinzheng Xu

Semi-supervised crowd counting is crucial for addressing the high annotation costs of densely populated scenes. Although several methods based on pseudo-labeling have been proposed, it remains challenging to effectively and accurately…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Maochen Yang , Zekun Li , Jian Zhang , Lei Qi , Yinghuan Shi

Chemical multisensor devices need calibration algorithms to estimate gas concentrations. Their possible adoption as indicative air quality measurements devices poses new challenges due to the need to operate in continuous monitoring modes…

人工智能 · 计算机科学 2020-02-14 S. De Vito , E. Esposito , M. Salvato , O. Popoola , F. Formisano , R. Jones , G. Di Francia

This paper deals with the Tobit Kalman filtering (TKF) process when the one-dimensional measurements are censored and the noises of the state-space model are coloured. Two improvements of the standard TKF process are proposed. Firstly, the…

统计方法学 · 统计学 2020-07-31 Kostas Loumponias

Current mathematical frameworks for predicting the flux state and macromolecular composition of the cell do not rely on thermodynamic constraints to determine the spontaneous direction of reactions. These predictions may be biologically…

最优化与控制 · 数学 2020-08-14 Amir Akbari , Bernhard O. Palsson

Ensuring safe water supplies requires effective water quality monitoring, especially in developing countries like Nepal, where contamination risks are high. This paper introduces various hybrid deep learning models to predict on the CCME…

机器学习 · 计算机科学 2025-10-28 Biplov Paneru , Bishwash Paneru

In this work, we present a framework to measure and mitigate intrinsic biases with respect to protected variables --such as gender-- in visual recognition tasks. We show that trained models significantly amplify the association of target…

计算机视觉与模式识别 · 计算机科学 2019-10-14 Tianlu Wang , Jieyu Zhao , Mark Yatskar , Kai-Wei Chang , Vicente Ordonez

Traditional statistical and machine learning methods typically assume that the training and test data follow the same distribution. However, this assumption is frequently violated in real-world applications, where the training data in the…

统计方法学 · 统计学 2025-07-08 Hanxuan Ye , Hongzhe Li