中文
相关论文

相关论文: Multi-Target Tobit Models for Completing Water Qua…

200 篇论文

Urban water quality is of great importance to our daily lives. Prediction of urban water quality help control water pollution and protect human health. However, predicting the urban water quality is a challenging task since the water…

计算机与社会 · 计算机科学 2016-11-01 Ye Liu , Yuxuan Liang , Shuming Liu , David S. Rosenblum , Yu Zheng

The gut microbiome plays a crucial role in human health, making it a corner stone of modern biomedical research. To study its structure and dynamics, machine learning models are increasingly used to identify key microbial patterns…

应用统计 · 统计学 2025-07-08 Alexandre Chaussard , Anna Bonnet , Sylvain Le Corff , Harry Sokol

Models of biological systems often have many unknown parameters that must be determined in order for model behavior to match experimental observations. Commonly-used methods for parameter estimation that return point estimates of the…

定量方法 · 定量生物学 2018-01-31 Sanjana Gupta , Liam Hainsworth , Justin S. Hogg , Robin E. C. Lee , James R. Faeder

The growing need to analyze large collections of documents has led to great developments in topic modeling. Since documents are frequently associated with other related variables, such as labels or ratings, much interest has been placed on…

机器学习 · 统计学 2018-08-20 Filipe Rodrigues , Mariana Lourenço , Bernardete Ribeiro , Francisco Pereira

In many applications, data can be heterogeneous in the sense of spanning latent groups with different underlying distributions. When predictive models are applied to such data the heterogeneity can affect both predictive performance and…

机器学习 · 统计学 2022-05-04 Thomas Lartigue , Sach Mukherjee

Accurate control of quantum systems requires precise measurement of the parameters that govern the dynamics, including control fields and interactions with the environment. Parameters will drift in time and experiments interleave protocols…

量子物理 · 物理学 2020-02-11 Swarnadeep Majumder , Leonardo Andreta de Castro , Kenneth R. Brown

We reduce measurement errors in a quantum computer using machine learning techniques. We exploit a simple yet versatile neural network to classify multi-qubit quantum states, which is trained using experimental data. This flexible approach…

Data imbalance is easily found in annotated data when the observations of certain continuous label values are difficult to collect for regression tasks. When they come to molecule and polymer property predictions, the annotated graph…

机器学习 · 计算机科学 2023-05-23 Gang Liu , Tong Zhao , Eric Inae , Tengfei Luo , Meng Jiang

Camera traps have become a common tool for wildlife monitoring efforts in ecological research and biodiversity conservation. Wildlife classification models have benefited from the increase in wildlife visual data. These models reach high…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Mufhumudzi Muthivhi , Jiahao Huo , Fredrik Gustafsson , Terence L. van Zyl

Two-class classification problems are often characterized by an imbalance between the number of majority and minority datapoints resulting in poor classification of the minority class in particular. Traditional approaches, such as…

机器学习 · 计算机科学 2025-07-11 Karen Medlin , Sven Leyffer , Krishnan Raghavan

Accurately predicting the time of occurrence of an event of interest is a critical problem in longitudinal data analysis. One of the main challenges in this context is the presence of instances whose event outcomes become unobservable after…

机器学习 · 计算机科学 2017-12-26 Ping Wang , Yan Li , Chandan K. Reddy

We study the problem of estimating the probability density function of a circular random variable subject to censoring. To this end, we propose a fully computable quotient estimator that combines a projection estimator on linear sieves with…

统计理论 · 数学 2025-08-11 Nicolas Conanec

We study robust regression under a contamination model in which covariates are clean while the responses may be corrupted in an adaptive manner. Unlike the classical Huber's contamination model, where both covariates and responses may be…

统计理论 · 数学 2026-04-07 Ilias Diakonikolas , Chao Gao , Daniel M. Kane , Ankit Pensia , Dong Xie

Multi-label classification poses challenges due to imbalanced and noisy labels in training data. We propose a unified data augmentation method, named BalanceMix, to address these challenges. Our approach includes two samplers for imbalanced…

机器学习 · 计算机科学 2023-12-13 Hwanjun Song , Minseok Kim , Jae-Gil Lee

Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may have been trained on proprietary datasets or the same benchmarks used for evaluation. This…

机器学习 · 计算机科学 2026-05-11 Pengrun Huang , Kamalika Chaudhuri , Yu-Xiang Wang

Machine learning models have dual-use potential, potentially serving both beneficial and malicious purposes. The development of open-source models in chemistry has specifically surfaced dual-use concerns around toxicological data and…

机器学习 · 计算机科学 2025-10-28 Quintina L. Campbell , Jonathan Herington , Andrew D. White

Water pollution is a major global environmental problem, and it poses a great environmental risk to public health and biological diversity. This work is motivated by assessing the potential environmental threat of coal mining through…

统计方法学 · 统计学 2019-05-17 Amal Agarwal , Lingzhou Xue

The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core proposal is to spike the training data by intentionally contaminating some test examples at…

统计方法学 · 统计学 2026-05-26 Johnny Tian-Zheng Wei , Jerry Li , Ameya Godbole , Robin Jia

Indiscriminate data poisoning attacks aim to decrease a model's test accuracy by injecting a small amount of corrupted training data. Despite significant interest, existing attacks remain relatively ineffective against modern machine…

机器学习 · 计算机科学 2023-06-07 Yiwei Lu , Gautam Kamath , Yaoliang Yu

Large language models generate high-quality responses with potential misinformation, underscoring the need for regulation by distinguishing AI-generated and human-written texts. Watermarking is pivotal in this context, which involves…

机器学习 · 计算机科学 2024-06-07 Mingjia Huo , Sai Ashish Somayajula , Youwei Liang , Ruisi Zhang , Farinaz Koushanfar , Pengtao Xie