中文
相关论文

相关论文: Log-Paradox: Necessary and sufficient conditions f…

200 篇论文

Given only data generated by a standard confounding graph with unobserved confounder, the Average Treatment Effect (ATE) is not identifiable. To estimate the ATE, a practitioner must then either (a) collect deconfounded data;(b) run a…

机器学习 · 统计学 2021-03-09 Kyra Gan , Andrew A. Li , Zachary C. Lipton , Sridhar Tayur

We investigate the likelihood ratio test for a large block-diagonal covariance matrix with an increasing number of blocks under the null hypothesis. While so far the likelihood ratio statistic has only been studied for normal populations,…

统计理论 · 数学 2024-08-01 Nina Dörnemann

The classical approach to analyze time-to-event data, e.g. in clinical trials, is to fit Kaplan-Meier curves yielding the treatment effect as the hazard ratio between treatment groups. Afterwards commonly a log-rank test is performed in…

统计方法学 · 统计学 2020-09-16 Kathrin Möllenhoff , Achim Tresch

This work is motivated by learning the individualized minimal clinically important difference, a vital concept to assess clinical importance in various biomedical studies. We formulate the scientific question into a high-dimensional…

统计方法学 · 统计学 2023-03-28 Huijie Feng , Jingyi Duan , Yang Ning , Jiwei Zhao

Distribution alignment has many applications in deep learning, including domain adaptation and unsupervised image-to-image translation. Most prior work on unsupervised distribution alignment relies either on minimizing simple non-parametric…

机器学习 · 计算机科学 2020-10-27 Ben Usman , Avneesh Sud , Nick Dufour , Kate Saenko

There is a considerable literature in case-control logistic regression on whether or not non-confounding covariates should be adjusted for. However, only limited and ad hoc theoretical results are available on this important topic. A…

统计理论 · 数学 2023-05-19 Siliang Zhang , Jinbo Chen , Zhiliang Ying , Hong Zhang

There is a well-known problem in Null Hypothesis Significance Testing: many statistically significant results fail to replicate in subsequent experiments. We show that this problem arises because standard `point-form null' significance…

统计方法学 · 统计学 2025-02-06 Fintan Costello , Paul Watts

Logs are essential for diagnosing failures and conducting retrospective studies, leading many software organizations to retain log messages for a long time. Nevertheless, the volume of generated log data grows rapidly as software systems…

软件工程 · 计算机科学 2026-03-24 Shiwen Shan , Yintong Huo , Hongzhan Zhong , Zhining Wang , Yuxin Su , Zibin Zheng

Counterfactual explanations can be obtained by identifying the smallest change made to a feature vector to qualitatively influence a prediction; for example, from 'loan rejected' to 'awarded' or from 'high risk of cardiovascular disease' to…

机器学习 · 计算机科学 2020-05-05 Martin Pawelczyk , Johannes Haug , Klaus Broelemann , Gjergji Kasneci

Inference based on the penalized density ratio model is proposed and studied. The model under consideration is specified by assuming that the log--likelihood function of two unknown densities is of some parametric form. The model has been…

统计理论 · 数学 2008-07-17 Konstantinos Fokianos

Trained models are often composed with post-hoc transforms such as temperature scaling (TS), ensembling and stochastic weight averaging (SWA) to improve performance, robustness, uncertainty estimation, etc. However, such transforms are…

机器学习 · 计算机科学 2024-10-07 Rishabh Ranjan , Saurabh Garg , Mrigank Raman , Carlos Guestrin , Zachary Lipton

We consider large-scale studies in which thousands of significance tests are performed simultaneously. In some of these studies, the multiple testing procedure can be severely biased by latent confounding factors such as batch effects and…

统计方法学 · 统计学 2016-06-21 Jingshu Wang , Qingyuan Zhao , Trevor Hastie , Art B. Owen

Statistical agencies and other institutions collect data under the promise to protect the confidentiality of respondents. When releasing microdata samples, the risk that records can be identified must be assessed. To this aim, a widely…

应用统计 · 统计学 2015-06-03 Cinzia Carota , Maurizio Filippone , Roberto Leombruni , Silvia Polettini

[See paper for full abstract] Meta-analysis is a crucial tool for answering scientific questions. It is usually conducted on a relatively small amount of ``trusted'' data -- ideally from randomized, controlled trials -- which allow causal…

机器学习 · 统计学 2024-07-15 Shiva Kaul , Geoffrey J. Gordon

We consider the problem of detection of sparse anomalies when monitoring a large number of data streams continuously in time. This problem is addressed using anytime-valid tests. In the context of a normal-means model and for a fixed…

统计理论 · 数学 2025-07-01 Muriel F. Pérez-Ortiz , Rui M. Castro

Supervised learning has become a cornerstone of modern machine learning, yet a comprehensive theory explaining its effectiveness remains elusive. Empirical phenomena, such as neural analogy-making and the linear representation hypothesis,…

Citations are increasingly used for research evaluations. It is therefore important to identify factors affecting citation scores that are unrelated to scholarly quality or usefulness so that these can be taken into account. Regression is…

数字图书馆 · 计算机科学 2015-11-02 Mike Thelwall , Paul Wilson

The ability to reliably distinguish human-written text from that generated by large language models is of profound societal importance. The dominant approach to this problem exploits the likelihood hypothesis: that machine-generated text…

计算与语言 · 计算机科学 2026-05-08 Tom Kempton , Viktor Drobnyi , Maeve Madigan , Stuart Burrell

Many practical studies rely on hypothesis testing procedures applied to data sets with missing information. An important part of the analysis is to determine the impact of the missing data on the performance of the test, and this can be…

统计方法学 · 统计学 2011-02-15 Dan L. Nicolae , Xiao-Li Meng , Augustine Kong

To solve key biomedical problems, experimentalists now routinely measure millions or billions of features (dimensions) per sample, with the hope that data science techniques will be able to build accurate data-driven inferences. Because…