English
Related papers

Related papers: The Search for Truth through Data: NP Decision Pro…

200 papers

Penalized regression models such as the Lasso have proved useful for variable selection in many fields - especially for situations with high-dimensional data where the numbers of predictors far exceeds the number of observations. These…

Methodology · Statistics 2014-03-19 Kasper Brink-Jensen , Claus Thorn Ekstrøm

A typical power calculation is performed by replacing unknown population-level quantities in the power function with what is observed in external studies. Many authors and practitioners view this as an assumed value of power and offer the…

Applications · Statistics 2026-05-12 Geoffrey S Johnson

Practical problems with missing data are common, and statistical methods have been developed concerning the validity and/or efficiency of statistical procedures. On a central focus, there have been longstanding interests on the mechanism…

Methodology · Statistics 2020-03-26 Rui Duan , C. Jason Liang , Pamela Shaw , Cheng Yong Tang , Yong Chen

Increasing accessibility of data to researchers makes it possible to conduct massive amounts of statistical testing. Rather than follow a carefully crafted set of scientific hypotheses with statistical analysis, researchers can now test…

Genomics · Quantitative Biology 2016-09-08 Olga A. Vsevolozhskaya , Chia-Ling Kuo , Gabriel Ruiz , Luda Diatchenko , Dmitri V. Zaykin

In this work, we consider a binary hypothesis testing problem involving a group of human decision-makers. Due to the nature of human behavior, each human decision-maker observes the phenomenon of interest sequentially up to a random length…

Signal Processing · Electrical Eng. & Systems 2023-01-26 Nandan Sriranga , Baocheng Geng , Pramod K. Varshney

This paper describes a decision theoretic formulation of learning the graphical structure of a Bayesian Belief Network from data. This framework subsumes the standard Bayesian approach of choosing the model with the largest posterior…

Artificial Intelligence · Computer Science 2013-02-01 Paola Sebastiani , Marco Ramoni

Two main procedures characterize the way in which social actors evaluate the qualities of the options in decision-making processes: they either seek to evaluate their intrinsic qualities (individual learners), or they rely on the opinion of…

Physics and Society · Physics 2024-07-31 Arkadiusz Jędrzejewski , Laura Hernández

It is almost always easier to find an accurate-but-complex model than an accurate-yet-simple model. Finding optimal, sparse, accurate models of various forms (linear models with integer coefficients, decision sets, rule lists, decision…

Machine Learning · Computer Science 2022-05-16 Lesia Semenova , Cynthia Rudin , Ronald Parr

Large language models (LLMs) are increasingly used as information sources, yet small changes in semantic framing can destabilize their truth judgments. We propose P-StaT (Perturbation Stability of Truth), an evaluation framework for testing…

Computation and Language · Computer Science 2026-01-21 Samantha Dies , Courtney Maynard , Germans Savcisens , Tina Eliassi-Rad

We discuss problems the null hypothesis significance testing (NHST) paradigm poses for replication and more broadly in the biomedical and social sciences as well as how these problems remain unresolved by proposals involving modified…

Methodology · Statistics 2021-07-21 Blakeley B. McShane , David Gal , Andrew Gelman , Christian Robert , Jennifer L. Tackett

The adequate use of information measured in a continuous manner along a period of time represents a methodological challenge. In the last decades, most of traditional statistical procedures have been extended for accommodating these…

Methodology · Statistics 2025-12-04 Pablo Martinez-Camblor

We prove that it is NP-hard to properly PAC learn decision trees with queries, resolving a longstanding open problem in learning theory (Bshouty 1993; Guijarro-Lavin-Raghavan 1999; Mehta-Raghavan 2002; Feldman 2016). While there has been a…

Computational Complexity · Computer Science 2023-07-11 Caleb Koch , Carmen Strassle , Li-Yang Tan

We use p-values to identify the threshold level at which a regression function takes off from its baseline value, a problem motivated by applications in toxicological and pharmacological dose-response studies and environmental statistics.…

Methodology · Statistics 2011-06-13 Atul Mallik , Bodhisattva Sen , Moulinath Banerjee , George Michailidis

Counterfactual reasoning is an important paradigm applicable in many fields, such as healthcare, economics, and education. In this work, we propose a novel method to address the issue of \textit{selection bias}. We learn two groups of…

Machine Learning · Computer Science 2019-12-20 Zichen Zhang , Qingfeng Lan , Lei Ding , Yue Wang , Negar Hassanpour , Russell Greiner

Despite extensive focus on techniques for evaluating the performance of two learning algorithms on a single dataset, the critical challenge of developing statistical tests to compare multiple algorithms across various datasets has been…

Machine Learning · Computer Science 2025-12-16 Mohammad Abu-Shaira , Weishi Shi

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, comprised of both a…

Machine Learning · Statistics 2022-06-08 Bin Yu , Karl Kumbier

We use p-values as a discrepancy criterion for identifying the threshold value at which a regression function takes off from its baseline value -- a problem that is motivated by applications in omics experiments, systems engineering,…

Methodology · Statistics 2010-08-26 Bodhisattva Sen , Moulinath Banerjee , George Michialidis

Attacks on the P-value are nothing new, but the recent attacks are increasingly more serious. They come from more mainstream sources, with widening targets such as a call to retire the significance testing altogether. While well meaning, I…

Other Statistics · Statistics 2022-01-11 Yudi Pawitan

Background: Well-designed phase II trials must have acceptable error rates relative to a pre-specified success criterion, usually a statistically significant p-value. Such standard designs may not always suffice from a clinical perspective…

Applications · Statistics 2020-02-10 Satrajit Roychoudhury , Nicolas Scheuer , Beat Neuenschwander

This thesis focuses on the discovery of stochastic differential equations (SDEs) and stochastic partial differential equations (SPDEs) from noisy and discrete time series. A major challenge is selecting the simplest possible correct model…

Machine Learning · Statistics 2025-07-08 Andonis Gerardos
‹ Prev 1 3 4 5 6 7 10 Next ›