English
Related papers

Related papers: Elementary methods provide more replicable results…

200 papers

We study the problem of detecting planted solutions in a random satisfiability formula. Adopting the formalism of hypothesis testing in statistical analysis, we describe the minimax optimal rates of detection. Our analysis relies on the…

Statistics Theory · Mathematics 2015-02-10 Quentin Berthet

Microstructure reconstruction is a key enabler of process-structure-property linkages, a central topic in materials engineering. Revisiting classical optimization-based reconstruction techniques,they are recognized as a powerful framework…

Materials Science · Physics 2021-03-19 Paul Seibert , Marreddy Ambati , Alexander Raßloff , Markus Kästner

Recent attacks of various viruses with having deep and extensive impact at a global scale has warranted that microbiome be studied extensively and in a robust analytic framework. Microbiome typically refers to the collective genomes of such…

Applications · Statistics 2023-03-30 M. Bhattacharjee

The population size ("abundance") of wildlife species has central interest in ecological research and management. Distance sampling is a dominant approach to the estimation of wildlife abundance for many vertebrate animal species. One…

Methodology · Statistics 2025-04-18 Benjamin R. Baer , Len Thomas , Stephen T. Buckland

Anomaly detection in time series is a complex task that has been widely studied. In recent years, the ability of unsupervised anomaly detection algorithms has received much attention. This trend has led researchers to compare only…

Machine Learning · Computer Science 2022-09-13 Julien Audibert , Pietro Michiardi , Frédéric Guyard , Sébastien Marti , Maria A. Zuluaga

Strong empirical evidence that one machine-learning algorithm A outperforms another one B ideally calls for multiple trials optimizing the learning pipeline over sources of variation such as data sampling, data augmentation, parameter…

Microbiome compositional data are often high-dimensional, sparse, and exhibit pervasive cross-sample heterogeneity. Generative modeling is a popular approach to analyze such data, and effective generative models must accurately characterize…

Methodology · Statistics 2025-01-03 Zhuoqun Wang , Jialiang Mao , Li Ma

High throughput sequencing (HTS)-based technology enables identifying and quantifying non-culturable microbial organisms in all environments. Microbial sequences have enhanced our understanding of the human microbiome, the soil and plant…

Applications · Statistics 2021-03-09 Pratheepa Jeganathan , Susan P. Holmes

Diffusion probabilistic models (DPMs) represent a class of powerful generative models. Despite their success, the inference of DPMs is expensive since it generally needs to iterate over thousands of timesteps. A key problem in the inference…

Machine Learning · Computer Science 2022-05-04 Fan Bao , Chongxuan Li , Jun Zhu , Bo Zhang

In this review, we present econometric and statistical methods for analyzing randomized experiments. For basic experiments we stress randomization-based inference as opposed to sampling-based inference. In randomization-based inference,…

Methodology · Statistics 2017-10-26 Susan Athey , Guido Imbens

To facilitate effective decision-making, precipitation datasets should include uncertainty estimates. Quantile regression with machine learning has been proposed for issuing such estimates. Distributional regression offers distinct…

Machine Learning · Computer Science 2025-01-07 Georgia Papacharalampous , Hristos Tyralis , Nikolaos Doulamis , Anastasios Doulamis

Large annotated datasets are crucial for the success of deep neural networks, but labeling data can be prohibitively expensive in domains such as medical imaging. This work tackles the subset selection problem: selecting a small set of the…

Machine Learning · Computer Science 2025-09-29 Noga Bar , Raja Giryes

Recent methodological advances are enabling better examination of speciation and extinction processes and patterns. A major open question is the origin of large discrepancies in species number between groups of the same age. Existing…

Populations and Evolution · Quantitative Biology 2015-09-01 Sacha Laurent , Marc Robinson-Rechavi , Nicolas Salamin

Asymptotic goodness-of-fit methods in contingency table analysis can struggle with sparse data, especially in multi-way tables where it can be infeasible to meet sample size requirements for a robust application of distributional…

Methodology · Statistics 2023-12-29 Shishir Agrawal , Luis David Garcia Puente , Minho Kim , Flavia Sancier-Barbosa

Multimodal single-cell technologies enable the simultaneous collection of diverse data types from individual cells, enhancing our understanding of cellular states. However, the integration of these datatypes and modeling the…

Machine Learning · Computer Science 2023-11-22 Bhavya Mehta , Nirmit Deliwala , Madhav Chandane

Despite the accelerating presence of exploratory causal analysis in modern science and medicine, the available non-experimental methods for validating causal models are not well characterized. One of the most popular methods is to evaluate…

Methodology · Statistics 2025-03-20 Ritwick Banerjee , Bryan Andrews , Erich Kummerfeld

Rerandomization is a strategy of increasing efficiency as compared to complete randomization. The idea with rerandomization is that of removing allocations with imbalance in the observed covariates and then randomizing within the set of…

Methodology · Statistics 2019-11-07 Junni L. Zhang , Per Johansson

Detecting differences in gene expression is an important part of single-cell RNA sequencing experiments, and many statistical methods have been developed for this aim. Most differential expression analyses focus on comparing expression…

Deep sequencing has become one of the most popular tools for transcriptome profiling in biomedical studies. While an abundance of computational methods exists for "normalizing" sequencing data to remove unwanted between-sample variations…

Genomics · Quantitative Biology 2022-01-14 Yannick Düren , Johannes Lederer , Li-Xuan Qin

Motivation: Microarray data has been recently been shown to be efficacious in distinguishing closely related cell types that often appear in the diagnosis of cancer. It is useful to determine the minimum number of genes needed to do such a…

Biological Physics · Physics 2007-05-23 J. M. Deutsch