English
Related papers

Related papers: Elementary methods provide more replicable results…

200 papers

The boom of DL technology leads to massive DL models built and shared, which facilitates the acquisition and reuse of DL models. For a given task, we encounter multiple DL models available with the same functionality, which are considered…

Software Engineering · Computer Science 2021-03-10 Linghan Meng , Yanhui Li , Lin Chen , Zhi Wang , Di Wu , Yuming Zhou , Baowen Xu

A new gradient-based adaptive sampling method is proposed for design of experiments applications which balances space filling, local refinement, and error minimization objectives while reducing reliance on delicate tuning parameters. High…

Methodology · Statistics 2024-05-09 Lucas Caparini , Gwynn J. Elfring , Mauricio Ponga

In microbial ecology studies, the most commonly used ways of investigating alpha (within-sample) diversity are either to apply count-only measures such as Simpson's index to Operational Taxonomic Unit (OTU) groupings, or to use classical…

Populations and Evolution · Quantitative Biology 2013-05-03 Connor O. McCoy , Frederick A. Matsen

Quantitative methods for studying biodiversity have been traditionally rooted in the classical theory of finite frequency tables analysis. However, with the help of modern experimental tools, like high throughput sequencing, we now begin to…

Methodology · Statistics 2015-12-22 Maciej Pietrzak , Grzegorz A. Rempała , Michał Seweryn , Jacek Wesołowski

Computational analysis methods including machine learning have a significant impact in the fields of genomics and medicine. High-throughput gene expression analysis methods such as microarray technology and RNA sequencing produce enormous…

Genomics · Quantitative Biology 2022-09-28 Nikita Bhandari , Rahee Walambe , Ketan Kotecha , Satyajeet Khare

Detecting predictive biomarkers from multi-omics data is important for precision medicine, to improve diagnostics of complex diseases and for better treatments. This needs substantial experimental efforts that are made difficult by the…

Quantitative Methods · Quantitative Biology 2021-06-08 Betül Güvenç Paltun , Samuel Kaski , Hiroshi Mamitsuka

Replication studies estimate the replicability rate of scientific results by aggregating binary verdicts of experiments. Exact replications are rarely attainable, so most replication sequences are non-exact. Experiments differ in ways that…

Applications · Statistics 2026-04-30 Berna Devezer , Erkan O. Buzbas

Background: In the metagenome assembly of a microbiome community, we may think abundant species would be easier to assemble due to their deeper coverage. However, this conjucture is rarely tested. We often do not know how many abundant…

Genomics · Quantitative Biology 2022-11-23 Xiaowen Feng , Heng Li

The increasing popularity of regression discontinuity methods for causal inference in observational studies has led to a proliferation of different estimating strategies, most of which involve first fitting non-parametric regression models…

Methodology · Statistics 2018-06-11 Guido Imbens , Stefan Wager

We report on an empirical study of the main strategies for quantile regression in the context of stochastic computer experiments. To ensure adequate diversity, six metamodels are presented, divided into three categories based on order…

Machine Learning · Statistics 2020-01-22 Léonard Torossian , Victor Picheny , Robert Faivre , Aurélien Garivier

Observational studies can play a useful role in assessing the comparative effectiveness of competing treatments. In a clinical trial the randomization of participants to treatment and control groups generally results in well-balanced groups…

Dendrograms are a way to represent evolutionary relationships between organisms. Nowadays, these are inferred based on the comparison of genes or protein sequences by taking into account their differences and similarities. The genetic…

Molecular Networks · Quantitative Biology 2019-09-09 Daniel Gamermann , Arnau Montagud , J. Alberto Conejero , Pedro Fernández de Córdoba , Javier F. Urchueguía

This paper evaluates six strategies for mitigating imbalanced data: oversampling, undersampling, ensemble methods, specialized algorithms, class weight adjustments, and a no-mitigation approach referred to as the baseline. These strategies…

Machine Learning · Computer Science 2023-11-13 Jacques Wainer

Missing values of varying patterns and rates in real-world tabular data pose a significant challenge in developing reliable data-driven models. The most commonly used statistical and machine learning methods for missing value imputation may…

Machine Learning · Computer Science 2025-03-26 Ibna Kowsar , Shourav B. Rabbani , Yina Hou , Manar D. Samad

We describe an automatic procedure for determining abundances from high resolution spectra. Such procedures are becoming increasingly important as large amounts of data are delivered from 8m telescopes and their high-multiplexing fiber…

Astrophysics · Physics 2009-11-07 Piercarlo Bonifacio , Elisabetta Caffau

High-throughput sequencing technology allows us to test the compositional difference of bacteria in different populations. One important feature of human microbiome data is that it often includes a large number of zeros. Such data can be…

Methodology · Statistics 2022-08-23 Wanjie Wang , Eric Z. Chen , Hongzhe Li

A series of ten plant species belonging to Magnoliopsida - Dicotyledons class were analyzed in terms of chemical compounds distribution of abundance, starting from the assumption that these distributions should give a picture of…

Applications · Statistics 2011-06-28 Lorentz Jäntschi , Sorana D. Bolboacă , Radu E. Sestraş

Several statistical approaches based on reproducing kernels have been proposed to detect abrupt changes arising in the full distribution of the observations and not only in the mean or variance. Some of these approaches enjoy good…

Statistics Theory · Mathematics 2017-10-13 Alain Celisse , Guillemette Marot , Morgane Pierre-Jean , Guillem Rigaill

We wish to estimate the total number of classes in a population based on sample counts, especially in the presence of high latent diversity. Drawing on probability theory that characterizes distributions on the integers by ratios of…

Methodology · Statistics 2014-12-10 A. Willis , J. Bunge

Linear and Quadratic Discriminant Analysis are well-known classical methods but can heavily suffer from non-Gaussian distributions and/or contaminated datasets, mainly because of the underlying Gaussian assumption that is not robust. To…

Machine Learning · Statistics 2022-01-11 Pierre Houdouin , Frédéric Pascal , Matthieu Jonckheere , Andrew Wang