English
Related papers

Related papers: PolyLinkR: A linkage-sensitive gene set enrichment…

200 papers

Recently-developed genotype imputation methods are a powerful tool for detecting untyped genetic variants that affect disease susceptibility in genetic association studies. However, existing imputation methods require individual-level…

Applications · Statistics 2010-11-15 Xiaoquan Wen , Matthew Stephens

The R package MixMashNet provides an integrated framework for estimating and analyzing single and multilayer networks using Mixed Graphical Models (MGMs), accommodating continuous, count, and categorical variables. In the multilayer…

Despite the success of neural networks (NNs), there is still a concern among many over their "black box" nature. Why do they work? Here we present a simple analytic argument that NNs are in fact essentially polynomial regression models.…

Machine Learning · Computer Science 2019-04-11 Xi Cheng , Bohdan Khomtchouk , Norman Matloff , Pete Mohanty

The R package CVEK introduces a suite of flexible machine learning models and robust hypothesis tests for learning the joint nonlinear effects of multiple covariates in limited samples. It implements the Cross-validated Ensemble of Kernels…

Computation · Statistics 2020-12-22 Wenying Deng , Jeremiah Zhe Liu , Erin Lake , Brent A. Coull

Modern high-throughput sequencing assays efficiently capture not only gene expression and different levels of gene regulation but also a multitude of genome variants. Focused analysis of alternative alleles of variable sites at homologous…

Estimating sample size and statistical power is an essential part of a good study design. This R package allows users to conduct power analysis based on Monte Carlo simulations in settings in which consideration of the correlations between…

Methodology · Statistics 2024-04-16 Phuc H. Nguyen , Stephanie M. Engel , Amy H. Herring

In partial label learning (PLL), each training sample is associated with a set of candidate labels, among which only one is valid. The core of PLL is to disambiguate the candidate labels to get the ground-truth one. In disambiguation, the…

Machine Learning · Computer Science 2023-12-19 Yuheng Jia , Chongjie Si , Min-ling Zhang

Real-life data are often non-IID due to complex distributions and interactions, and the sensitivity to the distribution of samples can differ among learning models. Accordingly, a key question for any supervised or unsupervised model is…

Machine Learning · Computer Science 2023-10-03 Zhilin Zhao , Longbing Cao

As in many other areas of science, systems biology makes extensive use of statistical association and significance estimates in contingency tables, a type of categorical data analysis known in this field as enrichment (also…

Quantitative Methods · Quantitative Biology 2011-11-10 Ricardo Vêncio , Ilya Shmulevich

This article introduces the R package hermiter which facilitates estimation of univariate and bivariate probability density functions and cumulative distribution functions along with full quantile functions (univariate) and nonparametric…

Computation · Statistics 2023-07-04 Michael Stephanou , Melvin Varughese

High-throughput pooled resequencing offers significant potential for whole genome population sequencing. However, its main drawback is the loss of haplotype information. In order to regain some of this information, we present LDx, a…

Genomics · Quantitative Biology 2015-06-11 Alison F. Feder , Dmitri A. Petrov , Alan O. Bergland

Renormalization group (RG) methods are emerging as tools in biology and computer science to support the search for simplifying structure in distributions over high-dimensional spaces. We show that mixture models can be thought of as having…

Statistical Mechanics · Physics 2024-02-09 Adam G. Kline , Stephanie E. Palmer

Detecting similar code fragments, usually referred to as code clones, is an important task. In particular, code clone detection can have significant uses in the context of vulnerability discovery, refactoring and plagiarism detection.…

Software Engineering · Computer Science 2021-05-26 Hongfa Xue , Yongsheng Mei , Kailash Gogineni , Guru Venkataramani , Tian Lan

NeXML is a powerful and extensible exchange standard recently proposed to better meet the expanding needs for phylogenetic data and metadata sharing. Here we present the RNeXML package, which provides users of the R programming language…

Quantitative Methods · Quantitative Biology 2015-06-10 Carl Boettiger , Scott Chamberlain , Rutger Vos , Hilmar Lapp

Gene-based testing is a commonly employed strategy in many genetic association studies. Gene-trait associations can be complex due to underlying population heterogeneity, gene-environment interactions, and various other reasons. Existing…

Methodology · Statistics 2020-12-15 Tianying Wang , Iuliana Ionita-Laza , Ying Wei

This paper proposes a novel pooling-based VGG-Lite model in order to mitigate class imbalance issues in Chest X-Ray (CXR) datasets. Automatic Pneumonia detection from CXR images by deep learning model has emerged as a prominent and dynamic…

Image and Video Processing · Electrical Eng. & Systems 2025-04-11 Santanu Roy , Ashvath Suresh , Palak Sahu , Tulika Rudra Gupta

Causal analyses for observational studies are often complicated by covariate imbalances among treatment groups, and matching methodologies alleviate this complication by finding subsets of treatment groups that exhibit covariate balance. It…

Methodology · Statistics 2021-04-26 Zach Branson

We extend a general approach to evaluating identification risk of synthesized variables in partially synthetic data. For multiple continuous synthesized variables, we introduce the use of a radius $r$ in the construction of identification…

Methodology · Statistics 2021-04-07 Ryan Hornby , Jingchen Hu

Convergent evolution provides a useful framework for testing whether independent origins of similar traits share common genetic mechanisms. Evolutionary Sparse Learning with Paired Species Contrast (ESL-PSC) is an approach to identify genes…

Populations and Evolution · Quantitative Biology 2026-05-28 John B. Allard , Sudhir Kumar

In \textit{computer-based testing} it has become standard to collect response accuracy (RA) and response times (RTs) for each test item. IRT models are used to measure a latent variable (e.g., ability, intelligence) using the RA…

Methodology · Statistics 2021-06-21 Jean-Paul Fox , Konrad Klotzke , Ahmet Salih Simsek
‹ Prev 1 3 4 5 6 7 10 Next ›