English
Related papers

Related papers: Small ancestry informative marker panels for compl…

200 papers

With the routine collection of massive-dimensional predictors in many application areas, screening methods that rapidly identify a small subset of promising predictors have become commonplace. We propose a new MOdular Bayes Screening (MOBS)…

Methodology · Statistics 2017-03-30 Yuhan Chen , David B. Dunson

Local community detection consists of finding a group of nodes closely related to the seeds, a small set of nodes of interest. Such group of nodes are densely connected or have a high probability of being connected internally than their…

Social and Information Networks · Computer Science 2020-05-11 Dany Kamuhanda , Meng Wang , Kun He

We propose a resampling-based fast variable selection technique for detecting relevant single nucleotide polymorphisms (SNP) in a multi-marker mixed effect model. Due to computational complexity, current practice primarily involves testing…

Applications · Statistics 2025-04-30 Subhabrata Majumdar , Saonli Basu , Matt McGue , Snigdhansu Chatterjee

Gene panel selection aims to identify the most informative genomic biomarkers in label-free genomic datasets. Traditional approaches, which rely on domain expertise, embedded machine learning models, or heuristic-based iterative…

Genomics · Quantitative Biology 2025-09-12 Meng Xiao , Weiliang Zhang , Xiaohan Huang , Hengshu Zhu , Min Wu , Xiaoli Li , Yuanchun Zhou

Standard approaches to analysing data in genome-wide association studies (GWAS) ignore any potential functional relationships between genetic markers. In contrast gene pathways analysis uses prior information on functional structure within…

Methodology · Statistics 2013-02-26 M. Silver , P. Chen , L. Ruoying , C. Y. Cheng , T. Y. Wong , E. Tai , Y. Y. Teo , G. Montana

For the vast majority of genome wide association studies (GWAS) published so far, statistical analysis was performed by testing markers individually. In this article we present some elementary statistical considerations which clearly show…

Applications · Statistics 2010-10-04 Florian Frommlet , Felix Ruhaltinger , Piotr Twarog , Malgorzata Bogdan

We study modeling and identification of stationary processes with a spectral density matrix of low rank. Equivalently, we consider processes having an innovation of reduced dimension for which Prediction Error Methods (PEM) algorithms are…

Systems and Control · Electrical Eng. & Systems 2023-01-18 Wenqi Cao , Giorgio Picci , Anders Lindquist

Large-scale modern data often involves estimation and testing for high-dimensional unknown parameters. It is desirable to identify the sparse signals, ``the needles in the haystack'', with accuracy and false discovery control. However, the…

Machine Learning · Computer Science 2021-11-08 Junhui Cai , Xu Han , Ya'acov Ritov , Linda Zhao

Sequencing costs currently prohibit the application of single-cell mRNA-seq to many biological and clinical analyses. Targeted single-cell mRNA-sequencing reduces sequencing costs by profiling reduced gene sets that capture biological…

Genomics · Quantitative Biology 2022-02-15 Xiaoqiao Chen , Sisi Chen , Matt Thomson

This paper describes a novel system for automatic classification of images obtained from Anti-Nuclear Antibody (ANA) pathology tests on Human Epithelial type 2 (HEp-2) cells using the Indirect Immunofluorescence (IIF) protocol. The IIF…

Cell Behavior · Quantitative Biology 2014-04-15 Arnold Wiliem , Conrad Sanderson , Yongkang Wong , Peter Hobson , Rodney F. Minchin , Brian C. Lovell

Weak signal identification and inference are very important in the area of penalized model selection, yet they are under-developed and not well-studied. Existing inference procedures for penalized estimators are mainly focused on strong…

Methodology · Statistics 2016-11-16 Peibei Shi , Annie Qu

In this paper a novel biclustering algorithm based on artificial intelligence (AI) is introduced. The method called EBIC aims to detect biologically meaningful, order-preserving patterns in complex data. The proposed algorithm is probably…

Machine Learning · Computer Science 2018-07-27 Patryk Orzechowski , Moshe Sipper , Xiuzhen Huang , Jason H. Moore

Species-sampling problems (SSPs) refer to a vast class of statistical problems calling for the estimation of (discrete) functionals of the unknown species composition of an unobservable population. A common feature of SSPs is their…

Methodology · Statistics 2024-02-13 Cecilia Balocchi , Federico Camerlenghi , Stefano Favaro

Demographic models built from genetic data play important roles in illuminating prehistorical events and serving as null models in genome scans for selection. We introduce an inference method based on the joint frequency spectrum of genetic…

Populations and Evolution · Quantitative Biology 2010-05-10 Ryan N. Gutenkunst , Ryan D. Hernandez , Scott H. Williamson , Carlos D. Bustamante

The proposed feature selection method builds a histogram of the most stable features from random subsets of a training set and ranks the features based on a classifier based cross-validation. This approach reduces the instability of…

Artificial Intelligence · Computer Science 2012-02-07 Alex Pappachen James , Akshay Maan

Motivation: Recent advances in technology for brain imaging and high-throughput genotyping have motivated studies examining the influence of genetic variation on brain structure. Wang et al. (Bioinformatics, 2012) have developed an approach…

Methodology · Statistics 2016-10-18 Keelin Greenlaw , Elena Szefer , Jinko Graham , Mary Lesperance , Farouk S. Nathoo

Traditionally, signal classification is a process in which previous knowledge of the signals is needed. Human experts decide which features are extracted from the signals, and used as inputs to the classification system. This requirement…

Neural and Evolutionary Computing · Computer Science 2019-04-11 Daniel Rivero , Enrique Fernandez-Blanco , Julian Dorado , Alejandro Pazos

Pedigree data contain family history information that is used to analyze hereditary diseases. These clinical data sets may contain duplicate records due to the same family visiting a clinic multiple times or a clinician entering multiple…

Applications · Statistics 2021-08-20 Theodore Huang , Matthew Ploenzke , Danielle Braun

Large-scale biobanks are being collected around the world in efforts to better understand human health and risk factors for disease. They often survey hundreds of thousands of individuals, combining questionnaires with clinical, genetic,…

Quantitative Methods · Quantitative Biology 2019-03-19 Qifan Yang , Gennady V. Roshchupkin , Wiro J. Niessen , Sarah E. Medland , Alyssa H. Zhu , Paul M. Thompson , Neda Jahanshad

Background: In recent years, researchers have made significant strides in understanding the heterogeneity of breast cancer and its various subtypes. However, the wealth of genomic and proteomic data available today necessitates efficient…

Quantitative Methods · Quantitative Biology 2024-03-04 Leandro Y. S. Okimoto , Rayol Mendonca-Neto , Fabíola G. Nakamura , Eduardo F. Nakamura , David Fenyö , Claudio T. Silva