中文
相关论文

相关论文: Gene ranking and biomarker discovery under correla…

200 篇论文

We study algorithms for estimating the statistical leverage scores of rectangular dense or sparse matrices of arbitrary rank. Our approach is based on combining rank revealing methods with compositions of dense and sparse randomized…

数据结构与算法 · 计算机科学 2022-03-08 Aleksandros Sobczyk , Efstratios Gallopoulos

Global constraints and reranking have not been used in cognates detection research to date. We propose methods for using global constraints by performing rescoring of the score matrices produced by state of the art cognates detection…

计算与语言 · 计算机科学 2017-08-22 Michael Bloodgood , Benjamin Strauss

A new algorithm, termed subspace evolution and transfer (SET), is proposed for solving the consistent matrix completion problem. In this setting, one is given a subset of the entries of a low-rank matrix, and asked to find one low-rank…

信息论 · 计算机科学 2010-02-03 Wei Dai , Olgica Milenkovic

Current face recognition systems achieve high progress on several benchmark tests. Despite this progress, recent works showed that these systems are strongly biased against demographic sub-groups. Consequently, an easily integrable solution…

计算机视觉与模式识别 · 计算机科学 2020-11-06 Philipp Terhörst , Jan Niklas Kolf , Naser Damer , Florian Kirchbuchner , Arjan Kuijper

The typical process for classifying and submitting a newly sequenced virus to the NCBI database involves two steps. First, a BLAST search is performed to determine likely family candidates. That is followed by checking the candidate…

基因组学 · 定量生物学 2016-03-22 Troy Hernandez , Jie Yang

In statistics and machine learning, feature selection is the process of picking a subset of relevant attributes for utilizing in a predictive model. Recently, rough set-based feature selection techniques, that employ feature dependency to…

机器学习 · 计算机科学 2020-03-30 Seyedeh Faezeh Farahbakhshian , Milad Taleby Ahvanooey

Estimating time-varying correlation matrices is challenging because existing methods may adapt slowly to structural changes, impose insufficient regularization, or produce diffuse posterior uncertainty. In moderate dimensions, an additional…

统计方法学 · 统计学 2026-05-11 Daniel Andrew Coulson , David S. Matteson , Martin T. Wells

In a clinical trial, the random allocation aims to balance prognostic factors between arms, preventing true confounders. However, residual differences due to chance may introduce near-confounders. Adjusting on prognostic factors is…

统计方法学 · 统计学 2024-11-18 Joe de Keizer , Rémi Lenain , Raphaël Porcher , Sarah Zoha , Arthur Chatton , Yohann Foucher

This paper discusses the problem of identifying differentially expressed groups of genes from a microarray experiment. The groups of genes are externally defined, for example, sets of gene pathways derived from biological databases. Our…

统计理论 · 数学 2009-09-29 Bradley Efron , Robert Tibshirani

Sequential recommendation has increasingly shifted toward generative recommenders that combine sequential patterns with semantic item information. Yet these methods are often evaluated on a small set of widely used benchmarks, raising a key…

In this paper, we study the trace regression when a matrix of parameters B* is estimated via the convex relaxation of a rank-regularized regression or via regularized non-convex optimization. It is known that these estimators satisfy…

机器学习 · 计算机科学 2023-08-31 Nima Hamidi , Mohsen Bayati

Randomized controlled trials (RCTs) are often underpowered to detect treatment heterogeneity in subgroups defined by cross-classifications of multiple covariates, due to sparse sample sizes in some strata. External RCT data can help, but…

统计方法学 · 统计学 2026-04-23 Youqi Yang , Walter Dempsey , Bhramar Mukherjee

Rerandomization enforces covariate balance across treatment groups in the design stage of experiments. Despite its intuitive appeal, its theoretical justification remains unsatisfying because its benefits of improving efficiency for…

统计理论 · 数学 2025-05-05 Xin Lu , Peng Ding

Reconstructing a gene network from high-throughput molecular data is often a challenging task, as the number of parameters to estimate easily is much larger than the sample size. A conventional remedy is to regularize or penalize the model…

Background With microarray technology becoming mature and popular, the selection and use of a small number of relevant genes for accurate classification of samples is a hot topic in the circles of biostatistics and bioinformatics. However,…

统计方法学 · 统计学 2014-03-05 Suyan Tian , Mayte Suárez-Fariñas

The ability to quickly and accurately identify microbial species in a sample, known as metagenomic profiling, is critical across various fields, from healthcare to environmental science. This paper introduces a novel method to profile…

基因组学 · 定量生物学 2025-04-10 Riselda Kodra , Hadjer Benmeziane , Irem Boybat , William Andrew Simon

BACKGROUND: Breast cancer has emerged as one of the most prevalent cancers among women leading to a high mortality rate. Due to the heterogeneous nature of breast cancer, there is a need to identify differentially expressed genes associated…

机器学习 · 计算机科学 2021-11-30 Sheetal Rajpal , Ankit Rajpal , Manoj Agarwal , Naveen Kumar

Reconstruction of gene regulatory networks is the process of identifying gene dependency from gene expression profile through some computation techniques. In our human body, though all cells pose similar genetic material but the activation…

Comparing large covariance matrices has important applications in modern genomics, where scientists are often interested in understanding whether relationships (e.g., dependencies or co-regulations) among a large number of genes vary…

统计方法学 · 统计学 2017-04-04 Jinyuan Chang , Wen Zhou , Wen-Xin Zhou , Lan Wang

This paper introduces a novel revealed-preference approach to ranking colleges and professional schools based on applicants' choices and standardized test scores. Unlike traditional rankings that rely on data supplied by institutions or…

综合经济学 · 经济学 2025-07-17 Federico Echenique , Michael Olabisi