中文
相关论文

相关论文: Identifying statistical dependence in genomic sequ…

200 篇论文

This paper provides partial identification of various binary choice models with misreported dependent variables. We propose two distinct approaches by exploiting different instrumental variables respectively. In the first approach, the…

计量经济学 · 经济学 2024-01-31 Orville Mondal , Rui Wang

Meta-analysis of multiple genome-wide association studies (GWAS) is effective for detecting single or multi marker associations with complex traits. We develop a flexible procedure ("STAMP") based on mixture models to perform region based…

统计方法学 · 统计学 2018-01-01 Andriy Derkach , Ruth M. Pfeiffer

Recent advances in molecular biology allow the quantification of the transcriptome and scoring transcripts as differentially or equally expressed between two biological conditions. Although these two tasks are closely linked, the available…

统计方法学 · 统计学 2017-02-08 Panagiotis Papastamoulis , Magnus Rattray

An informative sampling design leads to the selection of units whose inclusion probabilities are correlated with the response variable of interest. Model inference performed on the resulting observed sample will be biased for the population…

统计方法学 · 统计学 2018-06-29 Matthew R. Williams , Terrance D. Savitsky

Motivation: Modelling methods that find structure in data are necessary with the current large volumes of genomic data, and there have been various efforts to find subsets of genes exhibiting consistent patterns over subsets of treatments.…

机器学习 · 计算机科学 2016-09-15 Kerstin Bunte , Eemeli Leppäaho , Inka Saarinen , Samuel Kaski

Recently, much attention has been given to understanding recombination events along a chromosome in a variety of field. For instance, many population genetics problems are limited by the inaccuracy of inferred evolutionary histories of…

定量方法 · 定量生物学 2017-10-31 Jacqueline Kane , Joseph Rusinko , Katherine Thompson

Learning the differential statistical dependency network between two contexts is essential for many real-life applications, mostly in the high dimensional low sample regime. In this paper, we propose a novel differential network estimator…

机器学习 · 计算机科学 2022-04-25 Arshdeep Sekhon , Zhe Wang , Yanjun Qi

A standard approach for assessing the performance of partition models is to create synthetic data sets with a prespecified clustering structure, and assess how well the model reveals this structure. A common format is that subjects are…

统计方法学 · 统计学 2025-07-08 Michail Papathomas

Generative models are invaluable in many fields of science because of their ability to capture high-dimensional and complicated distributions, such as photo-realistic images, protein structures, and connectomes. How do we evaluate the…

Understanding how stochastic gene expression is regulated in biological systems using snapshots of single-cell transcripts requires state-of-the-art methods of computational analysis and statistical inference. A Bayesian approach to…

定量方法 · 定量生物学 2018-12-10 Yen Ting Lin , Nicolas E. Buchler

Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait.…

机器学习 · 统计学 2012-05-31 Chamont Wang , Jana Gevertz , Chaur-Chin Chen , Leonardo Auslender

This paper introduces a nonparametric copula-based index for detecting the strength and monotonicity structure of linear and nonlinear statistical dependence between pairs of random variables or stochastic signals. Our index, termed Copula…

机器学习 · 统计学 2020-02-25 Kiran Karra , Lamine Mili

The paper presents a new copula based method for measuring dependence between random variables. Our approach extends the Maximum Mean Discrepancy to the copula of the joint distribution. We prove that this approach has several advantageous…

机器学习 · 计算机科学 2019-08-15 Barnabas Poczos , Zoubin Ghahramani , Jeff Schneider

Statistical analysis of evolutionary-related protein sequences provides insights about their structure, function, and history. We show that Restricted Boltzmann Machines (RBM), designed to learn complex high-dimensional data and their…

定量方法 · 定量生物学 2019-02-28 Jérôme Tubiana , Simona Cocco , Rémi Monasson

We develop a model-based methodology for integrating gene-set information with an experimentally-derived gene list. The methodology uses a previously reported sampling model, but takes advantage of natural constraints in the…

统计方法学 · 统计学 2015-06-02 Zhishi Wang , Qiuling He , Bret Larget , Michael A. Newton

Understanding how genetic variants influence cellular-level processes is an important step towards understanding how they influence important organismal-level traits, or "phenotypes", including human disease susceptibility. To this end…

统计方法学 · 统计学 2013-07-30 Heejung Shim , Matthew Stephens

Deciphering complex gene-gene interactions remains challenging in transcriptomics as traditional methods often miss higher-order and nonlinear dependencies. This study introduces a quantum-inspired framework leveraging tensor networks (TNs)…

The detection of similarities between long DNA and protein sequences is studied using concepts of statistical physics. It is shown that mutual similarities can be detected by sequence alignment methods only if their amount exceeds a…

凝聚态物理 · 物理学 2009-10-28 Terence Hwa , Michael Lassig

Assessing the correctness of genome assemblies is an important step in any genome project. Several methods exist, but most are computationally intensive and, in some cases, inappropriate. Here I present baa.pl, a fast and easy-to-use…

基因组学 · 定量生物学 2014-02-10 Joseph F. Ryan

Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the…

机器学习 · 统计学 2016-01-21 Simone Romano , Nguyen Xuan Vinh , James Bailey , Karin Verspoor