中文
相关论文

相关论文: A solution for the rare type match problem when us…

200 篇论文

The likelihood ratio (LR) is a commonly used measure for determining the strength of forensic match evidence. When a forensic expert determines a high LR for DNA found at a crime scene matching the DNA profile of a suspect they typically…

应用统计 · 统计学 2020-09-21 Norman Fenton , Allan Jamieson , Sara Gomes , Martin Neil

Gene-gene interactions are often regarded as playing significant roles in influencing variabilities of complex traits. Although much research has been devoted to this area, to date a comprehensive statistical model that addresses the…

应用统计 · 统计学 2018-04-18 Durba Bhattacharya , Sourabh Bhattacharya

Missing genotypes can affect the efficacy of machine learning approaches to identify the risk genetic variants of common diseases and traits. The problem occurs when genotypic data are collected from different experiments with different DNA…

Density Ratio Estimation (DRE) is an important machine learning technique with many downstream applications. We consider the challenge of DRE with missing not at random (MNAR) data. In this setting, we show that using standard DRE methods…

机器学习 · 统计学 2023-02-22 Josh Givens , Song Liu , Henry W J Reeve

Insights into complex, high-dimensional data can be obtained by discovering features of the data that match or do not match a model of interest. To formalize this task, we introduce the "data selection" problem: finding a lower-dimensional…

统计方法学 · 统计学 2021-09-10 Eli N. Weinstein , Jeffrey W. Miller

In computational biology, gene expression datasets are characterized by very few individual samples compared to a large number of measurements per sample. Thus, it is appealing to merge these datasets in order to increase the number of…

统计方法学 · 统计学 2011-08-18 Meili Baragatti

Deep learning-based classification of rare anemia disorders is challenged by the lack of training data and instance-level annotations. Multiple Instance Learning (MIL) has shown to be an effective solution, yet it suffers from low accuracy…

机器学习 · 计算机科学 2022-07-06 Salome Kazeminia , Ario Sadafi , Asya Makhro , Anna Bogdanova , Shadi Albarqouni , Carsten Marr

Next-generation sequencing technologies now constitute a method of choice to measure gene expression. Data to analyze are read counts, commonly modeled using Negative Binomial distributions. A relevant issue associated with this…

统计方法学 · 统计学 2014-11-10 Elisabetta Bonafede , Franck Picard , Stéphane Robin , Cinzia Viroli

In this paper we propose a bayesian approach for near-duplicate image detection, and investigate how different probabilistic models affect the performance obtained. The task of identifying an image whose metadata are missing is often…

计算机视觉与模式识别 · 计算机科学 2021-08-23 Lucas Moutinho Bueno , Eduardo Valle , Ricardo da Silva Torres

In immunological studies, the characterization of small, functionally distinct cell subsets from blood and tissue is crucial to decipher system level biological changes. An increasing number of studies rely on assays that provide…

Single individual haplotyping is an NP-hard problem that emerges when attempting to reconstruct an organism's inherited genetic variations using data typically generated by high-throughput DNA sequencing platforms. Genomes of diploid…

机器学习 · 计算机科学 2019-09-04 Somsubhra Barik , Haris Vikalo

Mixture interpretation is a central challenge in forensic science, where evidence often contains contributions from multiple sources. In the context of DNA analysis, biological samples recovered from crime scenes may include genetic…

统计方法学 · 统计学 2025-05-05 Taylor Petty , Jan Hannig , Hari Iyer

Next-generation sequencing technologies provide a revolutionary tool for generating gene expression data. Starting with a fixed RNA sample, they construct a library of millions of differentially abundant short sequence tags or "reads",…

定量方法 · 定量生物学 2014-05-13 Dimitrios V. Vavoulis , Julian Gough

In this paper, we propose a new practical association rule mining algorithm for anomaly detection in Intrusion Detection System (IDS). First, with a view of anomaly cases being relatively rarely occurred in network packet database, we…

密码学与安全 · 计算机科学 2016-10-17 Hyeok Kong , Cholyong Jong , Unhyok Ryang

We expand Mendelian Randomization (MR) methodology to deal with randomly missing data on either the exposure or the outcome variable, and furthermore with data from nonindependent individuals (eg components of a family). Our method rests on…

Following the discovery of the brightest high-energy neutrino sources in the sky, the further detection of fainter sources is more challenging. A natural solution is to combine fainter source candidates, and instead of individual…

高能天体物理现象 · 物理学 2025-06-03 I. Bartos , M. Ackermann , M. Kowalski

Discovering statistically significant patterns from databases is an important challenging problem. The main obstacle of this problem is in the difficulty of taking into account the selection bias, i.e., the bias arising from the fact that…

机器学习 · 统计学 2016-03-10 Shinya Suzumura , Kazuya Nakagawa , Mahito Sugiyama , Koji Tsuda , Ichiro Takeuchi

Bi-clustering is a useful approach in analyzing biological data when observations come from heterogeneous groups and have a large number of features. We outline a general Bayesian approach in tackling bi-clustering problems in moderate to…

应用统计 · 统计学 2021-02-11 Han Yan , Jiexing Wu , Yang Li , Jun S. Liu

A maximum likelihood based model selection of discrete Bayesian networks is considered. The model selection is performed through scoring function $S$, which, for a given network $G$ and $n$-sample $D_n$, is defined to be the maximum…

统计理论 · 数学 2013-04-18 Nikolay H. Balov

DNA storage technology offers new possibilities for addressing massive data storage due to its high storage density, long-term preservation, low maintenance cost, and compact size. To improve the reliability of stored information, base…

机器学习 · 计算机科学 2024-09-24 Bowen Liu , Jiankun Li