English
Related papers

Related papers: New stopping criteria for segmenting DNA sequences

200 papers

Modeling biological sequences such as DNA, RNA, and proteins is crucial for understanding complex processes like gene regulation and protein synthesis. However, most current models either focus on a single type or treat multiple types of…

Genomics · Quantitative Biology 2024-10-16 Weixi Xiang , Xueting Han , Xiujuan Chai , Jing Bai

Disease subtype identification (clustering) is an important problem in biomedical research. Gene expression profiles are commonly utilized to infer disease subtypes, which often lead to biologically meaningful insights into disease. Despite…

Methodology · Statistics 2016-09-27 Jiehuan Sun , Joshua L. Warren , Hongyu Zhao

Model selection and order selection problems frequently arise in statistical practice. A popular approach to addressing these problems in the frequentist setting involves information criteria based on penalised maxima of log-likelihoods for…

Statistics Theory · Mathematics 2025-10-29 Hien Duy Nguyen , Mayetri Gupta , Jacob Westerhout , TrungTin Nguyen

Real-life statistical samples are often plagued by selection bias, which complicates drawing conclusions about the general population. When learning causal relationships between the variables is of interest, the sample may be assumed to be…

Statistics Theory · Mathematics 2018-11-15 Angelos P. Armen , Robin J. Evans

Symmetry principles play an important role in geometry, and physics, allowing for the reduction of complicated systems to simpler, more comprehensible models that preserve the system's features of interest. Biological systems are often…

Cell Behavior · Quantitative Biology 2025-02-26 Luis A. Álvarez-García , Wolfram Liebermeister , Ian Leifer , Hernán A. Makse

Clustering mixed-type data remains a major challenge in biomedical research to uncover clinically meaningful subgroups within heterogeneous patient populations. Most existing clustering methods impose restrictive assumptions like local…

Applications · Statistics 2026-04-23 Yueting Wang , Shu Wang , Jonathan G. Yabes , Chung-Chou H. Chang

It is shown that metric representation of DNA sequences is one-to-one. By using the metric representation method, suppression of nucleotide strings in the DNA sequences is determined. For a DNA sequence, an optimal string length to display…

Biological Physics · Physics 2007-05-23 Zuo-Bing Wu

Finding out statistically significant words in DNA and protein sequences forms the basis for many genetic studies. By applying the maximal entropy principle, we give one systematic way to study the nonrandom occurrence of words in DNA or…

Biological Physics · Physics 2009-11-06 Rui Hu , Bin Wang

The Bayes factor, the data-based updating factor from prior to posterior odds, is a principled measure of relative evidence for two competing hypotheses. It is naturally suited to sequential data analysis in settings such as clinical trials…

Methodology · Statistics 2026-01-07 Samuel Pawel , Leonhard Held

While there have been numerous sequential algorithms developed to estimate community structure in networks, there is little available guidance and study of what significance level or stopping parameter to use in these sequential testing…

Methodology · Statistics 2022-09-19 Riddhi Pratim Ghosh , Ian Barnett

Many sequential decision settings in healthcare feature funnel structures characterized by a series of stages, such as screenings or evaluations, where the number of patients who advance to each stage progressively decreases and decisions…

Machine Learning · Computer Science 2025-11-25 Shuvom Sadhuka , Sophia Lin , Bonnie Berger , Emma Pierson

Essential protein plays a crucial role in the process of cell life. The identification of essential proteins can not only promote the development of drug target technology, but also contribute to the mechanism of biological evolution. There…

Molecular Networks · Quantitative Biology 2020-05-20 Pengli Lu , JingJuan Yu

This work addresses the problem of segmentation in time series data with respect to a statistical parameter of interest in Bayesian models. It is common to assume that the parameters are distinct within each segment. As such, many Bayesian…

Machine Learning · Computer Science 2017-10-27 Alireza Ahrabian , Shirin Enshaeifar , Clive Cheong-Took , Payam Barnaghi

High-throughput genetic and epigenetic data are often screened for associations with an observed phenotype. For example, one may wish to test hundreds of thousands of genetic variants, or DNA methylation sites, for an association with…

Methodology · Statistics 2017-10-20 Eric F. Lock , David B. Dunson

The Bayesian learning rule is a natural-gradient variational inference method, which not only contains many existing learning algorithms as special cases but also enables the design of new algorithms. Unfortunately, when variational…

Machine Learning · Statistics 2020-10-27 Wu Lin , Mark Schmidt , Mohammad Emtiyaz Khan

This paper introduces a novel nonparametric criterion for determining the appropriate number of clusters, which is derived from the spatial median. The method is constructed to reconcile two competing objectives of cluster analysis: the…

Computation · Statistics 2025-09-26 Hend Gabr , Brian H Willis , Mohammed Baragilly

We introduce a new criterion to determine the order of an autoregressive model fitted to time series data. It has the benefits of the two well-known model selection techniques, the Akaike information criterion and the Bayesian information…

Statistics Theory · Mathematics 2016-08-25 Jie Ding , Vahid Tarokh , Yuhong Yang

The most common gene regulation mechanism is when a transcription factor protein binds to a regulatory sequence to increase or decrease RNA transcription. However, transcription factors face two main challenges when searching for these…

Quantitative Methods · Quantitative Biology 2023-11-21 Lucas Hedström , Ludvig Lizana

A new method to identify all sufficiently long repeating substrings in one or several symbol sequences is proposed. The method is based on a specific gauge applied to symbol sequences that guarantees identification of the repeating…

Genomics · Quantitative Biology 2016-04-07 Sergey Tsarev , Michael Sadovsky

Next-generation sequencing technologies provide a revolutionary tool for generating gene expression data. Starting with a fixed RNA sample, they construct a library of millions of differentially abundant short sequence tags or "reads",…

Quantitative Methods · Quantitative Biology 2014-05-13 Dimitrios V. Vavoulis , Julian Gough
‹ Prev 1 4 5 6 7 8 10 Next ›