基因组学
In this paper, we study the equilibrium behavior of Eigen's quasispecies equations for an arbitrary gene network. We consider a genome consisting of $ N $ genes, so that each gene sequence $ \sigma $ may be written as $ \sigma = \sigma_1…
One of the major successes in computational biology has been the unification, using the graphical model formalism, of a multitude of algorithms for annotating and comparing biological sequences. Graphical models that have been applied…
We introduce a random bit-string model of post-transcriptional genetic regulation based on sequence matching. The model spontaneously yields a scale free network with power law scaling with $ \gamma=-1$ and also exhibits log-periodic…
We study global fluctuations of the guanine and cytosine base content (GC%) in mouse genomic DNA using spectral analyses. Power spectra S(f) of GC% fluctuations in all nineteen autosomal and two sex chromosomes are observed to have the…
We have simulated the evolution of sexually reproducing populations composed of individuals represented by diploid genomes. A series of eight bits formed an allele occupying one of 128 loci of one haploid genome (chromosome). The…
Individual expression profiles from EBV transformed cell lines are an emerging resource for genomic investigation. In this study we characterize the effects of age, sex, and genetic variation on gene expression by surveying public datasets…
Multiple genome alignment remains a challenging problem. Effects of recombination including rearrangement, segmental duplication, gain, and loss can create a mosaic pattern of homology even among closely related organisms. We describe a…
Starting with the results of Li et al. in 1992 there is valuable interest in finding long range correlations in dna sequences since it raises questions about the role of introns and intron-containing genes. In the present paper we studied…
At the onset of X Chromosomes Inactivation, the vital process whereby female mammal cells equalize X products with respect to males, the X chromosomes are colocalized along their Xic (X-Inactivation Center) regions. The mechanism inducing…
Summary: GeneSupport implements a genome-scale algorithm: Maximum Gene-Support Tree to estimate species tree from gene trees based on multilocus sequences. It provides a new option for multiple genes to infer species tree. It is…
The search for sample-variable associations is an important problem in the exploratory analysis of high dimensional data. Biclustering methods search for sample-variable associations in the form of distinguished submatrices of the data…
The presence of clusters of rare codons is known to negatively impact the efficiency and accuracy of protein production. In this paper, we demonstrate a statistical method of identifying such clusters in the coding sequence of a gene. Using…
We have presented the basic knowledge on the structure of molecules coding the genetic information, mechanisms of transfer of this information from DNA to proteins and phenomena connected with replication of DNA. In particular, we have…
Motivation:Microarray experiments result in large scale data sets that require extensive mining and refining to extract useful information. We demonstrate the usefulness of (nonmetric) multidimensional scaling (MDS) method in analyzing a…
Isochores are long genome segments relatively homogeneous in G+C. A heuristic algorithm based on entropic segmentation has been developed by our group, and a web server implementing all the required components is available. However, a…
Proteomics can be defined as the large-scale analysis of proteins. Due to the complexity of biological systems, it is required to concatenate various separation techniques prior to mass spectrometry. These techniques, dealing with proteins…
In recent times whole-genome gene expression analysis has turned out to be a highly important tool to study the coordinated function of a very large number of genes within their corresponding cellular environment, especially in relation to…
The explanation is that in aortic tissue (both diseased and nondiseased) a BAK1 pseudogene is expressed; while in the matching blood samples the actual BAK1 gene is expressed. This explanation was reached after we realized that BAK1 has two…
Detection of rare variants by resequencing is important for the identification of individuals carrying disease variants. Rapid sequencing by new technologies enables low-cost resequencing of target regions, although it is still prohibitive…
We investigated the error-minimization properties of putative primordial codes that consisted of 16 supercodons, with the third base being completely redundant, using a previously derived cost function and the error minimization percentage…