English
Related papers

Related papers: Sequence alignment, mutual information, and dissim…

200 papers

Within bioinformatics, the textual alignment of amino acid sequences has long dominated the determination of similarity between proteins, with all that implies for shared structure, function and evolutionary descent. Despite the relative…

Quantitative Methods · Quantitative Biology 2016-02-10 Amit K Chattopadhyay , Diar Nasiev , Darren R Flower

A hybrid evolutionary algorithm with importance sampling method is proposed for multi-dimensional optimization problems in this paper. In order to make use of the information provided in the search process, a set of visited solutions is…

Neural and Evolutionary Computing · Computer Science 2013-08-26 Guanghui Huang , Zhifeng Pan

Single-cell data integration can provide a comprehensive molecular view of cells, and many algorithms have been developed to remove unwanted technical or biological variations and integrate heterogeneous single-cell datasets. Despite their…

Quantitative Methods · Quantitative Biology 2024-03-04 Rong Ma , Eric D. Sun , David Donoho , James Zou

The Earth Mover's Distance (EMD) is a state-of-the art metric for comparing discrete probability distributions, but its high distinguishability comes at a high cost in computational complexity. Even though linear-complexity approximation…

Machine Learning · Computer Science 2019-05-29 Kubilay Atasu , Thomas Mittelholzer

In this study, we introduce a new approach to combine multi-classifiers in an ensemble system. Instead of using numeric membership values encountered in fixed combining rules, we construct interval membership values associated with each…

Machine Learning · Computer Science 2017-03-17 Tien Thanh Nguyen , Xuan Cuong Pham , Alan Wee-Chung Liew , Witold Pedrycz

Since its introduction, the partial information decomposition (PID) has emerged as a powerful, information-theoretic technique useful for studying the structure of (potentially higher-order) interactions in complex systems. Despite its…

Information Theory · Computer Science 2023-12-11 Thomas F. Varley

Maximum mutual information (MMI) is a model selection criterion used for hidden Markov model (HMM) parameter estimation that was developed more than twenty years ago as a discriminative alternative to the maximum likelihood criterion for…

Computation and Language · Computer Science 2010-02-04 Steven Wegmann

Accurate estimation of evolutionary distances between taxa is important for many phylogenetic reconstruction methods. In the case of bacteria, distances can be estimated using a range of different evolutionary models, from single nucleotide…

Populations and Evolution · Quantitative Biology 2017-04-17 Stuart Serdoz , Attila Egri-Nagy , Jeremy Sumner , Barbara R. Holland , Peter D. Jarvis , Mark M. Tanaka , Andrew R. Francis

In this paper, we propose a new time-aware dissimilarity measure that takes into account the temporal dimension. Observations that are close in the description space, but distant in time are considered as dissimilar. We also propose a…

Machine Learning · Computer Science 2016-01-13 Marian-Andrei Rizoiu , Julien Velcin , Stéphane Lallich

The Information Bottleneck (IB) principle offers an information-theoretic framework for analyzing the training process of deep neural networks (DNNs). Its essence lies in tracking the dynamics of two mutual information (MI) values: between…

Machine Learning · Computer Science 2024-05-10 Ivan Butakov , Alexander Tolmachev , Sofia Malanchuk , Anna Neopryatnaya , Alexey Frolov , Kirill Andreev

Many learning algorithms such as kernel machines, nearest neighbors, clustering, or anomaly detection, are based on the concept of 'distance' or 'similarity'. Before similarities are used for training an actual machine learning model, we…

Conditional Mutual Information (CMI) is a measure of conditional dependence between random variables X and Y, given another random variable Z. It can be used to quantify conditional dependence among variables in many data-driven inference…

Machine Learning · Computer Science 2019-06-10 Sudipto Mukherjee , Himanshu Asnani , Sreeram Kannan

Understanding the contribution of individual features in predictive models remains a central goal in interpretable machine learning, and while many model-agnostic methods exist to estimate feature importance, they often fall short in…

Machine Learning · Computer Science 2025-07-08 Ivan Lazic , Chiara Barà , Marta Iovino , Sebastiano Stramaglia , Niksa Jakovljevic , Luca Faes

We point out a limitation of the mutual information neural estimation (MINE) where the network fails to learn at the initial training phase, leading to slow convergence in the number of training iterations. To solve this problem, we propose…

Information Theory · Computer Science 2019-06-03 Chung Chan , Ali Al-Bashabsheh , Hing Pang Huang , Michael Lim , Da Sun Handason Tam , Chao Zhao

Different combinations of input parameters to filament identification algorithms, such as Disperse and FilFinder, produce numerous different output skeletons. The skeletons are a one pixel wide representation of the filamentary structure in…

Instrumentation and Methods for Astrophysics · Physics 2017-05-24 C. -E. Green , M. R. Cunningham , J. R. Dawson , P. A. Jones , G. Novak , L. M. Fissel

We introduce new methods for phylogenetic tree quartet construction by using machine learning to optimize the power of phylogenetic invariants. Phylogenetic invariants are polynomials in the joint probabilities which vanish under a model of…

Populations and Evolution · Quantitative Biology 2007-05-23 Nicholas Eriksson , Yuan Yao

Mutual information (MI) is a fundamental quantity in information theory and machine learning. However, direct estimation of MI is intractable, even if the true joint probability density for the variables of interest is known, as it involves…

Machine Learning · Computer Science 2024-04-29 Rob Brekelmans , Sicong Huang , Marzyeh Ghassemi , Greg Ver Steeg , Roger Grosse , Alireza Makhzani

Recommendation systems for different Document Networks (DN) such as the World Wide Web (WWW) and Digital Libraries, often use distance functions extracted from relationships among documents and keywords. For instance, documents in the WWW…

Information Retrieval · Computer Science 2007-05-23 L. M. Rocha

We introduce the Mutual Information Machine (MIM), a probabilistic auto-encoder for learning joint distributions over observations and latent variables. MIM reflects three design principles: 1) low divergence, to encourage the encoder and…

Machine Learning · Computer Science 2020-02-24 Micha Livne , Kevin Swersky , David J. Fleet

While the linear Pearson correlation coefficient represents a well-established normalized measure to quantify the interrelation of two stochastic variables $X$ and $Y$, it fails for multidimensional variables such as Cartesian coordinates.…

Data Analysis, Statistics and Probability · Physics 2024-09-19 Daniel Nagel , Georg Diez , Gerhard Stock