English
Related papers

Related papers: Correlation-Compressed Direct Coupling Analysis

200 papers

Finding interdependency relations between (possibly multivariate) time series provides valuable knowledge about the processes that generate the signals. Information theory sets a natural framework for non-parametric measures of several…

Information Theory · Computer Science 2016-02-09 German Gomez-Herrero , Wei Wu , Kalle Rutanen , Miguel C. Soriano , Gordon Pipa , Raul Vicente

The inverse Potts problem to infer a Boltzmann distribution for homologous protein sequences from their single-site and pairwise amino acid frequencies recently attracts a great deal of attention in the studies of protein structure and…

Populations and Evolution · Quantitative Biology 2024-07-23 Sanzo Miyazawa

Discovering causal models from observational and interventional data is an important first step preceding what-if analysis or counterfactual reasoning. As has been shown before, the direction of pairwise causal relations can, under certain…

Machine Learning · Computer Science 2017-01-04 Karamjit Singh , Garima Gupta , Lovekesh Vig , Gautam Shroff , Puneet Agarwal

Deep learning models that leverage large datasets are often the state of the art for modelling molecular properties. When the datasets are smaller (< 2000 molecules), it is not clear that deep learning approaches are the right modelling…

Computational Engineering, Finance, and Science · Computer Science 2022-12-07 Gary Tom , Riley J. Hickman , Aniket Zinzuwadia , Afshan Mohajeri , Benjamin Sanchez-Lengeling , Alan Aspuru-Guzik

High-dimensional compositional data arise naturally in many applications such as metagenomic data analysis. The observed data lie in a high-dimensional simplex, and conventional statistical methods often fail to produce sensible results due…

Methodology · Statistics 2016-01-19 Yuanpei Cao , Wei Lin , Hongzhe Li

Direct Coupling Analysis (DCA) is a now widely used method to leverage statistical information from many similar biological systems to draw meaningful conclusions on each system separately. DCA has been applied with great success to…

Populations and Evolution · Quantitative Biology 2018-08-13 Chen-Yi Gao , Fabio Cecconi , Angelo Vulpiani , Hai-Jun Zhou , Erik Aurell

The No Unmeasured Confounding Assumption is widely used to identify causal effects in observational studies. Recent work on proximal inference has provided alternative identification results that succeed even in the presence of unobserved…

Machine Learning · Statistics 2022-10-17 Benjamin Kompa , David R. Bellamy , Thomas Kolokotrones , James M. Robins , Andrew L. Beam

Preferential sampling is a common feature in geostatistics and occurs when the locations to be sampled are chosen based on information about the phenomena under study. In this case, point pattern models are commonly used as the probability…

Methodology · Statistics 2022-10-27 Douglas Mateus da Silva , Dani Gamerman

The Conway-Maxwell-Poisson (CMP) or COM-Poison regression is a popular model for count data due to its ability to capture both under dispersion and over dispersion. However, CMP regression is limited when dealing with complex nonlinear…

Methodology · Statistics 2020-04-27 Suneel Babu Chatla , Galit Shmueli

Predicting the physical interaction of proteins is a cornerstone problem in computational biology. New classes of learning-based algorithms are actively being developed, and are typically trained end-to-end on protein complex structures…

Biomolecules · Quantitative Biology 2022-12-08 Siddharth Bhadra-Lobo , Georgy Derevyanko , Guillaume Lamoureux

Identifying causal relationships from observational time series data is a key problem in disciplines such as climate science or neuroscience, where experiments are often not possible. Data-driven causal inference is challenging since…

Methodology · Statistics 2019-12-03 Jakob Runge , Peer Nowack , Marlene Kretschmer , Seth Flaxman , Dino Sejdinovic

As predictive algorithms grow in popularity, using the same dataset to both train and test a new model has become routine across research, policy, and industry. Sample-splitting attains valid inference on model properties by using separate…

Econometrics · Economics 2025-11-27 Bruno Fava

Detection with high dimensional multimodal data is a challenging problem when there are complex inter- and intra- modal dependencies. While several approaches have been proposed for dependent data fusion (e.g., based on copula theory),…

Applications · Statistics 2018-02-14 Thakshila Wimalajeewa , Pramod K. Varshney

Predicting compound-protein affinity is critical for accelerating drug discovery. Recent progress made by machine learning focuses on accuracy but leaves much to be desired for interpretability. Through molecular contacts underlying…

Biomolecules · Quantitative Biology 2020-01-01 Mostafa Karimi , Di Wu , Zhangyang Wang , Yang Shen

It is common to conduct causal inference in matched observational studies by proceeding as though treatment assignments within matched sets are assigned uniformly at random and using this distribution as the basis for inference. This…

Methodology · Statistics 2023-11-14 Samuel D. Pimentel , Yaxuan Huang

Background: Selecting feature genes to predict phenotypes is one of the typical tasks in analyzing genomics data. Though many general-purpose algorithms were developed for prediction, dealing with highly correlated genes in the prediction…

Applications · Statistics 2022-04-11 Li Xing , Songwan Joun , Kurt Mackay , Mary Lesperance , Xuekui Zhang

There has been a lot of work fitting Ising models to multivariate binary data in order to understand the conditional dependency relationships between the variables. However, additional covariates are frequently recorded together with the…

Machine Learning · Statistics 2012-09-28 Jie Cheng , Elizaveta Levina , Pei Wang , Ji Zhu

Accurate interpolation of seismic data is crucial for improving the quality of imaging and interpretation. In recent years, deep learning models such as U-Net and generative adversarial networks have been widely applied to seismic data…

The standard methods for detecting differential gene expression are mostly designed for analyzing a single gene expression experiment. When data from multiple related gene expression studies are available, separately analyzing each study is…

Methodology · Statistics 2013-11-07 Yingying Wei , Hongkai Ji

The current capacity of computers makes it possible to perform simulations of small systems with portable, explicit-solvent potentials achieving high degree of accuracy. However, simplified models must be employed to exploit the behaviour…

Biomolecules · Quantitative Biology 2015-06-18 R. Capelli , C. Paissoni , P. Sormanni , G. Tiana