Related papers: Information Clustering and Pathogen Evolution
Data-stream clustering is an ever-expanding subdomain of knowledge extraction. Most of the past and present research effort aims at efficient scaling up for the huge data repositories. Our approach focuses on qualitative improvement, mainly…
Identifying viral pathogens and characterizing their transmission is essential to developing effective public health measures in response to a pandemic. Phylogenetics, though currently the most popular tool used to characterize the likely…
Despite recent molecular technique improvements, biological knowledge remains incomplete. Reasoning on living systems hence implies to integrate heterogeneous and partial informations. Although current investigations successfully focus on…
During the COVID pandemic, periods of exponential growth of the disease have been mitigated by containment measures that in different occasions have resulted in a power-law growth of the number of cases. The first observation of such…
Epidemiology aims at identifying subpopulations of cohort participants that share common characteristics (e.g. alcohol consumption) to explain risk factors of diseases in cohort study data. These data contain information about the…
Standard approaches to tackle high-dimensional supervised classification problem often include variable selection and dimension reduction procedures. The novel methodology proposed in this paper combines clustering of variables and feature…
There is a keen interest in characterizing variation in the microbiome across cancer patients, given increasing evidence of its important role in determining treatment outcomes. Here our goal is to discover subgroups of patients with…
The dynamics of contact networks and epidemics of infectious diseases often occur on comparable time scales. Ignoring one of these time scales may provide an incomplete understanding of the population dynamics of the infection process. We…
Clustering is commonly performed as an initial analysis step for uncovering structure in 'omics datasets, e.g. to discover molecular subtypes of disease. The high-throughput, high-dimensional nature of these datasets means that they provide…
For a reliable prediction of an epidemic or information spreading pattern in complex systems, well-defined measures are essential. In the susceptible-infected model on heterogeneous networks, the cluster of infected nodes in the…
The spread of a contagious disease clearly depends on when infected individuals come into contact with susceptible ones. Such effects, however, have remained largely unexplored in the study of epidemic outbreaks. In particular, it remains…
We begin by reviewing some probabilistic results about the Dirichlet Process and its close relatives, focussing on their implications for statistical modelling and analysis. We then introduce a class of simple mixture models in which…
Outbreaks are complex multi-scale processes that are impacted not only by cellular dynamics and the ability of pathogens to effectively reproduce and spread, but also by population-level dynamics and the effectiveness of mitigation…
The emergence and spread of deadly pandemics has repeatedly occurred throughout history, causing widespread infections and loss of life. The rapid spread of pandemics have made governments across the world adopt a range of actions,…
To understand the contact patterns of a population -- who is in contact with whom, and when the contacts happen -- is crucial for modeling outbreaks of infectious disease. Traditional theoretical epidemiology assumes that any individual can…
Irregularly sampled time series data are common in a variety of fields. Many typical methods for drawing insight from data fail in this case. Here we attempt to generalize methods for clustering trajectories to irregularly and sparsely…
Early and accurate detection of anomalies in time series data is critical, given the significant risks associated with false or missed detections. While MLP-based mixer models have shown promise in time series analysis, they lack a…
Recent work on overfitting Bayesian mixtures of distributions offers a powerful framework for clustering multivariate data using a latent Gaussian model which resembles the factor analysis model. The flexibility provided by overfitting…
Clustering is a technique for the analysis of datasets obtained by empirical studies in several disciplines with a major application for biomedical research. Essentially, clustering algorithms are executed by machines aiming at finding…
The problem of organizing data that evolves over time into clusters is encountered in a number of practical settings. We introduce evolutionary subspace clustering, a method whose objective is to cluster a collection of evolving data points…