Related papers: Intra-Chromosomal Potentials from Nucleosomal Posi…
The k-means algorithm is a partitional clustering method. Over 60 years old, it has been successfully used for a variety of problems. The popularity of k-means is in large part a consequence of its simplicity and efficiency. In this paper…
It is known that, by accounting for the multiboson interferences up to a finite order, the output distribution of noisy Boson Sampling, with distinguishability of bosons serving as noise, can be approximately sampled from in a time…
Source localization and spectral estimation are among the most fundamental problems in statistical and array signal processing. Methods which rely on the orthogonality of the signal and noise subspaces, such as Pisarenko's method, MUSIC,…
Due to advances in sensors, growing large and complex medical image data have the ability to visualize the pathological change in the cellular or even the molecular level or anatomical changes in tissues and organs. As a consequence, the…
The structure of nanoclusters is complex to describe due to their noncrystallinity, even though bonding and packing constraints limit the local atomic arrangements to only a few types. A computational scheme is presented to extract…
The deployment of language models brings challenges in generating reliable information, especially when these models are fine-tuned using human preferences. To extract encoded knowledge without (potentially) biased human labels,…
Finding a set of nested partitions of a dataset is useful to uncover relevant structure at different scales, and is often dealt with a data-dependent methodology. In this paper, we introduce a general two-step methodology for model-based…
We present a new clustering algorithm that is based on searching for natural gaps in the components of the lowest energy eigenvectors of the Laplacian of a graph. In comparing the performance of the proposed method with a set of other…
Observations of present and future X-ray telescopes include a large number of serendipidious sources of unknown types. They are a rich source of knowledge about X-ray dominated astronomical objects, their distribution, and their evolution.…
We consider the problem of selecting a subset of alternatives given noisy evaluations of the relative strength of different alternatives. We wish to select a k-subset (for a given k) that provides a maximum likelihood estimate for one of…
Image registration is useful for quantifying morphological changes in longitudinal MR images from prostate cancer patients. This paper describes a development in improving the learning-based registration algorithms, for this challenging…
The modeling of solute chemistry at low-symmetry defects in materials is historically challenging, due to the computation cost required to evaluate thermodynamic properties from first principles. Here, we offer a hybrid multiscale approach…
An extension of the latent class model is presented for clustering categorical data by relaxing the classical "class conditional independence assumption" of variables. This model consists in grouping the variables into inter-independent and…
Algorithms for clustering points in metric spaces is a long-studied area of research. Clustering has seen a multitude of work both theoretically, in understanding the approximation guarantees possible for many objective functions such as…
Most of the existing research in assembly pathway prediction/analysis of virus cap- sids makes the simplifying assumption that the configuration of the intermediate states can be extracted directly from the final configuration of the entire…
Cluster analysis, or clustering, plays a crucial role across numerous scientific and engineering domains. Despite the wealth of clustering methods proposed over the past decades, each method is typically designed for specific scenarios and…
For problems relating to fracture, a consistent embedding of a quantum (QM) domain in its classical (CM) environment requires that the classical system should yield the same structure and elastic properties as the QM domain for states near…
The analysis of cancer genomic data has long suffered "the curse of dimensionality". Sample sizes for most cancer genomic studies are a few hundreds at most while there are tens of thousands of genomic features studied. Various methods have…
We introduce and study a set of training-free methods of information-theoretic and algorithmic complexity nature applied to DNA sequences to identify their potential capabilities to determine nucleosomal binding sites. We test our measures…
We present a new scheme to extract numerically ``optimal'' interatomic potentials from large amounts of data produced by first-principles calculations. The method is based on fitting the potential to ab initio atomic forces of many atomic…