Related papers: Finding the Right Bricks for Molecular Lego: A Dat…
We develop a new methodology for model-based clustering. Optimizing the log-likelihood provides a principled statistical framework for clustering, with solutions found via the EM algorithm. However, because the log-likelihood is nonconvex,…
Predicting charge transport in organic molecular crystals is notoriously challenging. Carrier mobility calculations in organic semiconductors are dominated by quantum chemistry methods based on charge hopping, which are laborious and only…
Advancements of both computational and experimental tools have recently led to significant progress in the development of new advanced and functional materials, paralleled by a quick growth of the overall amount of data and information on…
Charge transport in two zinc metal-organic frameworks (MOFs) has been investigated using periodic semiempirical molecular orbital calculations with the AM1* Hamiltonian. Restricted Hartree-Fock calculations underestimate the band gap…
In order to efficiently explore the chemical space of all possible small molecules, a common approach is to compress the dimension of the system to facilitate downstream machine learning tasks. Towards this end, we present a data driven…
Probabilistic clustering models (or equivalently, mixture models) are basic building blocks in countless statistical models and involve latent random variables over discrete spaces. For these models, posterior inference methods can be…
Traditional load analysis is facing challenges with the new electricity usage patterns due to demand response as well as increasing deployment of distributed generations, including photovoltaics (PV), electric vehicles (EV), and energy…
Materials informatics, data-enabled investigation, is a "fourth paradigm" in materials science research after the conventional empirical approach, theoretical science, and computational research. Materials informatics has two essential…
An emerging paradigm in modern electronics is that of CMOS + $\sf X$ requiring the integration of standard CMOS technology with novel materials and technologies denoted by $\sf X$. In this context, a crucial challenge is to develop accurate…
Efficiently retrieving an enormous chemical library to design targeted molecules is crucial for accelerating drug discovery, organic chemistry, and optoelectronic materials. Despite the emergence of generative models to produce novel…
Charge transport through a short DNA oligomer (Dickerson dodecamer) in presence of structural fluctuations is investigated using a hybrid computational methodology based on a combination of quantum mechanical electronic structure…
Clustering is a crucial component of many data mining systems involving the analysis and exploration of various data. Data diversity calls for clustering algorithms to be accurate while providing stable (i.e., deterministic and robust)…
Clustering is one of the major tasks in data mining. In the last few years, Clustering of spatial data has received a lot of research attention. Spatial databases are components of many advanced information systems like geographic…
This research demonstrates that Ising machines can effectively solve optimal elemental configuration searches in crystals, with Au-Cu alloys serving as an example. The energy function is derived using the cluster expansion method in the…
The structure of nanoclusters is complex to describe due to their noncrystallinity, even though bonding and packing constraints limit the local atomic arrangements to only a few types. A computational scheme is presented to extract…
Molecular crystals compose the current state of the art when it comes to organic-based optoelectronic applications. Charge transport is a crucial aspect of their performance. The ability to predict accurate electron mobility is needed in…
Proposed here is a dynamic Monte-Carlo algorithm that is efficient in simulating dense systems of long flexible chain molecules. It expands on the configurational-bias Monte-Carlo method through the simultaneous generation of a large set of…
We describe a robust, fast, and memory-efficient procedure that can cluster millions of structures derived from molecular dynamics simulations. The essence of the method is based on a peak-picking algorithm applied to three- and…
As scientific data repositories and filesystems grow in size and complexity, they become increasingly disorganized. The coupling of massive quantities of data with poor organization makes it challenging for scientists to locate and utilize…
Determining the number of clusters in a dataset is a fundamental issue in data clustering. Many methods have been proposed to solve the problem of selecting the number of clusters, considering it to be a problem with regard to model…