Related papers: Compact atomic descriptors enable accurate predict…
Because of the advance in technologies, modern statistical studies often encounter linear models with the number of explanatory variables much larger than the sample size. Estimation and variable selection in these high-dimensional problems…
The availability of precise and accurate simulation is a limiting factor for interpreting and forecasting data in many fields of science and engineering. Often, one or more distinct simulation software applications are developed, each with…
Hardness is a materials' property with implications in several industrial fields, including oil and gas, manufacturing, and others. However, the relationship between this macroscale property and atomic (i.e., microscale) properties is…
Penalized logistic regression is extremely useful for binary classification with large number of covariates (higher than the sample size), having several real life applications, including genomic disease classification. However, the…
Representation based classification (RC) methods such as sparse RC (SRC) have shown great potential in face recognition in recent years. Most previous RC methods are based on the conventional regression models, such as lasso regression,…
The formulation of descriptors of the local chemical environment, enabling the construction of machine-learning models, is usually obtained by studying the properties of the expansion coefficients of a neighborhood density. In this work, we…
Robust local feature representations are essential for spatial intelligence tasks such as robot navigation and augmented reality. Establishing reliable correspondences requires descriptors that provide both high discriminative power and…
Machine learning (ML) enables the development of interatomic potentials that promise the accuracy of first principles methods while retaining the low cost and parallel efficiency of empirical potentials. While ML potentials traditionally…
Dictionary learning aims at seeking a dictionary under which the training data can be sparsely represented. Methods in the literature typically formulate the dictionary learning problem as an optimization w.r.t. two variables, i.e.,…
3D action recognition was shown to benefit from a covariance representation of the input data (joint 3D positions). A kernel machine feed with such feature is an effective paradigm for 3D action recognition, yielding state-of-the-art…
Computational screening for new and improved catalyst materials relies on accurate and low-cost predictions of key parameters such as adsorption energies. Here, we use recently developed compressed sensing methods to identify descriptors…
This work introduces various approaches to include connected three-body terms in unitary many-body theories, focusing a representative example on the driven similarity renormalization group (DSRG). Starting from the least approximate method…
Predicting accurate protein-ligand binding affinity is important in drug discovery but remains a challenge even with computationally expensive biophysics-based energy scoring methods and state-of-the-art deep learning approaches. Despite…
A quantitative descriptor of local atomic environments is often required for the analysis of atomistic data. Descriptors of the local atomic environment ideally provide physically and chemically intuitive insight. This requires descriptors…
We show how to speed up global optimization of molecular structures using machine learning methods. To represent the molecular structures we introduce the auto-bag feature vector that combines: i) a local feature vector for each atom, ii)…
The choice of structural resolution is a fundamental aspect of protein modelling, determining the balance between descriptive power and interpretability. Although atomistic simulations provide maximal detail, much of this information is…
Deep Click-Through Rate (CTR) prediction models play an important role in modern industrial recommendation scenarios. However, high memory overhead and computational costs limit their deployment in resource-constrained environments.…
Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets…
Statistical learning algorithms are finding more and more applications in science and technology. Atomic-scale modeling is no exception, with machine learning becoming commonplace as a tool to predict energy, forces and properties of…
In this abstract paper, we introduce a new kernel learning method by a nonparametric density estimator. The estimator consists of a group of k-centroids clusterings. Each clustering randomly selects data points with randomly selected…