相关论文: Using Multivariate Imputation by Chained Equations…
Continual learning (CL) remains a significant challenge for deep neural networks, as it is prone to forgetting previously acquired knowledge. Several approaches have been proposed in the literature, such as experience rehearsal,…
Classifying incomplete multi-view data is inevitable since arbitrary view missing widely exists in real-world applications. Although great progress has been achieved, existing incomplete multi-view methods are still difficult to obtain a…
Due to the over-fitting problem caused by imbalance samples, there is still room to improve the performance of data-driven automatic modulation classification (AMC) in noisy scenarios. By fully considering the signal characteristics, an AMC…
Machine learning (ML) is playing an increasingly important role in scientific research. In conjunction with classical statistical approaches, ML-assisted analytical strategies have shown great promise in accelerating research findings. This…
Electronic Health Records (EHRs) exhibit a high amount of missing data due to variations of patient conditions and treatment needs. Imputation of missing values has been considered an effective approach to deal with this challenge. Existing…
Nested Monte Carlo is widely used for risk estimation, but its efficiency is limited by the discontinuity of the indicator function and high computational cost. This paper proposes a nested Multilevel Monte Carlo (MLMC) method combined with…
A new approach to estimating photometric redshifts - using Artificial Neural Networks (ANNs) - is investigated. Unlike the standard template-fitting photometric redshift technique, a large spectroscopically-identified training set is…
This paper presents details of our winning solutions to the task IV of NIPS 2017 Competition Track entitled Classifying Clinically Actionable Genetic Mutations. The machine learning task aims to classify genetic mutations based on text…
When deploying machine learning estimators in science and engineering (SAE) domains, it is critical to avoid failed estimations that can have disastrous consequences, e.g., in aero engine design. This work focuses on detecting and…
We present a census of local active galactic nuclei (AGN) at a redshift of $z\leq0.025$ selected using the high-ionization [Ne v] $\lambda14.32\,\mu$m emission line from the Infrared Database of Extragalactic Observables from Spitzer…
Data acquisition and recording in the form of databases are routine operations. The process of collecting data, however, may experience irregularities, resulting in databases with missing data. Missing entries might alter analysis…
Over the past 30 years, numerous large-scale photometric astronomical surveys have been conducted, including SDSS, Pan-STARRS, Gaia,2MASS, WISE, and others. These surveys provide extensive photometric measurements that can be used to infer…
Uncertainty estimation is a key issue when considering the application of deep neural network methods in science and engineering. In this work, we introduce a novel algorithm that quantifies epistemic uncertainty via Monte Carlo sampling…
Multiple Instance Learning (MIL) is a weak supervision learning paradigm that allows modeling of machine learning problems in which labels are available only for groups of examples called bags. A positive bag may contain one or more…
Multiple imputation is a well-established general technique for analyzing data with missing values. A convenient way to implement multiple imputation is sequential regression multiple imputation (SRMI), also called chained equations…
We explore the accuracy of the clustering-based redshift inference within the MICE2 simulation. This method uses the spatial clustering of galaxies between a spectroscopic reference sample and an unknown sample. The goal of this study is to…
We consider the problem of handling missing data with deep latent variable models (DLVMs). First, we present a simple technique to train DLVMs when the training set contains missing-at-random data. Our approach, called MIWAE, is based on…
We present the result of a spectroscopic campaign targeting Active Galactic Nucleus (AGN) candidates selected using a novel unsupervised machine-learning (ML) algorithm trained on optical and mid-infrared (mid-IR) photometry. AGN candidates…
We use the Random Forest (RF) algorithm to develop a tool for automated activity classification of galaxies into 5 different classes: Star-forming (SF), AGN, LINER, Composite, and Passive. We train the algorithm on a combination of mid-IR…
\Multiple imputation (MI) is a popular and well-established method for handling missing data in multivariate data sets, but its practicality for use in massive and complex data sets has been questioned. One such data set is the Panel Study…