Related papers: Autoclassification of the Variable 3XMM Sources Us…
We propose a Fourier-based learning algorithm for highly nonlinear multiclass classification. The algorithm is based on a smoothing technique to calculate the probability distribution of all classes. To obtain the probability distribution,…
The Optical Monitor Catalogue of serendipitous sources (OMCat) contains entries for every source detected in the publicly available XMM-Newton Optical Monitor (OM) images taken in either the imaging or ``fast'' modes. Since the OM is…
In this paper, a new approach for classification of target task using limited labeled target data as well as enormous unlabeled source data is proposed which is called self-taught learning. The target and source data can be drawn from…
The Third Catalog of Hard Fermi Large Area Telescope Sources (3FHL) reports the detection of 1556 objects at E > 10 GeV. However, 177 sources remain unassociated and 23 are associated with a ROSAT X-ray detection of unknown origin. Pointed…
Most machine learning (ML) algorithms have several stochastic elements, and their performances are affected by these sources of randomness. This paper uses an empirical study to systematically examine the effects of two sources: randomness…
This study revisits the findings of Carl et al., who evaluated the pre-trained Google Inception-ResNet-v2 model for automated detection of European wild mammal species in camera trap images. To assess the reproducibility and…
Random forests are a scheme proposed by Leo Breiman in the 2000's for building a predictor ensemble with a set of decision trees that grow in randomly selected subspaces of data. Despite growing interest and practical use, there has been…
We perform an MMT/Hectospec redshift survey of the North Ecliptic Pole Wide (NEPW) field covering 5.4 square degrees, and use it to estimate the photometric redshifts for the sources without spectroscopic redshifts. By combining 2572 newly…
We present preliminary results from our on-going study: Comparing and optimizing source detection procedures for XMM images. By constructing realistic spatial and spectral source distributions and ``observing'' these through the XMM Science…
Random forests are a statistical learning method widely used in many areas of scientific research because of its ability to learn complex relationships between input and output variables and also its capacity to handle high-dimensional…
With the development of technology, the chemical production process is becoming increasingly complex and large-scale, making fault detection particularly important. However, current detective methods struggle to address the complexities of…
This paper proposes a novel type of random forests called a denoising random forests that are robust against noises contained in test samples. Such noise-corrupted samples cause serious damage to the estimation performances of random…
The XMM-RM project was designed to provide X-ray coverage of the Sloan Digital Sky Survey Reverberation Mapping (SDSS-RM) field. 41 XMM-Newton exposures, placed surrounding the Chandra AEGIS field, were taken, covering an area of 6.13 deg^2…
In this paper, we modify the proof methods of some previously weakly consistent variants of random forests into strongly consistent proof methods, and improve the data utilization of these variants in order to obtain better theoretical…
In its first four years of operation, the Fermi Large Area Telescope (LAT) detected 3033 $\gamma$-ray emitting sources. In the Fermi-LAT Third Source Catalogue (3FGL) about 50% of the sources have no clear association with a likely…
Random Forests are one of the most popular classifiers in machine learning. The larger they are, the more precise is the outcome of their predictions. However, this comes at a cost: their running time for classification grows linearly with…
Deep Learning methods are notorious for relying on extensive labeled datasets to train and assess their performance. This can cause difficulties in practical situations where models should be trained for new applications for which very…
This paper explores interpretability techniques for two of the most successful learning algorithms in medical decision-making literature: deep neural networks and random forests. We applied these algorithms in a real-world medical dataset…
The advent of large astronomical surveys has made available large and complex data sets. However, the process of discovery and interpretation of each potentially new astronomical source is, many times, still handcrafted. In this context,…
In the digital era, the exponential growth of scientific publications has made it increasingly difficult for researchers to efficiently identify and access relevant work. This paper presents an automated framework for research article…