Related papers: Computational AstroStatistics: Fast and Efficient …
The Sloan Digital Sky Survey (SDSS) started a new phase in August 2008, with new instrumentation and new surveys focused on Galactic structure and chemical evolution, measurements of the baryon oscillation feature in the clustering of…
Often the relation between the variables constituting a multivariate data space might be characterized by one or more of the terms: ``nonlinear'', ``branched'', ``disconnected'', ``bended'', ``curved'', ``heterogeneous'', or, more general,…
The All-Sky Automated Survey (ASAS) appeared to be extremely useful in establishing the census of bright variable stars in the sky. A short review of the characteristics of the ASAS data and discoveries based on these data and related to…
Nonlinear component analysis such as kernel Principle Component Analysis (KPCA) and kernel Canonical Correlation Analysis (KCCA) are widely used in machine learning, statistics and data analysis, but they can not scale up to big datasets.…
The nature of scientific and technological data collection is evolving rapidly: data volumes and rates grow exponentially, with increasing complexity and information content, and there has been a transition from static data sets to data…
We introduce a software suite developed for galaxy cluster cosmological analysis with the Dark Energy Survey Data. Cosmological analyses based on galaxy cluster number counts and weak-lensing measurements need efficient software…
Large-scale astrophysics datasets present an opportunity for new machine learning techniques to identify regions of interest that might otherwise be overlooked by traditional searches. To this end, we use Classification Without Labels…
It is commonplace in cosmology to analyze fields projected onto the celestial sphere, and in particular density fields that are defined by a set of points e.g. galaxies. When performing an harmonic-space analysis of such data (e.g. an…
Developing algorithms to search through data efficiently is a challenging part of searching for signs of technology beyond our solar system. We have built a digital signal processing system and computer cluster on the backend of the Karl G.…
Principal Component Analysis (PCA) is being extensively used in Astronomy but not yet exhaustively exploited for variability search. The aim of this work is to investigate the effectiveness of using the PCA as a method to search for…
We present an overview and statistical analysis of the data included in WEBDA. This database includes valuable information such as coordinates, rectangular positions, proper motions, photometric as well as spectroscopic data, radial and…
We revisit the method of cumulants for analysing dynamic light scattering data in particle sizing applications. Here the data, in the form of the time correlation function of scattered light, is written as a series involving the first few…
The Sloan Digital Sky Survey has surveyed 14,555 square degrees of the sky, and delivered over a trillion pixels of imaging data. We present the large-scale clustering of 1.6 million quasars between z = 0.5 and z = 2.5 that have been…
Obtaining accurate photometric redshift estimations is an important aspect of cosmology, remaining a prerequisite of many analyses. In creating novel methods to produce redshift estimations, there has been a shift towards using machine…
scida is a Python package for reading and analyzing large scientific data sets with support for various cosmological and galaxy formation simulations out-of-the-box. Data access is provided through a hierarchical dictionary-like data…
We develop a cosmological parameter estimation code for (tomographic) angular power spectra analyses of galaxy number counts, for which we include, for the first time, redshift-space distortions (RSD) in the Limber approximation. This…
Mixture models combine multiple components into a single probability density function. They are a natural statistical model for many situations in astronomy, such as surveys containing multiple types of objects, cluster analysis in various…
We present a machine learning (ML) framework for the detection of wide binary star systems using Gaia DR3 data. By training supervised ML models on established wide binary catalogues, we efficiently classify wide binaries and employ…
Data assimilation (DA) improves prediction of chaotic systems by combining model forecasts with sparse, noisy observations. Many DA methods are inherently probabilistic, but accurate probabilistic DA is often computationally expensive…
We use principal component analysis (PCA) to estimate stellar masses, mean stellar ages, star formation histories (SFHs), dust extinctions and stellar velocity dispersions for ~290,000 galaxies with stellar masses greater than $10^{11}Msun…