Related papers: SDSS-RASS: Next Generation of Cluster-Finding Algo…
Discovering and clustering subspaces in high-dimensional data is a fundamental problem of machine learning with a wide range of applications in data mining, computer vision, and pattern recognition. Earlier methods divided the problem into…
In recent years, machine learning (ML) algorithms have been successfully employed in Astronomy for analyzing and interpreting the data collected from various surveys. The need for new robust and efficient data analysis tools in Astronomy is…
We present recent results from the Laboratory for Cosmological Data Mining (http://lcdm.astro.uiuc.edu) at the National Center for Supercomputing Applications (NCSA) to provide robust classifications and photometric redshifts for objects in…
Solid observational evidences indicate a strong dependence of the galaxy formation and evolution on the environment. In order to study in particular the interaction between the intracluster medium and the evolution of cluster galaxies, we…
We study 203 (of 442) Swift AGN and Cluster Survey extended X-ray sources located in the SDSS DR8 footprint to search for galaxy over-densities in three dimensional space using SDSS galaxy photometric redshifts and positions near the Swift…
We use a sample of galaxies from the Two Micron All Sky Survey (2MASS) Extended Source Catalog to refine a matched filter method of finding galaxy clusters that takes into account each galaxy's position, magnitude, and redshift if…
This paper presents stellar mass functions and i-band luminosity functions for Sloan Digital Sky Survey (SDSS) galaxies at $i < 21$ using clustering redshifts, from which we also compute targeting completeness measurements for the Baryon…
This paper considers the problem of clustering a collection of unlabeled data points assumed to lie near a union of lower-dimensional planes. As is common in computer vision or unsupervised learning applications, we do not know in advance…
Co-clustering simultaneously clusters rows and columns, revealing more fine-grained groups. However, existing co-clustering methods suffer from poor scalability and cannot handle large-scale data. This paper presents a novel and scalable…
We present a catalog of galaxy clusters detected in a new ROSAT PSPC survey. The survey is optimized to sample, at high redshifts, the mass range corresponding to T> keV clusters at z=0. Technically, our survey is the extension of the 160…
We present a description of the observations and data reduction procedures for an extensive spectroscopic and multi-band photometric study of nine high redshift, optically-selected cluster candidates. The primary goal of the survey is to…
In this paper, we present a deep extension of Sparse Subspace Clustering, termed Deep Sparse Subspace Clustering (DSSC). Regularized by the unit sphere distribution assumption for the learned deep features, DSSC can infer a new data…
Since the late 1970's, redshift surveys have been vital for progress in understanding large-scale structure in the Universe. The original CfA redshift survey collected spectra of 20-30 galaxies per clear night on a 1.5 meter telescope; over…
Low-latency instance segmentation of LiDAR point clouds is crucial in real-world applications because it serves as an initial and frequently-used building block in a robot's perception pipeline, where every task adds further delay.…
Galaxy clusters are usually detected in blind optical surveys via suitable filtering methods. We present an optimal matched filter which maximizes their signal-to-noise ratio by taking advantage of the knowledge we have of their intrinsic…
This paper introduces {\em fusion subspace clustering}, a novel method to learn low-dimensional structures that approximate large scale yet highly incomplete data. The main idea is to assign each datum to a subspace of its own, and minimize…
Clusters of galaxies are the most massive objects in the Universe and mapping their location is an important astronomical problem. This paper describes an algorithm (based on statistical signal processing methods), a software architecture…
The problem of estimating the number of clusters (say k) is one of the major challenges for the partitional clustering. This paper proposes an algorithm named k-SCC to estimate the optimal k in categorical data clustering. For the…
Distributed data mining techniques and mainly distributed clustering are widely used in the last decade because they deal with very large and heterogeneous datasets which cannot be gathered centrally. Current distributed clustering…
In the theoretical framework of hierarchical structure formation, galaxy clusters evolve through continuous accretion and mergers of substructures. Cosmological simulations have revealed the best picture of the Universe as a 3-D filamentary…