Related papers: Scaling pair count to next galaxy surveys
Correlation clustering is a widely-used approach for clustering large data sets based only on pairwise similarity information. In recent years, there has been a steady stream of better and better classical algorithms for approximating this…
Scene graph generation (SGG) in satellite imagery (SAI) benefits promoting understanding of geospatial scenarios from perception to cognition. In SAI, objects exhibit great variations in scales and aspect ratios, and there exist rich…
It is commonplace in cosmology to analyze fields projected onto the celestial sphere, and in particular density fields that are defined by a set of points e.g. galaxies. When performing an harmonic-space analysis of such data (e.g. an…
With the emergence of social networks, online platforms dedicated to different use cases, and sensor networks, the emergence of large-scale graph community detection has become a steady field of research with real-world applications.…
Standard cosmology is based on the assumption that the universe is spatially homogeneous. However the consensus on a homogeneous matter structure, even on very large scales, has never been complete. The advantage of correlation dimension…
Modern redshift surveys enable the identification of large samples of galaxies in pairs, taken from many different environments. Meanwhile, cosmological simulations allow a detailed understanding of the statistical properties of the…
We have developed a new analytic method to calculate the galaxy two-point correlation functions (TPCFs) accurately and efficiently, applicable to surveys with finite, regular, and mask-free geometries. We have derived simple, accurate…
As galaxy surveys become larger and more complex, keeping track of the completeness, magnitude limit, and other survey parameters as a function of direction on the sky becomes an increasingly challenging computational task. For example,…
Identifying independence between two random variables or correlated given their samples has been a fundamental problem in Statistics. However, how to do so in a space-efficient way if the number of states is large is not quite well-studied.…
SCAN (Structural Clustering Algorithm for Networks) is a well-studied, widely used graph clustering algorithm. For large graphs, however, sequential SCAN variants are prohibitively slow, and parallel SCAN variants do not effectively share…
Spatial scan statistics are well-known methods for cluster detection and are widely used in epidemiology and medical studies for detecting and evaluating the statistical significance of disease hotspots. For the sake of simplicity, the…
Images have become an important data source in many scientific and commercial domains. Analysis and exploration of image collections often requires the retrieval of the best subregions matching a given query. The support of such…
We use pair and environmental classifications of ~ 211,000 star-forming galaxies from the Sloan Digital Sky Survey, along with a suite of merger simulations, to investigate the enhancement of star formation as a function of separation in…
Real-time multi-camera 3D reconstruction is crucial for 3D perception, immersive interaction, and robotics. Existing methods struggle with multi-view fusion, camera extrinsic uncertainty, and scalability for large camera setups. We propose…
The growing field of large-scale time domain astronomy requires methods for probabilistic data analysis that are computationally tractable, even with large datasets. Gaussian Processes are a popular class of models used for this purpose…
Cross-matching operation, which is to find corresponding data for the same celestial object or region from multiple catalogues,is indispensable to astronomical data analysis and research. Due to the large amount of astronomical catalogues…
We consider big spatial data, which is typically produced in scientific areas such as geological or seismic interpretation. The spatial data can be produced by observation (e.g. using sensors or soil instrument) or numerical simulation…
Realistic lightcone mocks are important in the clustering analyses of large galaxy surveys. For simulations where only the snapshots are available, it is common to create approximate lightcones by joining together the snapshots in spherical…
The existing algorithms for processing skyline queries cannot adapt to big data. This paper proposes two approximate skyline algorithms based on sampling. The first algorithm obtains a fixed size sample and computes the approximate skyline…
Context. Studies of galaxy pairs can provide valuable information to jointly understand the formation and evolution of galaxies and galaxy groups. Consequently, taking into account the new high precision photo-z surveys, it is important to…