Related papers: Multi-Scale CLEAN: A comparison of its performance…
The increase in the observed volume in cosmological surveys imposes various challenges on simulation preparations. Firstly, the volume of the simulations required increases proportionally to the observations. However, large-volume…
We present a comparative analysis of estimators and Bayesian methods for determining the number count dipole from cosmological surveys. The increase in discordance between the number count dipole and the CMB's kinematic dipole has presented…
In modern large-scale distributed systems, analytics jobs submitted by various users often share similar work, for example scanning and processing the same subset of data. Instead of optimizing jobs independently, which may result in…
We study the problem of applying spectral clustering to cluster multi-scale data, which is data whose clusters are of various sizes and densities. Traditional spectral clustering techniques discover clusters by processing a similarity…
Benchmark datasets in computer vision often contain off-topic images, near duplicates, and label errors, leading to inaccurate estimates of model performance. In this paper, we revisit the task of data cleaning and formalize it as either a…
Correlation clustering is a widely-used approach for clustering large data sets based only on pairwise similarity information. In recent years, there has been a steady stream of better and better classical algorithms for approximating this…
Future large scale structure surveys will measure the locations and shapes of billions of galaxies. The precision of such catalogs will require meticulous treatment of systematic contamination of the observed fields. We compare several…
DBSCAN is a classical density-based clustering procedure with tremendous practical relevance. However, DBSCAN implicitly needs to compute the empirical density for each sample point, leading to a quadratic worst-case time complexity, which…
The Small-Correlated-Against-Large Estimator (SCALE) for small-scale lensing of the cosmic microwave background (CMB) provides a novel method for measuring the amplitude of CMB lensing power without the need for reconstruction of the…
Cross-correlating the data of neutral hydrogen (HI) 21cm intensity mapping with galaxy surveys is an effective method to extract astrophysical and cosmological information. In this work, we investigate the cross-correlation of MeerKAT…
Matrix multiplication is a foundational operation in scientific computing and machine learning, yet its computational complexity makes it a significant bottleneck for large-scale applications. The shift to parallel architectures, primarily…
Data analysis require a pairwise proximity measure over objects. Recent work has extended this to situations where the distance information between objects is given as comparison results of distances between three objects (triplets). Humans…
Classification is a popular task in the field of Machine Learning (ML) and Artificial Intelligence (AI), and it happens when outputs are categorical variables. There are a wide variety of models that attempts to draw some conclusions from…
We explore three different methods based on weak lensing to extract cosmological constraints from the large-scale structure. In the first approach (method I), small-scale galaxy lensing measurements of their halo mass provide a constraint…
We present SKiLLS, a suite of multi-band image simulations for the weak lensing analysis of the complete Kilo-Degree Survey (KiDS), dubbed KiDS-Legacy analysis. The resulting catalogues enable joint shear and redshift calibration, enhancing…
Cosmological measurements require the calculation of nontrivial quantities over large datasets. The next generation of survey telescopes (such as DES, PanSTARRS, and LSST) will yield measurements of billions of galaxies. The scale of these…
Typically, binary classification lens-finding schemes are used to discriminate between lens candidates and non-lenses. However, these models often suffer from substantial false-positive classifications. Such false positives frequently occur…
One of the major performance and scalability bottlenecks in large scientific applications is parallel reading and writing to supercomputer I/O systems. The usage of parallel file systems and consistency requirements of POSIX, that all the…
Scientific discoveries are increasingly driven by analyzing large volumes of image data. Many new libraries and specialized database management systems (DBMSs) have emerged to support such tasks. It is unclear, however, how well these…
Weak gravitational lensing is a powerful probe for constraining cosmological parameters, but its success relies on accurate shear measurements. In this paper, we use image simulations to investigate how a joint analysis of high-resolution…