Related papers: A Processing Pipeline for High Volume Pulsar Data …
Coming high-cadence wide-field optical telescopes will image hundreds of thousands of sources per minute. Besides inspecting the near real-time data streams for transient and variability events, the accumulated data archive is a wealthy…
The ongoing exponential growth of computational power, and the growth of the commercial High Performance Computing (HPC) industry, has led to a point where ten commercial systems currently exceed the performance of the highest-used HPC…
Machine learning (ML) offers powerful methods for detecting and modeling associations often in data with large feature spaces and complex associations. Many useful tools/packages (e.g. scikit-learn) have been developed to make the various…
The Pulsar Virtual Observatory will provide a means for scientists in all fields to access and analyze the large data sets stored in pulsar surveys without specific knowledge about the data or the processing mechanisms. This is achieved by…
In the past couple of decades, the computational abilities of supercomput- ers have increased tremendously. Leadership scale supercomputers now are capable of petaflops. Likewise, the problem size targeted by applications running on such…
Reducing a set of numbers to a single value is a fundamental operation in applications such as signal processing, data compression, scientific computing, and neural networks. Accumulation, which involves summing a dataset to obtain a single…
Stellar streams are potentially a very sensitive observational probe of galactic astrophysics, as well as the dark matter population in the Milky Way. On the other hand, performing a detailed, high-fidelity statistical analysis of these…
Generation of science-ready data from processed data products is one of the major challenges in next-generation radio continuum surveys with the Square Kilometre Array (SKA) and its precursors, due to the expected data volume and the need…
The extension of on-board data processing capabilities is an attractive option to reduce telemetry for scientific instruments on deep space missions. The challenges that this presents, however, require a comprehensive software system, which…
As the complexity of enterprise systems increases, the need for monitoring and analyzing such systems also grows. A number of companies have built sophisticated monitoring tools that go far beyond simple resource utilization reports. For…
This paper presents the DDF Pipeline, a radio astronomy data processing tool initially designed for the LOw-Frequency ARray (LO- FAR) radio-telescope and a candidate for processing data from the Square Kilometre Array (SKA). This work…
Data is rapidly increasing in volume and velocity and the Internet of Things (IoT) is one important source of this data. The IoT is a collection of connected devices (things) which are constantly recording data from their surroundings using…
Three pulsar timing arrays are now producing high quality data sets. As reviewed in this paper, these data sets are been processed to 1) develop a pulsar-based time standard, 2) search for errors in the solar system planetary ephemeris and…
The sensitivity of ongoing searches for gravitational wave (GW) sources in the ultra-low frequency regime ($10^{-9}$ Hz to $10^{-7}$ Hz) using Pulsar Timing Arrays (PTAs) will continue to increase in the future as more well-timed pulsars…
Many questions in computational social science rely on datasets assembled from heterogeneous online sources, a process that is often labor-intensive, costly, and difficult to reproduce. Recent advances in large language models enable…
Data preparation, especially data cleaning, is very important to ensure data quality and to improve the output of automated decision systems. Since there is no single tool that covers all steps required, a combination of tools -- namely a…
We introduce an advanced information extraction pipeline to automatically process very large collections of unstructured textual data for the purpose of investigative journalism. The pipeline serves as a new input processor for the upcoming…
We describe the software requirement and design specifications for all-sky panoramic astronomical pipelines. The described software aims to meet the specific needs of super-wide angle optics, and includes cosmic-ray hit rejection, image…
In this paper, we evaluate Apache Spark for a data-intensive machine learning problem. Our use case focuses on policy diffusion detection across the state legislatures in the United States over time. Previous work on policy diffusion has…
Weak gravitational lensing measurements are traditionally made at optical wavelengths where many highly resolved galaxy images are readily available. However, the Square Kilometre Array (SKA) holds great promise for this type of measurement…