Related papers: ChromaStarDB: SQL database-driven spectrum synthes…
Simulation-based inference (SBI) has become an important tool in cosmology for extracting additional information from observational data using simulations. However, all cosmological simulations are approximations of the actual universe, and…
In the past decade, the modeling community has produced many feature-rich modeling editors and tool prototypes not only for modeling standards but particularly also for many domain-specific languages. More recently, however, web-based…
Skyline queries are frequently used in data analytics and multi-criteria decision support applications to filter relevant information from big amounts of data. Apache Spark is a popular framework for processing big, distributed data. The…
Constraint-based pattern discovery is at the core of numerous data mining tasks. Patterns are extracted with respect to a given set of constraints (frequency, closedness, size, etc). In the context of sequential pattern mining, a large…
Fast and accurate crystal structure prediction (CSP) algorithms and web servers are highly desirable for exploring and discovering new materials out of the infinite design space. However, currently, the computationally expensive first…
Crystal structure prediction (CSP), which aims to predict the three-dimensional atomic arrangement of a crystal from its composition, is central to materials discovery and mechanistic understanding. However, given the composition in a unit…
We present an interactive IDL program for viewing and analyzing astronomical spectra in the context of modern imaging surveys. SpecPro's interactive design lets the user simultaneously view spectroscopic, photometric, and imaging data,…
Data-driven models of stellar spectra are useful tools to study non-stellar information, such as the Diffuse Interstellar Bands (DIBs) caused by intervening interstellar material. Using $\sim 55000$ spectra of $\sim 17000$ red clump stars…
In this paper we present a new family of Intensional RDBs (IRDBs) which extends the traditional RDBs with the Big Data and flexible and 'Open schema' features, able to preserve the user-defined relational database schemas and all…
Binary star DataBase (BDB) is the database of binary/multiple systems of various observational types. BDB contains data on physical and positional parameters of 260,000 components of 120,000 stellar systems of multiplicity 2 to more than…
Sparse Subspace Clustering (SSC) has achieved state-of-the-art clustering quality by performing spectral clustering over a $\ell^{1}$-norm based similarity graph. However, SSC is a transductive method which does not handle with the data not…
We present a new pipeline for the efficient generation of synthetic observations of the extragalactic microwave sky, tailored to large ground-based CMB experiments such as the Simons Observatory, Advanced ACTPol, SPT-3G, and CMB-S4. Such…
Text-to-SQL, which translates a natural language question into an SQL query, has advanced with in-context learning of Large Language Models (LLMs). However, existing methods show little improvement in performance compared to randomly chosen…
In this paper, we present a deep extension of Sparse Subspace Clustering, termed Deep Sparse Subspace Clustering (DSSC). Regularized by the unit sphere distribution assumption for the learned deep features, DSSC can infer a new data…
Structured Query Language (SQL) remains the standard language used in Relational Database Management Systems (RDBMSs) and has found applications in healthcare (patient registries), businesses (inventories, trend analysis), military,…
Relational databases (RDBs) underpin the majority of global data management systems, where information is structured into multiple interdependent tables. To effectively use the knowledge within RDBs for predictive tasks, recent advances…
The Chinese Space Station Survey Telescope (CSST) aims to map the universe across an unprecedented dynamic range of stellar densities, spanning from extragalactic voids to the crowded Galactic center (e.g. a few stars and galaxies in the…
Molecular line intensity calculations are not a straightforward task. We present a description of the basics for including molecular linesin synthetic spectra, and of the input data needed. We aim both at describing ways in which molecular…
Graphs are a popular data type found in many domains. Numerous techniques have been proposed to find interesting patterns in graphs to help understand the data and support decision-making. However, there are generally two limitations that…
We present an approach to computing consistent answers to queries possibly involving an aggregation operator in databases operating under a star schema and possibly containing missing values and inconsistent data. Our approach is based on…