Related papers: Using LSDB to enable large-scale catalog distribut…
In the age of big data, it is important for primary research data to follow the FAIR principles of findability, accessibility, interoperability, and reusability. Data harmonization enhances interoperability and reusability by aligning…
Data stream learning is a very relevant paradigm because of the increasing real-world scenarios generating data at high velocities and in unbounded sequences. Stream learning aims at developing models that can process instances as they…
Clustering analysis is of substantial significance for data mining. The properties of big data raise higher demand for more efficient and economical distributed clustering methods. However, existing distributed clustering methods mainly…
The data analysis software (DAS) for VLT ESPRESSO is aimed to set a new benchmark in the treatment of spectroscopic data towards the extremely-large-telescope era, providing carefully designed, fully interactive recipes to take care of…
Data discovery is crucial for data management and analysis and can benefit from better utilization of metadata. For example, users may want to search data using queries like ``find the tables created by Alex and endorsed by Mike that…
Large time series models (LTMs) have emerged as powerful tools for universal forecasting, yet they often struggle with the inherent diversity and nonstationarity of real-world time series data, leading to an unsatisfactory trade-off between…
As urban populations grow, cities are becoming more complex, driving the deployment of interconnected sensing systems to realize the vision of smart cities. These systems aim to improve safety, mobility, and quality of life through…
Reconstructing an accurate and consistent large-scale LiDAR point cloud map is crucial for robotics applications. The existing solution, pose graph optimization, though it is time-efficient, does not directly optimize the mapping…
Development of the Rubin Observatory Legacy Survey of Space and Time (LSST) includes a series of Data Challenges (DC) arranged by various LSST Scientific Collaborations (SC) that are taking place during the projects preoperational phase.…
Reliable LiDAR perception requires robustness across sensors, environments, and adverse weather. However, existing datasets rarely provide physically consistent observations of the same scene under varying sensor configurations and weather…
We present the latest results of our search for planets with HARPS-N at the 3.6 m Telescopio Nazionale Galileo under the Tracking Advanced Planetary Systems project: an in-depth study of the 15 most Li abundant giants from the PennState -…
Combining multiple datasets enables performance boost on many computer vision tasks. But similar trend has not been witnessed in object detection when combining multiple datasets due to two inconsistencies among detection datasets: taxonomy…
In the high luminosity era of the Large Hadron Collider, the HL-LHC, the instantaneous luminosity is expected to reach unprecedented values, resulting in about 200 proton-proton interactions in a typical bunch crossing. To cope with the…
We present htof, an open-source tool for interpreting and fitting the intermediate astrometric data (IAD) from both the 1997 and 2007 reductions of Hipparcos, the scanning-law of Gaia, and future missions such as the Nancy Grace Roman Space…
Large language models (LLMs) have shown remarkable reasoning capabilities, yet aligning such abilities to small language models (SLMs) remains a challenge due to distributional mismatches and limited model capacity. Existing reasoning…
In most process control systems nowadays, process measurements are periodically collected and archived in historians. Analytics applications process the data, and provide results offline or in a time period that is considerably slow in…
The exponential growth of data is making query processing increasingly critical for modern computing infrastructure, yet the environmental impact of database operations remains poorly understood and largely overlooked. This paper presents…
Data engineering workflows require reliable differencing across files, databases, and query outputs, yet existing tools falter under schema drift, heterogeneous types, and limited explainability. SmartDiff is a unified system that combines…
Detecting small sets of relevant patterns from a given dataset is a central challenge in data mining. The relevance of a pattern is based on user-provided criteria; typically, all patterns that satisfy certain criteria are considered…
Lasair is the UK Community Broker for transient alerts from the Legacy Survey of Space and Time (LSST) from the Vera C. Rubin Observatory. We explain the system's capabilities, how users can achieve their scientific goals, and how Lasair is…