English
Related papers

Related papers: Dataversifying Natural Sciences: Pioneering a Data…

200 papers

Data lakes have emerged as a flexible and scalable solution for storing and analyzing large volumes of heterogeneous data, including structured, semi-structured, and unstructured formats. Despite their growing adoption in both industry and…

Databases · Computer Science 2026-01-28 Yi Lyu , Pei-Chieh Lo , Natan Lidukhover

The continuous growth of data production in almost all scientific areas raises new problems in data access and management, especially in a scenario where the end-users, as well as the resources that they can access, are worldwide…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-08-16 Tommaso Tedeschi , Diego Ciangottini , Marco Baioletti , Valentina Poggioni , Daniele Spiga , Loriano Storchi , Mirco Tracolli

Data commons collate data with cloud computing infrastructure and commonly used software services, tools and applications to create biomedical resources for the large-scale management, analysis, harmonization, and sharing of biomedical…

Genomics · Quantitative Biology 2018-12-27 Robert L. Grossman

With the rise of big data, business intelligence had to find solutions for managing even greater data volumes and variety than in data warehouses, which proved ill-adapted. Data lakes answer these needs from a storage point of view, but…

Databases · Computer Science 2018-07-12 Iuri Nogueira , Maram Romdhane , Jérôme Darmont

This manuscript provides a systemic and data-centric view of what we term essential data science, as a natural ecosystem with challenges and missions stemming from the fusion of data universe with its multiple combinations of the 5D…

Machine Learning · Computer Science 2026-01-14 Emilio Porcu , Roy El Moukari , Laurent Najman , Francisco Herrera , Horst Simon

The proliferation of data across the system lifecycle presents both a significant opportunity and a challenge for Engineering Design and Systems Engineering (EDSE). While this "digital thread" has the potential to drive innovation, the…

Software Engineering · Computer Science 2026-03-19 H. Sinan Bank , Daniel R. Herber

In the last few years, the concept of data lake has become trendy for data storage and analysis. Thus, several design alternatives have been proposed to build data lake systems. However, these proposals are difficult to evaluate as there…

Databases · Computer Science 2021-10-05 Pegdwendé Sawadogo , Jérôme Darmont

The data volumes stored in telescope archives is constantly increasing due to the development and improvements in the instrumentation. Often the archives need to be stored over a distributed storage architecture, provided by independent…

Instrumentation and Methods for Astrophysics · Physics 2022-02-07 Y. G. Grange , V. N. Pandey , X. Espinal , R. Di Maria , A. P. Millar

Recent advancements in Earth system science have been marked by the exponential increase in the availability of diverse, multivariate datasets characterised by moderate to high spatio-temporal resolutions. Earth System Data Cubes (ESDCs)…

Scientific data governance should prioritize maximizing the utility of data throughout the research lifecycle. Research software systems that enable analysis reproducibility inform data governance policies and assist administrators in…

While automated experiments and high-throughput methods are becoming more mainstream in the age of data, empowering individual researchers to capture, collate, and contextualize their data faster and more reproducibly still remains a…

Computers and Society · Computer Science 2020-07-30 Ha-Kyung Kwon , Chirranjeevi Balaji Gopal , Jared Kirschner , Santiago Caicedo , Brian D. Storey

Data analytics stands to benefit from the increasing availability of datasets that are held without their conceptual relationships being explicitly known. When collected, these datasets form a data lake from which, by processes like data…

Databases · Computer Science 2020-11-23 Alex Bogatu , Alvaro A. A. Fernandes , Norman W. Paton , Nikolaos Konstantinou

We make a case for "planetary computing" -- infrastructure to handle the ingestion, transformation, analysis and publication of global data products for furthering environmental science and enabling better informed policy-making. We draw on…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-06-04 Patrick Ferris , Michael Dales , Sadiq Jaffer , Amelia Holcomb , Eleanor Toye Scott , Thomas Swinfield , Alison Eyres , Andrew Balmford , David Coomes , Srinivasan Keshav , Anil Madhavapeddy

There has been an increasing recognition of the value of data and of data-based decision making. As a consequence, the development of data science as a field of study has intensified in recent years. However, there is no systematic and…

Other Computer Science · Computer Science 2024-01-05 M. Tamer Özsu

Given a set of deep learning models, it can be hard to find models appropriate to a task, understand the models, and characterize how models are different one from another. Currently, practitioners rely on manually-written documentation to…

Databases · Computer Science 2025-02-24 Koyena Pal , David Bau , Renée J. Miller

Metadata management for distributed data sources is a long-standing but ever-growing problem. To counter this challenge in a research-data and library-oriented setting, this work constructs a data architecture, derived from the data-lake:…

Databases · Computer Science 2026-05-08 Christian Himpe

This paper presents a multifarious examination of natural resources and environmental scientists' adventures navigating the policy change towards open access and cultural shift in data management, sharing, and reuse. Situated in the…

Digital Libraries · Computer Science 2018-03-06 Yi Shen

Recent advances in large language models (LLMs) have enabled a new class of AI agents that automate multiple stages of the data science workflow by integrating planning, tool use, and multimodal reasoning across text, code, tables, and…

Research data are the foundation of Artificial Intelligence (AI)-driven science, yet current AI applications remain limited to a few fields with readily available, well-structured, digitized datasets. Achieving comprehensive AI empowerment…

Organizations routinely accumulate semi-structured log datasets generated as the output of code; these datasets remain unused and uninterpreted, and occupy wasted space - this phenomenon has been colloquially referred to as "data lake"…

Databases · Computer Science 2018-03-01 Yihan Gao , Silu Huang , Aditya Parameswaran