English
Related papers

Related papers: Pipeline Provenance for Analysis, Evaluation, Trus…

200 papers

Data mining is about obtaining new knowledge from existing datasets. However, the data in the existing datasets can be scattered, noisy, and even incomplete. Although lots of effort is spent on developing or fine-tuning data mining models…

Machine Learning · Computer Science 2019-06-21 Canchen Li

Organizations of all kinds, whether public or private, profit-driven or non-profit, and across various industries and sectors, rely on dashboards for effective data visualization. However, the reliability and efficacy of these dashboards…

Human-Computer Interaction · Computer Science 2023-09-19 Johne Jarske , Jorge Rady , Lucia V. L. Filgueiras , Leandro M. Velloso , Tania L. Santos

Artificial intelligence is increasingly deployed to synthesize large-scale public input in policy consultations and participatory processes. Yet no formal framework exists for auditing whether these summaries faithfully represent the source…

Artificial Intelligence · Computer Science 2026-04-23 Sachit Mahajan

Provenance has been thought of a mechanism to verify a workflow and to provide workflow reproducibility. This provenance of scientific workflows has been effectively carried out in Grid based scientific workflow systems. However, recent…

Databases · Computer Science 2015-12-01 Khawar Hasham , Kamran Munir , Richard McClatchey

The Kieker observability framework is a tool that provides users with the means to design a custom observability pipeline for their application. Originally tailored for Java, supporting Python with Kieker is worthwhile. Python's popularity…

Software Engineering · Computer Science 2025-08-06 Daphné Larrivain , Shinhyung Yang , Wilhelm Hasselbring

Scientists rely on simulations to study natural phenomena. Trusting the simulation results is vital to develop sciences in any field. One approach to build trust is to ensure the reproducibility and traceability of the simulations through…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-09-21 Paula Olaya , Jay Lofstead , Michela Taufer

FMRI data are noisy, complicated to acquire, and typically go through many steps of processing before they are used in a study or clinical practice. Being able to visualize and understand the data from the start through the completion of…

Neurons and Cognition · Quantitative Biology 2024-11-21 Richard C. Reynolds , Daniel R. Glen , Gang Chen , Ziad S. Saad , Robert W. Cox , Paul A. Taylor

To reduce and analyze astronomical images, astronomers can rely on a wide range of libraries providing low-level implementations of legacy algorithms. However, combining these routines into robust and functional pipelines requires a major…

Instrumentation and Methods for Astrophysics · Physics 2021-11-17 Lionel J. Garcia , Mathilde Timmermans , Francisco J. Pozuelos , Elsa Ducrot , Michaël Gillon , Laetitia Delrez , Robert D. Wells , Emmanuël Jehin

PaPy, which stands for parallel pipelines in Python, is a highly flexible framework that enables the construction of robust, scalable workflows for either generating or processing voluminous datasets. A workflow is created from user-written…

Programming Languages · Computer Science 2014-07-17 Marcin Cieslik , Cameron Mura

Discovering authoritative links between publications and the datasets that they use can be a labor-intensive process. We introduce a natural language processing pipeline that retrieves and reviews publications for informal references to…

Digital Libraries · Computer Science 2023-05-03 Sara Lafia , Lizhou Fan , Libby Hemphill

The recomputability and reproducibility of results from scientific software requires access to both the source code and all associated input and output data. However, the full collection of these resources often does not accompany the key…

Computational Engineering, Finance, and Science · Computer Science 2015-12-24 Christian T. Jacobs , Alexandros Avdis , Gerard J. Gorman , Matthew D. Piggott

The online data reduction service reductus transforms measurements in experimental science from laboratory coordinates into physically meaningful quantities with accurate estimation of uncertainties based on instrumental settings and…

Data Analysis, Statistics and Probability · Physics 2018-11-15 Brian Maranville , William Ratcliff , Paul Kienzle

Visualisation facilitates the understanding of scientific data both through exploration and explanation of visualised data. Provenance contributes to the understanding of data by containing the contributing factors behind a result. With the…

Databases · Computer Science 2015-02-06 Bilal Arshad , Kamran Munir , Richard McClatchey , Saad Liaquat

Science is conducted collaboratively, often requiring the sharing of knowledge about computational experiments. When experiments include only datasets, they can be shared using Uniform Resource Identifiers (URIs) or Digital Object…

Databases · Computer Science 2018-06-19 Zhihao Yuan , Dai Hai Ton That , Siddhant Kothari , Gabriel Fils , Tanu Malik

Data preparation, especially data cleaning, is very important to ensure data quality and to improve the output of automated decision systems. Since there is no single tool that covers all steps required, a combination of tools -- namely a…

Databases · Computer Science 2023-08-29 Valerie Restat

Recommender systems have demonstrated significant impact across diverse domains, yet ensuring the reproducibility of experimental findings remains a persistent challenge. A primary obstacle lies in the fragmented and often opaque data…

Sustainable urban environments based on Internet of Things (IoT) technologies require appropriate policy management. However, such policies are established as a result of underlying, potentially complex and long-term policy making…

Computers and Society · Computer Science 2018-04-20 Barkha Javed , Richard McClatchey , Zaheer Khan , Jetendr Shamdasani

Process mining techniques such as process discovery and conformance checking provide insights into actual processes by analyzing event data that are widely available in information systems. These data are very valuable, but often contain…

Cryptography and Security · Computer Science 2020-09-25 Majid Rafiei , Wil M. P. van der Aalst

One major challenge in science is to make all results potentially reproducible. Thus, along with the raw data, every step from basic processing of the data, evaluation, to the generation of the figures, has to be documented as clearly as…

Graphics · Computer Science 2020-07-31 Richard Gerum

Sales pipeline analysis is fundamental to proactive management of an enterprize's sales pipeline and critical for business success. In particular, win propensity prediction, which involves quantitatively estimating the likelihood that…

Computers and Society · Computer Science 2015-02-24 Junchi Yan , Min Gong , Changhua Sun , Jin Huang , Stephen M. Chu