English
Related papers

Related papers: LakeBench: Benchmarks for Data Discovery over Data…

200 papers

Commonsense question-answering (QA) tasks, in the form of benchmarks, are constantly being introduced for challenging and comparing commonsense QA systems. The benchmarks provide question sets that systems' developers can use to train and…

Artificial Intelligence · Computer Science 2020-12-23 Henrique Santos , Minor Gordon , Zhicheng Liang , Gretchen Forbush , Deborah L. McGuinness

Federated learning is a new machine learning paradigm. The goal is to build a machine learning model from the data sets distributed on multiple devices so-called an isolated data island, while keeping their data secure and private. Most…

Machine Learning · Computer Science 2021-03-15 Yuan Liang , Yange Guo , Yanxia Gong , Chunjie Luo , Jianfeng Zhan , Yunyou Huang

Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack…

Cloud data lakes provide a modern solution for managing large volumes of data. The fundamental principle behind these systems is the separation of compute and storage layers. In this architecture, inexpensive cloud storage is utilized for…

Databases · Computer Science 2025-10-20 Gregory , Weintraub

Microservice architectures have become a popular approach for designing scalable distributed applications. Despite their extensive use in industrial settings for over a decade, there is limited understanding of the data management…

Databases · Computer Science 2025-04-29 Rodrigo Laigner , Zhexiang Zhang , Yijian Liu , Leonardo Freitas Gomes , Yongluan Zhou

Data discovery - retrieving relevant tables from a data lake in response to user queries - is a fundamental building block for downstream analytics. In practice, data discovery must support different query modalities, including natural…

Relational Database Management Systems designed for Online Analytical Processing (RDBMS-OLAP) have been foundational to democratizing data and enabling analytical use cases such as business intelligence and reporting for many years.…

Databases · Computer Science 2023-10-16 Dipankar Mazumdar , Jason Hughes , JB Onofre

Assessing the quality and impact of individual data points is critical for improving model performance and mitigating undesirable biases within the training dataset. Several data valuation algorithms have been proposed to quantify data…

Machine Learning · Computer Science 2023-10-16 Kevin Fu Jiang , Weixin Liang , James Zou , Yongchan Kwon

Data products are reusable, self-contained assets designed for specific business use cases. Automating their discovery is of great industry interest, as it enables efficient data access in large data lakes and supports analytical workflows.…

Information Retrieval · Computer Science 2026-03-19 Liangliang Zhang , Nandana Mihindukulasooriya , Niharika S. D'Souza , Sola Shirai , Sarthak Dash , Yao Ma , Horst Samulowitz

Data management can be a complex challenge in fields such as bioinformatics and health sciences, which continuously generate extensive heterogeneous datasets. In the context of collaborative global health initiatives, secure storage and…

Software Engineering · Computer Science 2026-05-20 Danilo Silva , Monika Moir , Cheryl Baxter , Tulio de Oliveira , Joicymara Xavier , Marcel Dunaiski

Entity resolution (ER) is the process of identifying records that refer to the same entities within one or across multiple databases. Numerous techniques have been developed to tackle ER challenges over the years, with recent emphasis…

Databases · Computer Science 2023-11-14 George Papadakis , Nishadi Kirielle , Peter Christen , Themis Palpanas

In this paper, we present a new DBMS performance benchmark that can simulate user exploration with any specified dashboard design made of standard visualization and interaction components. The distinguishing feature of our SImulation-BAsed…

Human-Computer Interaction · Computer Science 2025-01-14 Joanna Purich , Anthony Wise , Leilani Battle

As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly urgent. Existing benchmarks and evaluations often focus on…

Effective processing, interpretation, and management of sensor data have emerged as a critical component of cyber-physical systems. Traditionally, processing sensor data requires profound theoretical knowledge and proficiency in…

Artificial Intelligence · Computer Science 2025-04-01 Pengrui Quan , Xiaomin Ouyang , Jeya Vikranth Jeyakumar , Ziqi Wang , Yang Xing , Mani Srivastava

Branchable databases are evolving from developer tools to infrastructure for agentic workloads characterized by speculative mutations and non-linear state exploration. Traditional RDBMS mechanisms such as nested transactions do not provide…

Databases · Computer Science 2026-04-21 Elaine Ang , Sam Weldon , In Keun Kim , Kevin Durand , Kostis Kaffes , Eugene Wu

It is challenging to determine whether datasets are findable, accessible, interoperable, and reusable (FAIR) because the FAIR Guiding Principles refer to highly idiosyncratic criteria regarding the metadata used to annotate datasets.…

Digital Libraries · Computer Science 2022-10-17 Mark A. Musen , Martin J. O'Connor , Erik Schultes , Marcos Martinez-Romero , Josef Hardi , John Graybeal

Foundation models have enabled rapid progress across many specialized domains by leveraging large-scale pre-training on unlabeled data, demonstrating strong generalization to a variety of downstream tasks. While such models have gained…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Mirali Purohit , Bimal Gajera , Vatsal Malaviya , Irish Mehta , Kunal Kasodekar , Jacob Adler , Steven Lu , Umaa Rebbapragada , Hannah Kerner

Existing text-to-SQL benchmarks have largely been constructed from public databases with well-structured schemas and simplistic question-SQL pairs. While large language models (LLMs) excel on these settings, their efficacy in complex…

Computation and Language · Computer Science 2026-05-14 Peter Baile Chen , Devin Yang , Weiyue Li , Fabian Wenz , Yi Zhang , Nesime Tatbul , Michael Cafarella , Çağatay Demiralp , Michael Stonebraker

Data and workload drift are key to evaluating database components such as caching, cardinality estimation, indexing, and query optimization. Yet, existing benchmarks are static, offering little to no support for modeling drift. This…

Databases · Computer Science 2025-10-14 Guanli Liu , Renata Borovica-Gajic

Benchmarking the performance of community detection methods on empirical social network data has been identified as critical for improving these methods. In particular, while most current research focuses on detecting communities in data…

Social and Information Networks · Computer Science 2013-02-05 Conrad Lee , Pádraig Cunningham