English

A Thorough Investigation of Content-Defined Chunking Algorithms for Data Deduplication

Distributed, Parallel, and Cluster Computing 2024-10-22 v3

Abstract

Data deduplication emerged as a powerful solution for reducing storage and bandwidth costs in cloud settings by eliminating redundancies at the level of chunks. This has spurred the development of numerous Content-Defined Chunking (CDC) algorithms over the past two decades. Despite advancements, the current state-of-the-art remains obscure, as a thorough and impartial analysis and comparison is lacking. We conduct a rigorous theoretical analysis and impartial experimental comparison of several leading CDC algorithms. Using four realistic datasets, we evaluate these algorithms against four key metrics: throughput, deduplication ratio, average chunk size, and chunk-size variance. Our analyses, in many instances, extend the findings of their original publications by reporting new results and putting existing ones into context. Moreover, we highlight limitations that have previously gone unnoticed. Our findings provide valuable insights that inform the selection and optimization of CDC algorithms for practical applications in data deduplication.

Keywords

Cite

@article{arxiv.2409.06066,
  title  = {A Thorough Investigation of Content-Defined Chunking Algorithms for Data Deduplication},
  author = {Marcel Gregoriadis and Leonhard Balduf and Björn Scheuermann and Johan Pouwelse},
  journal= {arXiv preprint arXiv:2409.06066},
  year   = {2024}
}

Comments

Submitted to IEEE Transactions on Cloud Computing for possible publication

R2 v1 2026-06-28T18:39:13.921Z