English

Accessibility Barriers in Multi-Terabyte Public Datasets: The Gap Between Promise and Practice

Computers and Society 2025-06-17 v1 Digital Libraries Information Retrieval

Abstract

The promise of "free and open" multi-terabyte datasets often collides with harsh realities. While these datasets may be technically accessible, practical barriers -- from processing complexity to hidden costs -- create a system that primarily serves well-funded institutions. This study examines accessibility challenges across web crawls, satellite imagery, scientific data, and collaborative projects, revealing a consistent two-tier system where theoretical openness masks practical exclusivity. Our analysis demonstrates that datasets marketed as "publicly accessible" typically require minimum investments of $1,000+ for meaningful analysis, with complex processing pipelines demanding $10,000-100,000+ in infrastructure costs. The infrastructure requirements -- distributed computing knowledge, domain expertise, and substantial budgets -- effectively gatekeep these datasets despite their "open" status, limiting practical accessibility to those with institutional support or substantial resources.

Keywords

Cite

@article{arxiv.2506.13256,
  title  = {Accessibility Barriers in Multi-Terabyte Public Datasets: The Gap Between Promise and Practice},
  author = {Marc Bara},
  journal= {arXiv preprint arXiv:2506.13256},
  year   = {2025}
}

Comments

5 pages, 28 references. Analysis of practical barriers to accessing multi-terabyte public datasets

R2 v1 2026-07-01T03:19:14.825Z