English
Related papers

Related papers: Creating Structure in Web Archives With Collection…

200 papers

The web is often treated as a durable record of institutional and social life, yet in practice it is fragile, revisable, and frequently ephemeral. Domains change, redesigns erase earlier material, institutions relocate, maintainers…

Digital Libraries · Computer Science 2026-05-22 Meliksah Yorulmazlar

We document strategies and lessons learned from sampling the web by collecting 27.3 million URLs with 3.8 billion archived pages spanning 26 years (1996-2021) from the Internet Archive's (IA) Wayback Machine. Our goal is to revisit…

Digital Libraries · Computer Science 2025-07-22 Kritika Garg , Sawood Alam , Dietrich Ayala , Mark Graham , Michele C. Weigle , Michael L. Nelson

Certain HTTP Cookies on certain sites can be a source of content bias in archival crawls. Accommodating Cookies at crawl time, but not utilizing them at replay time may cause cookie violations, resulting in defaced composite mementos that…

Digital Libraries · Computer Science 2019-06-18 Sawood Alam , Plinio Vargas , Michele C. Weigle , Michael L. Nelson

The Archives Unleashed project aims to improve scholarly access to web archives through a multi-pronged strategy involving tool creation, process modeling, and community building - all proceeding concurrently in mutually-reinforcing…

Digital Libraries · Computer Science 2020-01-16 Nick Ruest , Jimmy Lin , Ian Milligan , Samantha Fritz

Although user access patterns on the live web are well-understood, there has been no corresponding study of how users, both humans and robots, access web archives. Based on samples from the Internet Archive's public Wayback Machine, we…

Digital Libraries · Computer Science 2013-09-17 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

As the Distributed Collection Manager's work on building tools to support users maintaining collections of changing web-based resources has progressed, questions about the characteristics of people's collections of web pages have arisen.…

Digital Libraries · Computer Science 2011-01-05 Paul Logasa Bogen , Frank Shipman , Richard Furuta

The historical, cultural, and intellectual importance of archiving the web has been widely recognized. Today, all countries with high Internet penetration rate have established high-profile archiving initiatives to crawl and archive the…

Digital Libraries · Computer Science 2013-08-13 Zhiwu Xie , Herbert Van de Sompel , Jinyang Liu , Johann van Reenen , Ramiro Jordan

In a Web plagued by disappearing resources, Web archive collections provide a valuable means of preserving Web resources important to the study of past events ranging from elections to disease outbreaks. These archived collections start…

Digital Libraries · Computer Science 2019-05-30 Alexander C. Nwala , Michele C. Weigle , Michael L. Nelson

Document collections of various domains, e.g., legal, medical, or financial, often share some underlying collection-wide structure, which captures information that can aid both human users and structure-aware models. We propose to identify…

Computation and Language · Computer Science 2025-08-27 Gili Lior , Yoav Goldberg , Gabriel Stanovsky

In this paper, we present a meta-analysis of several Web content extraction algorithms, and make recommendations for the future of content extraction on the Web. First, we find that nearly all Web content extractors do not consider a very…

Information Retrieval · Computer Science 2015-08-19 Tim Weninger , Rodrigo Palacios , Valter Crescenzi , Thomas Gottron , Paolo Merialdo

Popular web pages are archived frequently, which makes it difficult to visualize the progression of the site through the years at web archives. The What Did It Look Like (WDILL) Twitter bot shows web page transitions by creating a timelapse…

Digital Libraries · Computer Science 2021-04-30 Dhruv Patel , Alexander C. Nwala , Michael L. Nelson , Michele C. Weigle

As the amount of data on the World Wide Web continues to grow exponentially, access to semantically structured information remains limited. The Semantic Web has emerged as a solution to enhance the machine-readability of data, making it…

Digital Libraries · Computer Science 2023-06-21 Muhammad Zohaib

Network Error Logging helps web server operators detect operational problems in real-time to provide fast and reliable services. HTTP Archive provides detail information of historical data on HTTP requests. This paper leverages the data and…

Networking and Internet Architecture · Computer Science 2023-05-03 Kamil Jeřábek , Libor Polčák

The definition of scholarly content has expanded to include the data and source code that contribute to a publication. While major archiving efforts to preserve conventional scholarly content, typically in PDFs (e.g., LOCKSS, CLOCKSS,…

Digital Libraries · Computer Science 2022-08-10 Emily Escamilla , Martin Klein , Talya Cooper , Vicky Rampin , Michele C. Weigle , Michael L. Nelson

Using APIs to develop software applications is the norm. APIs help developers to build applications faster as they do not need to reinvent the wheel. It is therefore important for developers to understand the APIs that they plan to use.…

Software Engineering · Computer Science 2023-04-06 Ferdian Thung , Kisub Kim , Ting Zhang , Ivana Clairine Irsan , Ratnadira Widyasari , Zhou Yang , David Lo

The Web publishing paradigm of Linked Data has been gaining traction in the cultural heritage sector: libraries, archives and museums. At first glance, the principles of Linked Data seem simple enough. However experienced Web developers,…

Digital Libraries · Computer Science 2013-06-21 Ed Summers , Dorothea Salo

We conducted a preliminary field study to understand the current state of personal digital archiving in practice. Our aim is to design a service for the long-term storage, preservation, and access of digital belongings by examining how…

Digital Libraries · Computer Science 2007-05-23 Catherine C. Marshall , Sara Bly , Francoise Brun-Cottan

Data archives are an important source of high quality data in many fields, making them ideal sites to study data reuse. By studying data reuse through citation networks, we are able to learn how hidden research communities - those that use…

Digital Libraries · Computer Science 2022-10-21 Sara Lafia , Lizhou Fan , Andrea Thomer , Libby Hemphill

As part of their efforts to consistently and reliably identify and preserve the archival records of their organizations, archivists are trying to address the issue of appraising and collecting World Wide Web documents. This paper discusses…

History and Philosophy of Physics · Physics 2007-05-23 Jean M. Deken

The creation of open archives i.e. archives where access is regulated by open licensing models (content, source, data), should be seen as part of a broader socio-economic phenomenon that finds legal expression in specific organizational and…

Digital Libraries · Computer Science 2011-09-06 Prodromos Tsiavos , Petros Stefaneas