English
Related papers

Related papers: The Ephemeral Web and the Case for Proactive Archi…

200 papers

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web…

Digital Libraries · Computer Science 2021-01-26 Shawn M. Jones , Michele C. Weigle , Michael L. Nelson

The web of data has brought forth the need to preserve and sustain evolving information within linked datasets; however, a basic requirement of data preservation is the maintenance of the datasets' structural characteristics as well. As…

Web archives capture the history of the Web and are therefore an important source to study how societal developments have been reflected on the Web. However, the large size of Web archives and their temporal nature pose many challenges to…

Digital Libraries · Computer Science 2016-12-20 Gerhard Gossen , Elena Demidova , Thomas Risse

Recent advances of preservation technologies have led to an increasing number of Web archive systems and collections. These collections are valuable to explore the past of the Web, but their value can only be uncovered with effective access…

Information Retrieval · Computer Science 2017-01-17 Khoi Duy Vo , Tuan Tran , Tu Ngoc Nguyen , Xiaofei Zhu , Wolfgang Nejdl

We describe challenges related to web archiving, replaying archived web resources, and verifying their authenticity. We show that Web Packaging has significant potential to help address these challenges and identify areas in which changes…

Networking and Internet Architecture · Computer Science 2019-06-18 Sawood Alam , Michele C. Weigle , Michael L. Nelson , Martin Klein , Herbert Van de Sompel

We conducted a preliminary field study to understand the current state of personal digital archiving in practice. Our aim is to design a service for the long-term storage, preservation, and access of digital belongings by examining how…

Digital Libraries · Computer Science 2007-05-23 Catherine C. Marshall , Sara Bly , Francoise Brun-Cottan

The Windsor Study Group on Digital Archiving was commissioned to recommend strategies, policies, and technologies necessary for ensuring the integrity and longevity of electronic publications. The goal of this work is to inform institutions…

Digital Libraries · Computer Science 2014-04-01 Sandy Payette

Backup or preservation of websites is often not considered until after a catastrophic event has occurred. In the face of complete website loss, "lazy" webmasters or concerned third parties may be able to recover some of their website from…

Information Retrieval · Computer Science 2011-11-09 Frank McCown , Joan A. Smith , Michael L. Nelson , Johan Bollen

The Web publishing paradigm of Linked Data has been gaining traction in the cultural heritage sector: libraries, archives and museums. At first glance, the principles of Linked Data seem simple enough. However experienced Web developers,…

Digital Libraries · Computer Science 2013-06-21 Ed Summers , Dorothea Salo

Web archives are large longitudinal collections that store webpages from the past, which might be missing on the current live Web. Consequently, temporal search over such collections is essential for finding prominent missing webpages and…

Information Retrieval · Computer Science 2017-02-07 Helge Holzmann , Wolfgang Nejdl , Avishek Anand

Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources.…

Digital Libraries · Computer Science 2017-05-18 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

The vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose…

Infrastructures are not inherently durable or fragile, yet all are fragile over the long term. Durability requires care and maintenance of individual components and the links between them. Astronomy is an ideal domain in which to study…

Digital Libraries · Computer Science 2016-11-02 Christine L. Borgman , Peter T. Darch , Ashley E. Sands , Milena S. Golshan

Archiving the web is socially and culturally critical, but presents problems of scale. The Internet Archive's Wayback Machine can replay captured web pages as they existed at a certain point in time, but it has limited ability to provide…

Information Retrieval · Computer Science 2013-06-12 Ahmed AlSum , Michael L. Nelson

Archival institutions and programs worldwide work to ensure that the records of governments, organizations, communities, and individuals are preserved for future generations as cultural heritage, as sources of rights, and as vehicles for…

Computers and Society · Computer Science 2022-03-15 Emanuele Frontoni , Marina Paolanti , Tracey P. Lauriault , Michael Stiber , Luciana Duranti , Abdul-Mageed Muhammad

The current Web has no general mechanisms to make digital artifacts --- such as datasets, code, texts, and images --- verifiable and permanent. For digital artifacts that are supposed to be immutable, there is moreover no commonly accepted…

Cryptography and Security · Computer Science 2015-07-08 Tobias Kuhn , Michel Dumontier

Web Engineering is the application of systematic, disciplined and quantifiable approaches to development, operation, and maintenance of Web-based applications. It is both a pro-active approach and a growing collection of theoretical and…

Software Engineering · Computer Science 2007-05-23 Yogesh Deshpande , San Murugesan , Athula Ginige , Steve Hansen , Daniel Schwabe , Martin Gaedke , Bebo White

Web archives, a key area of digital preservation, meet the needs of journalists, social scientists, historians, and government organizations. The use cases for these groups often require that they guide the archiving process themselves,…

Digital Libraries · Computer Science 2021-01-26 Shawn M. Jones , Alexander Nwala , Michele C. Weigle , Michael L. Nelson

We document strategies and lessons learned from sampling the web by collecting 27.3 million URLs with 3.8 billion archived pages spanning 26 years (1996-2021) from the Internet Archive's (IA) Wayback Machine. Our goal is to revisit…

Digital Libraries · Computer Science 2025-07-22 Kritika Garg , Sawood Alam , Dietrich Ayala , Mark Graham , Michele C. Weigle , Michael L. Nelson

The expanding need for an open information sharing infrastructure to promote scholarly communication led to the pioneering establishment of arXiv.org, now maintained by the Cornell University Library. To be sustainable, the repository…

Digital Libraries · Computer Science 2015-10-02 Oya Y. Rieger