English
Related papers

Related papers: Caching HTTP 404 Responses Eliminates Unnecessary …

200 papers

Certain HTTP Cookies on certain sites can be a source of content bias in archival crawls. Accommodating Cookies at crawl time, but not utilizing them at replay time may cause cookie violations, resulting in defaced composite mementos that…

Digital Libraries · Computer Science 2019-06-18 Sawood Alam , Plinio Vargas , Michele C. Weigle , Michael L. Nelson

Although web advertisements represent an inimitable part of digital cultural heritage, serious archiving and replay challenges persist. To explore these challenges, we created a dataset of 279 archived ads. We encountered five problems in…

Digital Libraries · Computer Science 2025-09-24 Travis Reid , Alex H. Poole , Hyung Wook Choi , Christopher Rauch , Mat Kelly , Michael L. Nelson , Michele C. Weigle

Web archiving services play an increasingly important role in today's information ecosystem, by ensuring the continuing availability of information, or by deliberately caching content that might get deleted or removed. Among these, the…

Computers and Society · Computer Science 2018-04-10 Savvas Zannettou , Jeremy Blackburn , Emiliano De Cristofaro , Michael Sirivianos , Gianluca Stringhini

When a user requests a web page from a web archive, the user will typically either get an HTTP 200 if the page is available, or an HTTP 404 if the web page has not been archived. This is because web archives are typically accessed by URI…

Digital Libraries · Computer Science 2019-08-09 Lulwah M. Alkwai , Michael L. Nelson , Michele C. Weigle

The Web is ephemeral. Many resources have representations that change over time, and many of those representations are lost forever. A lucky few manage to reappear as archived resources that carry their own URIs. For example, some content…

We want to make web archiving entertaining so that it can be enjoyed like a spectator sport. To this end, we have been working on a proof of concept that involves gamification of the web archiving process and integrating video games and web…

Digital Libraries · Computer Science 2022-11-07 Travis Reid , Michael L. Nelson , Michele C. Weigle

The historical, cultural, and intellectual importance of archiving the web has been widely recognized. Today, all countries with high Internet penetration rate have established high-profile archiving initiatives to crawl and archive the…

Digital Libraries · Computer Science 2013-08-13 Zhiwu Xie , Herbert Van de Sompel , Jinyang Liu , Johann van Reenen , Ramiro Jordan

Many web sites are transitioning how they construct their pages. The conventional model is where the content is embedded server-side in the HTML and returned to the client in an HTTP response. Increasingly, sites are moving to a model where…

Digital Libraries · Computer Science 2023-05-03 Michele C. Weigle , Michael L. Nelson , Sawood Alam , Mark Graham

As web technologies evolve, web archivists work to keep up so that our digital history is preserved. Recent advances in web technologies have introduced client-side executed scripts that load data without a referential identifier or that…

Digital Libraries · Computer Science 2019-05-17 Mat Kelly , Justin F. Brunelle , Michele C. Weigle , Michael L. Nelson

Performance in web applications is a key aspect of user experience and system scalability. Among the different techniques used to improve web application performance, caching has been widely used. While caching has been widely explored in…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-09 Mohammad Umar , Bharat Tripathi

HTTP/2 supersedes HTTP/1.1 to tackle the performance challenges of the modern Web. A highly anticipated feature is Server Push, enabling servers to send data without explicit client requests, thus potentially saving time. Although…

Networking and Internet Architecture · Computer Science 2018-10-15 Torsten Zimmermann , Benedikt Wolters , Oliver Hohlfeld , Klaus Wehrle

Retrieval and content management are assumed to be mutually exclusive. In this paper we suggest that they need not be so. In the usual information retrieval scenario, some information about queries leading to a website (due to `hits' or…

Information Retrieval · Computer Science 2019-08-29 C Ravindranath Chowdary , Anil Kumar Singh , Anil Nelakanti

Modern web pages have complex structures comprised of up to hundreds of different resources, such as scripts and images. Server push is an HTTP/2 feature enabling servers to preemptively send resources to clients before they realize they…

Networking and Internet Architecture · Computer Science 2022-07-14 Rui Meireles , Junrui Liu , Peter Steenkiste

HTTP/2 and HTTP/3 avoid concurrent connections but instead multiplex requests over a single connection. Besides enabling new features, this reduces overhead and enables fair bandwidth sharing. Redundant connections should hence be a story…

Networking and Internet Architecture · Computer Science 2021-10-28 Constantin Sander , Leo Blöcher , Klaus Wehrle , Jan Rüth

We describe challenges related to web archiving, replaying archived web resources, and verifying their authenticity. We show that Web Packaging has significant potential to help address these challenges and identify areas in which changes…

Networking and Internet Architecture · Computer Science 2019-06-18 Sawood Alam , Michele C. Weigle , Michael L. Nelson , Martin Klein , Herbert Van de Sompel

Two main trends in today's internet are of major interest for video streaming services: most content delivery platforms coincide towards using adaptive video streaming over HTTP and new network architectures allowing caching at intermediate…

Networking and Internet Architecture · Computer Science 2013-07-04 Reinhard Grandl , Kai Su , Cedric Westphal

Web archives are a historically valuable source of information. In some respects, web archives are the only record of the evolution of human society in the last two decades. They preserve a mix of personal and collective memories, the…

Digital Libraries · Computer Science 2021-08-04 Miguel Costa

Modern web applications rely heavily on client-side API calls to fetch data, render content, and communicate with backend services. However, the quality of these network interactions (redundant requests, missing cache headers, oversized…

Software Engineering · Computer Science 2026-02-19 Ali Hassaan Mughal , Muhammad Bilal , Noor Fatima

Web caches play a crucial role in web performance and scalability. However, detecting cached responses is challenging when web servers do not reliably communicate the cache status through standardized headers. This paper presents a novel…

Cryptography and Security · Computer Science 2024-07-24 Matteo Golinelli , Bruno Crispo

Text extraction from web pages has many applications, including web crawling optimization and document clustering. Though much has been written about the acquisition of content from live web pages, content acquisition of archived web pages,…

Digital Libraries · Computer Science 2016-02-24 Shawn M. Jones , Harihar Shankar
‹ Prev 1 2 3 10 Next ›