English
Related papers

Related papers: Longitudinal Sampling of URLs From the Wayback Mac…

200 papers

The Environmental Governance and Data Initiative (EDGI) regularly crawled US federal environmental websites between 2016 and 2020 to capture changes between two presidential administrations. However, because it does not include the previous…

Digital Libraries · Computer Science 2025-06-02 Lesley Frew , Michael L. Nelson , Michele C. Weigle

The Web has been around and maturing for 25 years. The popular websites of today have undergone vast changes during this period, with a few being there almost since the beginning and many new ones becoming popular over the years. This makes…

Digital Libraries · Computer Science 2017-02-07 Helge Holzmann , Wolfgang Nejdl , Avishek Anand

Text extraction from web pages has many applications, including web crawling optimization and document clustering. Though much has been written about the acquisition of content from live web pages, content acquisition of archived web pages,…

Digital Libraries · Computer Science 2016-02-24 Shawn M. Jones , Harihar Shankar

The dark web hosts a dynamic ecosystem of cybercrime forums and marketplaces that adapt to law enforcement pressure, technological change, and economic incentives. Prior research has extracted cyber threat intelligence from these platforms…

Cryptography and Security · Computer Science 2026-05-18 Roy Ricaldi , Maximilian Schafer , Philipp Zech , Luca Allodi , Raffaela Groner , Irdin Pekaric

Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources.…

Digital Libraries · Computer Science 2017-05-18 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

Internet-wide scans are a common active measurement approach to study the Internet, e.g., studying security properties or protocol adoption. They involve probing large address ranges (IPv4 or parts of IPv6) for specific ports or protocols.…

Networking and Internet Architecture · Computer Science 2019-01-23 Jan Rüth , Torsten Zimmermann , Oliver Hohlfeld

Understanding how people interact with the web is key for a variety of applications, e.g., from the design of effective web pages to the definition of successful online marketing campaigns. Browsing behavior has been traditionally…

Computers and Society · Computer Science 2021-05-05 Luca Vassio , Idilio Drago , Marco Mellia , Zied Ben Houidi , Mohamed Lamine Lamali

Web archives are large longitudinal collections that store webpages from the past, which might be missing on the current live Web. Consequently, temporal search over such collections is essential for finding prominent missing webpages and…

Information Retrieval · Computer Science 2017-02-07 Helge Holzmann , Wolfgang Nejdl , Avishek Anand

In this work, we study how URL extraction results depend on input format. We compiled a pilot dataset by extracting URLs from 10 arXiv papers and used the same heuristic method to extract URLs from four formats derived from the PDF files or…

Digital Libraries · Computer Science 2025-09-08 Rochana R. Obadage , Lamia Salsabil , Sawood Alam , Bipasha Banarjee , William A. Ingram , Edward A. Fox , Jian Wu

Upon replay, JavaScript on archived web pages can generate recurring HTTP requests that lead to unnecessary traffic to the web archive. In one example, an archived page averaged more than 1000 requests per minute. These requests are not…

Networking and Internet Architecture · Computer Science 2022-12-02 Kritika Garg , Himarsha R. Jayanetti , Sawood Alam , Michele C. Weigle , Michael L. Nelson

A large number of URLs are made public by various platforms for security analysis, archiving, and paste sharing -- such as VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt. These services may unintentionally expose…

Cryptography and Security · Computer Science 2026-02-26 Tarek Ramadan , AbdelRahman Abdou , Mohammad Mannan , Amr Youssef

Automated analysis of privacy policies has proved a fruitful research direction, with developments such as automated policy summarization, question answering systems, and compliance detection. Prior research has been limited to analysis of…

Computers and Society · Computer Science 2021-07-22 Ryan Amos , Gunes Acar , Eli Lucherini , Mihir Kshirsagar , Arvind Narayanan , Jonathan Mayer

Web archives capture the history of the Web and are therefore an important source to study how societal developments have been reflected on the Web. However, the large size of Web archives and their temporal nature pose many challenges to…

Digital Libraries · Computer Science 2016-12-20 Gerhard Gossen , Elena Demidova , Thomas Risse

Social graph construction from various sources has been of interest to researchers due to its application potential and the broad range of technical challenges involved. The World Wide Web provides a huge amount of continuously updated data…

Social and Information Networks · Computer Science 2017-01-13 Miroslav Shaltev , Jan-Hendrik Zab , Philipp Kemkes , Stefan Siersdorfer , Sergej Zerr

Web archives, a key area of digital preservation, meet the needs of journalists, social scientists, historians, and government organizations. The use cases for these groups often require that they guide the archiving process themselves,…

Digital Libraries · Computer Science 2021-01-26 Shawn M. Jones , Alexander Nwala , Michele C. Weigle , Michael L. Nelson

As conventional storage density reaches its physical limits, the cost of a gigabyte of storage is no longer plummeting, but rather has remained mostly flat for the past decade. Meanwhile, file sizes continue to grow, leading to ever fuller…

Operating Systems · Computer Science 2025-03-31 Kevin Saric , Gowri Sankar Ramachandran , Raja Jurdak , Surya Nepal

The historical, cultural, and intellectual importance of archiving the web has been widely recognized. Today, all countries with high Internet penetration rate have established high-profile archiving initiatives to crawl and archive the…

Digital Libraries · Computer Science 2013-08-13 Zhiwu Xie , Herbert Van de Sompel , Jinyang Liu , Johann van Reenen , Ramiro Jordan

Long-term Web archives comprise Web documents gathered over longer time periods and can easily reach hundreds of terabytes in size. Semantic annotations such as named entities can facilitate intelligent access to the Web archive data.…

Information Retrieval · Computer Science 2017-02-03 Tarcisio Souza , Elena Demidova , Thomas Risse , Helge Holzmann , Gerhard Gossen , Julian Szymanski

Is software obsolescence a significant risk? To explore this issue, we analysed a corpus of over 2.5 billion resources corresponding to the UK Web domain, as crawled between 1996 and 2010. Using the DROID and Apache Tika identification…

Digital Libraries · Computer Science 2012-10-08 Andrew N. Jackson

When a user views an archived page using the archive's user interface (UI), the user selects a datetime to view from a list. The archived web page, if available, is then displayed. From this display, the web archive UI attempts to simulate…

Digital Libraries · Computer Science 2013-09-24 Scott G. Ainsworth , Michael L. Nelson