English
Related papers

Related papers: Impact of URI Canonicalization on Memento Count

200 papers

To perform a longitudinal investigation of web archives and detecting variations and changes replaying individual archived pages, or mementos, we created a sample of 16,627 mementos from 17 public web archives. Over the course of our…

Digital Libraries · Computer Science 2021-08-16 Mohamed Aturban , Michael L. Nelson , Michele C. Weigle

In this paper we present the results of a study into the persistence and availability of web resources referenced from papers in scholarly repositories. Two repositories with different characteristics, arXiv and the UNT digital library, are…

Digital Libraries · Computer Science 2011-05-18 Robert Sanderson , Mark Phillips , Herbert Van de Sompel

The Web is ephemeral. Many resources have representations that change over time, and many of those representations are lost forever. A lucky few manage to reappear as archived resources that carry their own URIs. For example, some content…

As defined by the Memento Framework, TimeMaps are ma-chine-readable lists of time-specific copies -- called "mementos" -- of an archived original resource. In theory, as an archive acquires additional mementos over time, a TimeMap should be…

Digital Libraries · Computer Science 2013-07-23 Justin F. Brunelle , Michael L. Nelson

URI redirections are integral to web management, supporting structural changes, SEO optimization, and security. However, their complexities affect usability, SEO performance, and digital preservation. This study analyzed 11 million unique…

Digital Libraries · Computer Science 2025-07-30 Kritika Garg , Sawood Alam , Dietrich Ayala , Michele C. Weigle , Michael L. Nelson

We document the creation of a data set of 16,627 archived web pages, or mementos, of 3,698 unique live web URIs (Uniform Resource Identifiers) from 17 public web archives. We used four different methods to collect the dataset. First, we…

Digital Libraries · Computer Science 2019-05-13 Mohamed Aturban , Michael L. Nelson , Michele C. Weigle , Martin Klein , Herbert Van de Sompel

The Memento aggregator currently polls every known public web archive when serving a request for an archived web page, even though some web archives focus on only specific domains and ignore the others. Similar to query routing in…

Digital Libraries · Computer Science 2013-09-17 Ahmed AlSum , Michele C. Weigle , Michael L. Nelson , Herbert Van de Sompel

In this work we propose MementoMap, a flexible and adaptive framework to efficiently summarize holdings of a web archive. We described a simple, yet extensible, file format suitable for MementoMap. We used the complete index of the…

Digital Libraries · Computer Science 2019-05-30 Sawood Alam , Michele C. Weigle , Michael L. Nelson , Fernando Melo , Daniel Bicho , Daniel Gomes

Text extraction from web pages has many applications, including web crawling optimization and document clustering. Though much has been written about the acquisition of content from live web pages, content acquisition of archived web pages,…

Digital Libraries · Computer Science 2016-02-24 Shawn M. Jones , Harihar Shankar

As Digital Libraries (DL) become more aligned with the web architecture, their functional components need to be fundamentally rethought in terms of URIs and HTTP. Annotation, a core scholarly activity enabled by many DL solutions, exhibits…

Digital Libraries · Computer Science 2010-03-22 Robert Sanderson , Herbert Van de Sompel

Modeling of long history data suffers from long-context window attention dilution, system efficiency and catastrophic forgetting problems, where naive linear scaling approach like LastN would fail. We introduce Memento, a personalized…

Large Language Models (LLMs) demonstrate remarkable capabilities in question answering (QA), but metrics for assessing their reliance on memorization versus retrieval remain underdeveloped. Moreover, while finetuned models are…

Machine Learning · Computer Science 2025-06-17 Peter Carragher , Abhinand Jha , R Raghav , Kathleen M. Carley

The Memento protocol provides a uniform approach to query individual web archives. Soon after its emergence, Memento Aggregator infrastructure was introduced that supports querying across multiple archives simultaneously. An Aggregator…

Digital Libraries · Computer Science 2016-06-30 Nicolas J. Bornand , Lyudmila Balakireva , Herbert Van de Sompel

Memory-augmented neural networks (MANNs) refer to a class of neural network models equipped with external memory (such as neural Turing machines and memory networks). These neural networks outperform conventional recurrent neural networks…

Machine Learning · Computer Science 2017-11-13 Seongsik Park , Seijoon Kim , Seil Lee , Ho Bae , Sungroh Yoon

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web…

Digital Libraries · Computer Science 2021-01-26 Shawn M. Jones , Michele C. Weigle , Michael L. Nelson

Large language models and deep research agents supply citation URLs to support their claims, yet the reliability of these citations has not been systematically measured. We address six research questions about citation URL validity using 10…

Computation and Language · Computer Science 2026-04-06 Delip Rao , Eric Wong , Chris Callison-Burch

Generative models based on the Flow Matching objective, particularly Rectified Flow, have emerged as a dominant paradigm for efficient, high-fidelity image synthesis. However, while existing research heavily prioritizes generation quality…

Machine Learning · Computer Science 2026-03-17 Mingxing Rao , Daniel Moyer

Most archived HTML pages embed other web resources, such as images and stylesheets. Playback of the archived web pages typically provides only the capture date (or Memento-Datetime) of the root resource and not the Memento-Datetime of the…

Digital Libraries · Computer Science 2014-10-07 Scott G. Ainsworth , Michael L. Nelson , Herbert Van de Sompel

Although the Internet Archive's Wayback Machine is the largest and most well-known web archive, there have been a number of public web archives that have emerged in the last several years. With varying resources, audiences and collection…

Digital Libraries · Computer Science 2013-01-08 Scott G. Ainsworth , Ahmed AlSum , Hany SalahEldeen , Michele C. Weigle , Michael L. Nelson

Recent approaches to goal and plan recognition using classical planning domains have achieved state of the art results in terms of both recognition time and accuracy by using heuristics based on planning landmarks. To achieve such fast…

Artificial Intelligence · Computer Science 2020-05-07 Kin Max Piamolini Gusmão , Ramon Fraga Pereira , Felipe Meneguzzi
‹ Prev 1 2 3 10 Next ›