中文
相关论文

相关论文: Rules of Acquisition for Mementos and Their Conten…

200 篇论文

We document the creation of a data set of 16,627 archived web pages, or mementos, of 3,698 unique live web URIs (Uniform Resource Identifiers) from 17 public web archives. We used four different methods to collect the dataset. First, we…

数字图书馆 · 计算机科学 2019-05-13 Mohamed Aturban , Michael L. Nelson , Michele C. Weigle , Martin Klein , Herbert Van de Sompel

The Web is ephemeral. Many resources have representations that change over time, and many of those representations are lost forever. A lucky few manage to reappear as archived resources that carry their own URIs. For example, some content…

Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources.…

数字图书馆 · 计算机科学 2017-05-18 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

Memento aggregators enable users to query multiple web archives for captures of a URI in time through a single HTTP endpoint. While this one-to-many access point is useful for researchers and end-users, aggregators are in a position to…

数字图书馆 · 计算机科学 2023-01-10 Mat Kelly

Event collections are frequently built by crawling the live web on the basis of seed URIs nominated by human experts. Focused web crawling is a technique where the crawler is guided by reference content pertaining to the event. Given the…

数字图书馆 · 计算机科学 2018-04-06 Martin Klein , Lyudmila Balakireva , Herbert Van de Sompel

To perform a longitudinal investigation of web archives and detecting variations and changes replaying individual archived pages, or mementos, we created a sample of 16,627 mementos from 17 public web archives. Over the course of our…

数字图书馆 · 计算机科学 2021-08-16 Mohamed Aturban , Michael L. Nelson , Michele C. Weigle

As web technologies evolve, web archivists work to keep up so that our digital history is preserved. Recent advances in web technologies have introduced client-side executed scripts that load data without a referential identifier or that…

数字图书馆 · 计算机科学 2019-05-17 Mat Kelly , Justin F. Brunelle , Michele C. Weigle , Michael L. Nelson

Personal and private Web archives are proliferating due to the increase in the tools to create them and the realization that Internet Archive and other public Web archives are unable to capture personalized (e.g., Facebook) and private…

数字图书馆 · 计算机科学 2018-06-05 Mat Kelly , Michael L. Nelson , Michele C. Weigle

For traditional library collections, archivists can select a representative sample from a collection and display it in a featured physical or digital library space. Web archive collections may consist of thousands of archived pages, or…

数字图书馆 · 计算机科学 2021-03-23 Shawn M. Jones , Martin Klein , Michele C. Weigle , Michael L. Nelson

Web browsers provide a user-friendly means of navigating the web. Users rely on their web browser to provide information about the websites they are visiting, such as the security state. Browsers also provide a user interface (UI) with…

数字图书馆 · 计算机科学 2021-04-28 Abigail Mabe

The World Wide Web caters to the needs of billions of users in heterogeneous groups. Each user accessing the World Wide Web might have his / her own specific interest and would expect the web to respond to the specific requirements. The…

信息检索 · 计算机科学 2017-11-22 K. S. Kuppusamy , G. Aghila

Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in…

数字图书馆 · 计算机科学 2017-07-31 Gerhard Gossen , Elena Demidova , Thomas Risse

In this paper we present the results of a study into the persistence and availability of web resources referenced from papers in scholarly repositories. Two repositories with different characteristics, arXiv and the UNT digital library, are…

数字图书馆 · 计算机科学 2011-05-18 Robert Sanderson , Mark Phillips , Herbert Van de Sompel

In this paper, we present a meta-analysis of several Web content extraction algorithms, and make recommendations for the future of content extraction on the Web. First, we find that nearly all Web content extractors do not consider a very…

信息检索 · 计算机科学 2015-08-19 Tim Weninger , Rodrigo Palacios , Valter Crescenzi , Thomas Gottron , Paolo Merialdo

The information available on web pages mostly contains semi-structured text documents which are represented either in XML, or HTML, or XHTML format that lacks formatted document structure. The document does not discriminate between the text…

信息检索 · 计算机科学 2014-03-11 Sandeep Sirsat

Recent advances of preservation technologies have led to an increasing number of Web archive systems and collections. These collections are valuable to explore the past of the Web, but their value can only be uncovered with effective access…

信息检索 · 计算机科学 2017-01-17 Khoi Duy Vo , Tuan Tran , Tu Ngoc Nguyen , Xiaofei Zhu , Wolfgang Nejdl

Events of various kinds are mentioned and discussed in text documents, whether they are books, news articles, blogs or microblog feeds. The paper starts by giving an overview of how events are treated in linguistics and philosophy. We…

计算与语言 · 计算机科学 2016-01-18 Jugal Kalita

In this paper, we focused on the problem of extracting information from web pages containing many records, a task of growing importance in the era of massive web data. Recently, the development of neural network methods has improved the…

计算与语言 · 计算机科学 2025-02-21 Alexander Kustenkov , Maksim Varlamov , Alexander Yatskov

Web archives capture the history of the Web and are therefore an important source to study how societal developments have been reflected on the Web. However, the large size of Web archives and their temporal nature pose many challenges to…

数字图书馆 · 计算机科学 2016-12-20 Gerhard Gossen , Elena Demidova , Thomas Risse

Web images come in hand with valuable contextual information. Although this information has long been mined for various uses such as image annotation, clustering of images, inference of image semantic content, etc., insufficient attention…

多媒体 · 计算机科学 2020-05-21 F. Fauzi , H. J. Long , M. Belkhatir
‹ 上一页 1 2 3 10 下一页 ›