中文
相关论文

相关论文: ArchiveWeb: collaboratively extending and explorin…

200 篇论文

Curated web archive collections contain focused digital contents which are collected by archiving organizations to provide a representative sample covering specific topics and events to preserve them for future exploration and analysis. In…

数字图书馆 · 计算机科学 2017-02-02 Zeon Trevor Fernando , Ivana Marenzi , Wolfgang Nejdl , Rishita Kalyani

Web archives capture the history of the Web and are therefore an important source to study how societal developments have been reflected on the Web. However, the large size of Web archives and their temporal nature pose many challenges to…

数字图书馆 · 计算机科学 2016-12-20 Gerhard Gossen , Elena Demidova , Thomas Risse

Web archive analytics is the exploitation of publicly accessible web pages and their evolution for research purposes -- to the extent organizationally possible for researchers. In order to better understand the complexity of this task, the…

数字图书馆 · 计算机科学 2021-07-05 Michael Völske , Janek Bevendorff , Johannes Kiesel , Benno Stein , Maik Fröbe , Matthias Hagen , Martin Potthast

Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources.…

数字图书馆 · 计算机科学 2017-05-18 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

Web archiving is the process of collecting portions of the Web to ensure that the information is preserved for future exploitation. However, despite the increasing number of web archives worldwide, the absence of efficient and meaningful…

数字图书馆 · 计算机科学 2018-10-25 Pavlos Fafalios , Helge Holzmann , Vaibhav Kasturia , Wolfgang Nejdl

The Archives Unleashed project aims to improve scholarly access to web archives through a multi-pronged strategy involving tool creation, process modeling, and community building - all proceeding concurrently in mutually-reinforcing…

数字图书馆 · 计算机科学 2020-01-16 Nick Ruest , Jimmy Lin , Ian Milligan , Samantha Fritz

Event collections are frequently built by crawling the live web on the basis of seed URIs nominated by human experts. Focused web crawling is a technique where the crawler is guided by reference content pertaining to the event. Given the…

数字图书馆 · 计算机科学 2018-04-06 Martin Klein , Lyudmila Balakireva , Herbert Van de Sompel

As web archives' holdings grow, archivists subdivide them into collections so they are easier to understand and manage. In this work, we review the collection structures of eight web archive platforms: : Archive-It, Conifer, the Croatian…

Web archives, a key area of digital preservation, meet the needs of journalists, social scientists, historians, and government organizations. The use cases for these groups often require that they guide the archiving process themselves,…

数字图书馆 · 计算机科学 2021-01-26 Shawn M. Jones , Alexander Nwala , Michele C. Weigle , Michael L. Nelson

Our cultural discourse is increasingly carried in the web. With the initial emergence of the web many years ago, there was a period where conventional mediums (e.g., music, movies, books, scholarly publications) were primary and the web was…

数字图书馆 · 计算机科学 2012-09-13 Michael L. Nelson

The web of data has brought forth the need to preserve and sustain evolving information within linked datasets; however, a basic requirement of data preservation is the maintenance of the datasets' structural characteristics as well. As…

The vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose…

Web archives are a valuable resource for researchers of various disciplines. However, to use them as a scholarly source, researchers require a tool that provides efficient access to Web archive data for extraction and derivation of smaller…

数字图书馆 · 计算机科学 2017-02-06 Helge Holzmann , Vinay Goel , Avishek Anand

We describe challenges related to web archiving, replaying archived web resources, and verifying their authenticity. We show that Web Packaging has significant potential to help address these challenges and identify areas in which changes…

网络与互联网体系结构 · 计算机科学 2019-06-18 Sawood Alam , Michele C. Weigle , Michael L. Nelson , Martin Klein , Herbert Van de Sompel

Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in…

数字图书馆 · 计算机科学 2017-07-31 Gerhard Gossen , Elena Demidova , Thomas Risse

The historical, cultural, and intellectual importance of archiving the web has been widely recognized. Today, all countries with high Internet penetration rate have established high-profile archiving initiatives to crawl and archive the…

数字图书馆 · 计算机科学 2013-08-13 Zhiwu Xie , Herbert Van de Sompel , Jinyang Liu , Johann van Reenen , Ramiro Jordan

As web technologies evolve, web archivists work to keep up so that our digital history is preserved. Recent advances in web technologies have introduced client-side executed scripts that load data without a referential identifier or that…

数字图书馆 · 计算机科学 2019-05-17 Mat Kelly , Justin F. Brunelle , Michele C. Weigle , Michael L. Nelson

The field of web archiving provides a unique mix of human and automated agents collaborating to achieve the preservation of the web. Centuries old theories of archival appraisal are being transplanted into the sociotechnical environment of…

数字图书馆 · 计算机科学 2016-11-09 Ed Summers , Ricardo Punzalan

As the Distributed Collection Manager's work on building tools to support users maintaining collections of changing web-based resources has progressed, questions about the characteristics of people's collections of web pages have arisen.…

数字图书馆 · 计算机科学 2011-01-05 Paul Logasa Bogen , Frank Shipman , Richard Furuta

The Data Web refers to the vast and rapidly increasing quantity of scientific, corporate, government and crowd-sourced data published in the form of Linked Open Data, which encourages the uniform representation of heterogeneous data items…

‹ 上一页 1 2 3 10 下一页 ›