中文
相关论文

相关论文: The Many Shapes of Archive-It

200 篇论文

As web archives' holdings grow, archivists subdivide them into collections so they are easier to understand and manage. In this work, we review the collection structures of eight web archive platforms: : Archive-It, Conifer, the Croatian…

Web archiving is the process of collecting portions of the Web to ensure that the information is preserved for future exploitation. However, despite the increasing number of web archives worldwide, the absence of efficient and meaningful…

数字图书馆 · 计算机科学 2018-10-25 Pavlos Fafalios , Helge Holzmann , Vaibhav Kasturia , Wolfgang Nejdl

Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources.…

数字图书馆 · 计算机科学 2017-05-18 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

Curated web archive collections contain focused digital content which is collected by archiving organizations, groups, and individuals to provide a representative sample covering specific topics and events to preserve them for future…

数字图书馆 · 计算机科学 2017-02-03 Zeon Trevor Fernando , Ivana Marenzi , Wolfgang Nejdl

The field of web archiving provides a unique mix of human and automated agents collaborating to achieve the preservation of the web. Centuries old theories of archival appraisal are being transplanted into the sociotechnical environment of…

数字图书馆 · 计算机科学 2016-11-09 Ed Summers , Ricardo Punzalan

Curated web archive collections contain focused digital contents which are collected by archiving organizations to provide a representative sample covering specific topics and events to preserve them for future exploration and analysis. In…

数字图书馆 · 计算机科学 2017-02-02 Zeon Trevor Fernando , Ivana Marenzi , Wolfgang Nejdl , Rishita Kalyani

Web archiving services play an increasingly important role in today's information ecosystem, by ensuring the continuing availability of information, or by deliberately caching content that might get deleted or removed. Among these, the…

计算机与社会 · 计算机科学 2018-04-10 Savvas Zannettou , Jeremy Blackburn , Emiliano De Cristofaro , Michael Sirivianos , Gianluca Stringhini

Web archive analytics is the exploitation of publicly accessible web pages and their evolution for research purposes -- to the extent organizationally possible for researchers. In order to better understand the complexity of this task, the…

数字图书馆 · 计算机科学 2021-07-05 Michael Völske , Janek Bevendorff , Johannes Kiesel , Benno Stein , Maik Fröbe , Matthias Hagen , Martin Potthast

Recent advances of preservation technologies have led to an increasing number of Web archive systems and collections. These collections are valuable to explore the past of the Web, but their value can only be uncovered with effective access…

信息检索 · 计算机科学 2017-01-17 Khoi Duy Vo , Tuan Tran , Tu Ngoc Nguyen , Xiaofei Zhu , Wolfgang Nejdl

Web archives capture the history of the Web and are therefore an important source to study how societal developments have been reflected on the Web. However, the large size of Web archives and their temporal nature pose many challenges to…

数字图书馆 · 计算机科学 2016-12-20 Gerhard Gossen , Elena Demidova , Thomas Risse

In a Web plagued by disappearing resources, Web archive collections provide a valuable means of preserving Web resources important to the study of past events ranging from elections to disease outbreaks. These archived collections start…

数字图书馆 · 计算机科学 2019-05-30 Alexander C. Nwala , Michele C. Weigle , Michael L. Nelson

Archived collections of documents (like newspaper and web archives) serve as important information sources in a variety of disciplines, including Digital Humanities, Historical Science, and Journalism. However, the absence of efficient and…

信息检索 · 计算机科学 2021-07-30 Pavlos Fafalios , Vaibhav Kasturia , Wolfgang Nejdl

The vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose…

Event collections are frequently built by crawling the live web on the basis of seed URIs nominated by human experts. Focused web crawling is a technique where the crawler is guided by reference content pertaining to the event. Given the…

数字图书馆 · 计算机科学 2018-04-06 Martin Klein , Lyudmila Balakireva , Herbert Van de Sompel

Archived collections of documents (like newspaper archives) serve as important information sources for historians, journalists, sociologists and other interested parties. Semantic Layers over such digital archives allow describing and…

信息检索 · 计算机科学 2022-10-19 Pavlos Fafalios , Vaibhav Kasturia , Wolfgang Nejdl

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web…

数字图书馆 · 计算机科学 2021-01-26 Shawn M. Jones , Michele C. Weigle , Michael L. Nelson

Retrieval-augmented generation over semi-structured sources such as HTML is constrained by a mismatch between document structure and the flat, sequence-based interfaces of today's embedding and generative models. Retrieval pipelines often…

信息检索 · 计算机科学 2026-04-24 Mike Rainey , Umut Acar , Muhammed Sezer

As the amount of data on the World Wide Web continues to grow exponentially, access to semantically structured information remains limited. The Semantic Web has emerged as a solution to enhance the machine-readability of data, making it…

数字图书馆 · 计算机科学 2023-06-21 Muhammad Zohaib

In this paper we explore visually the structure of the collection of a digital research data archive in terms of metadata for deposited datasets. We look into the distribution of datasets over different scientific fields; the role of main…

数字图书馆 · 计算机科学 2012-04-17 Andrea Scharnhorst , Olav ten Bosch , Peter Doorn

Data archives are an important source of high quality data in many fields, making them ideal sites to study data reuse. By studying data reuse through citation networks, we are able to learn how hidden research communities - those that use…

数字图书馆 · 计算机科学 2022-10-21 Sara Lafia , Lizhou Fan , Andrea Thomer , Libby Hemphill
‹ 上一页 1 2 3 10 下一页 ›