English
Related papers

Related papers: The Many Shapes of Archive-It

200 papers

Web archive data usually contains high-quality documents that are very useful for creating specialized collections of documents, e.g., scientific digital libraries and repositories of technical reports. In doing so, there is a substantial…

Information Retrieval · Computer Science 2020-09-03 Krutarth Patel , Cornelia Caragea , Mark Phillips , Nathaniel Fox

With the fast growth of the Internet, more and more information is available on the Web. The Semantic Web has many features which cannot be handled by using the traditional search engines. It extracts metadata for each discovered Web…

Artificial Intelligence · Computer Science 2011-11-30 Ahmed Tolba , Nabila Eladawi , Mohammed Elmogy

We document strategies and lessons learned from sampling the web by collecting 27.3 million URLs with 3.8 billion archived pages spanning 26 years (1996-2021) from the Internet Archive's (IA) Wayback Machine. Our goal is to revisit…

Digital Libraries · Computer Science 2025-07-22 Kritika Garg , Sawood Alam , Dietrich Ayala , Mark Graham , Michele C. Weigle , Michael L. Nelson

The Archives Unleashed project aims to improve scholarly access to web archives through a multi-pronged strategy involving tool creation, process modeling, and community building - all proceeding concurrently in mutually-reinforcing…

Digital Libraries · Computer Science 2020-01-16 Nick Ruest , Jimmy Lin , Ian Milligan , Samantha Fritz

Significant parts of cultural heritage are produced on the web during the last decades. While easy accessibility to the current web is a good baseline, optimal access to the past web faces several challenges. This includes dealing with…

Digital Libraries · Computer Science 2017-01-31 Nattiya Kanhabua , Philipp Kemkes , Wolfgang Nejdl , Tu Ngoc Nguyen , Felipe Reis , Nam Khanh Tran

Longitudinal corpora like legal, corporate and newspaper archives are of immense value to a variety of users, and time as an important factor strongly influences their search behavior in these archives. While many systems have been…

Information Retrieval · Computer Science 2018-10-25 Jaspreet Singh , Avishek Anand

The clear, social, and dark web have lately been identified as rich sources of valuable cyber-security information that -given the appropriate tools and methods-may be identified, crawled and subsequently leveraged to actionable…

Cryptography and Security · Computer Science 2021-09-16 Paris Koloveas , Thanasis Chantzios , Christos Tryfonopoulos , Spiros Skiadopoulos

The article examines the theoretical, methodological, and technical foundations of research on audiovisual corpora within the field of digital humanities. It outlines the main transversal issues underlying the processes of constructing,…

Digital Libraries · Computer Science 2025-11-07 Peter Stockinger

Web service composition is the process of synthesizing a new composite service using a set of available Web services in order to satisfy a client request that cannot be treated by any available Web services. The Web services space is a…

Artificial Intelligence · Computer Science 2013-05-02 Chantal Cherifi , Yvan Rivierre , Jean-Francois Santucci

The web of data has brought forth the need to preserve and sustain evolving information within linked datasets; however, a basic requirement of data preservation is the maintenance of the datasets' structural characteristics as well. As…

Approaches form the foundation for conducting scientific research. Querying approaches from a vast body of scientific papers is extremely time-consuming, and without a well-organized management framework, researchers may face significant…

Computation and Language · Computer Science 2025-06-16 Bing Ma , Hai Zhuge

Secondary analysis or the reuse of existing survey data is a common practice among social scientists. Searching for relevant datasets in Digital Libraries is a somehow unfamiliar behaviour for this community. Dataset retrieval, especially…

Digital Libraries · Computer Science 2020-10-13 Zeljko Carevic , Dwaipayan Roy , Philipp Mayr

This article presents the OpenCitations Index, a collection of open citation data maintained by OpenCitations, an independent, not-for-profit infrastructure organisation for open scholarship dedicated to publishing open bibliographic and…

Digital Libraries · Computer Science 2024-10-01 Ivan Heibi , Arianna Moretti , Silvio Peroni , Marta Soricetti

In this paper we introduce the concept of network semantic segmentation for social network analysis. We consider the GitHub social coding network which has been a center of attention for both researchers and software developers. Network…

Social and Information Networks · Computer Science 2019-03-08 Neda Hajiakhoond Bidoki , Gita Sukthankar

Following a particular news story online is an important but difficult task, as the relevant information is often scattered across different domains/sources (e.g., news articles, blogs, comments, tweets), presented in various formats and…

Computation and Language · Computer Science 2018-08-20 Bichen Shi , Thanh-Binh Le , Neil Hurley , Georgiana Ifrim

Web archives are a valuable resource for researchers of various disciplines. However, to use them as a scholarly source, researchers require a tool that provides efficient access to Web archive data for extraction and derivation of smaller…

Digital Libraries · Computer Science 2017-02-06 Helge Holzmann , Vinay Goel , Avishek Anand

Long-term Web archives comprise Web documents gathered over longer time periods and can easily reach hundreds of terabytes in size. Semantic annotations such as named entities can facilitate intelligent access to the Web archive data.…

Information Retrieval · Computer Science 2017-02-03 Tarcisio Souza , Elena Demidova , Thomas Risse , Helge Holzmann , Gerhard Gossen , Julian Szymanski

Structured and semi-structured data describing entities, taxonomies and ontologies appears in many domains. There is a huge interest in integrating structured information from multiple sources; however integrating structured data to infer…

Artificial Intelligence · Computer Science 2010-05-28 Anon Plangprasopchok , Kristina Lerman , Lise Getoor

The Semantic Web is an extension of the current web in which information is given well-defined meaning. The perspective of Semantic Web is to promote the quality and intelligence of the current web by changing its contents into machine…

Artificial Intelligence · Computer Science 2012-08-06 Hamed Hassanzadeh , MohammadReza Keyvanpour

In today's Web, Web Services are created and updated on the fly. For answering complex needs of users, the construction of new web services based on existing ones is required. It has received a great attention from different communities.…

Software Engineering · Computer Science 2014-02-11 Narges Hesami Rostami , Esmaeil Kheirkhah , Mehrdad Jalali