English
Related papers

Related papers: Web Archive Analytics

200 papers

Collections of research article data harvested from the web have become common recently since they are important resources for experimenting on tasks such as named entity recognition, text summarization, or keyword generation. In fact,…

Information Retrieval · Computer Science 2022-05-24 Erion Çano , Benjamin Roth

Scientists, governments, and companies increasingly publish datasets on the Web. Google's Dataset Search extracts dataset metadata -- expressed using schema.org and similar vocabularies -- from Web pages in order to make datasets…

Information Retrieval · Computer Science 2020-06-15 Omar Benjelloun , Shiyu Chen , Natasha Noy

Because of its willingness to share data with academia and industry, Twitter has been the primary social media platform for scientific research as well as for consulting businesses and governments in the last decade. In recent years, a…

Social and Information Networks · Computer Science 2023-04-12 Juergen Pfeffer , Angelina Mooseder , Jana Lasser , Luca Hammer , Oliver Stritzel , David Garcia

In the digital era, user interactions with various resources such as databases, data warehouses, websites, and knowledge graphs (KGs) are increasingly mediated through digital platforms. These interactions leave behind digital traces,…

Databases · Computer Science 2025-08-20 Dihia Lanasri

This presentation focuses on the importance of web crawling and page ranking algorithms in dealing with the massive amount of data present on the World Wide Web. As the web continues to grow exponentially, efficient search and retrieval…

Information Retrieval · Computer Science 2023-06-22 Nithin T K , Chandana S , Barani G , Chavva Dharani , M S Karishma

Nowadays, invoking third party code increasingly involves calling web services via their web APIs, as opposed to the more traditional scenario of downloading a library and invoking the library's API. However, there are also new challenges…

Software Engineering · Computer Science 2017-05-19 Erik Wittern , Annie Ying , Yunhui Zheng , Jim A. Laredo , Julian Dolby , Christopher C. Young , Aleksander A. Slominski

In this paper we review studies of the growth of the Internet and technologies that are useful for information search and retrieval on the Web. Search engines are retrieve the efficient information. We collected data on the Internet from…

Information Retrieval · Computer Science 2013-10-18 Avinash N Bhute , B. B. Meshram

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

General Economics · Economics 2023-08-07 Jens Foerderer

The Web is a ubiquitous economic, educational, and collaborative space. However, it also serves as a haven for personal information harvesting. Existing decentralised Web-based ecosystems, such as Solid, aim to combat personal data…

Databases · Computer Science 2020-08-17 Ruben Taelman , Simon Steyskal , Sabrina Kirrane

When a user requests a web page from a web archive, the user will typically either get an HTTP 200 if the page is available, or an HTTP 404 if the web page has not been archived. This is because web archives are typically accessed by URI…

Digital Libraries · Computer Science 2019-08-09 Lulwah M. Alkwai , Michael L. Nelson , Michele C. Weigle

Scientists across disciplines often use data from the internet to conduct research, generating valuable insights about human behavior. However, as generative AI relying on massive text corpora becomes increasingly valuable, platforms have…

Computers and Society · Computer Science 2024-12-20 Megan A. Brown , Andrew Gruen , Gabe Maldoff , Solomon Messing , Zeve Sanderson , Michael Zimmer

One in five arXiv articles published in 2021 contained a URI to a Git Hosting Platform (GHP), which demonstrates the growing prevalence of GHP URIs in scholarly publications. However, GHP URIs are vulnerable to the same reference rot that…

Digital Libraries · Computer Science 2024-01-11 Emily Escamilla , Martin Klein , Talya Cooper , Vicky Rampin , Michele C. Weigle , Michael L. Nelson

Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the…

Networking and Internet Architecture · Computer Science 2018-02-12 Mohammed S. Hadi , Ahmed Q. Lawey , Taisir E. H. El-Gorashi , Jaafar M. H. Elmirghani

The World Wide Web is a popular and interactive medium to distribute information in this scenario. The web is huge, diverse, ever changing, widely disseminated global information service center. We are familiar with terms like e-commerce,…

Other Computer Science · Computer Science 2012-08-30 Priyanka Rahi

A hidden database refers to a dataset that an organization makes accessible on the web by allowing users to issue queries through a search interface. In other words, data acquisition from such a source is not by following static…

Databases · Computer Science 2012-08-02 Cheng Sheng , Nan Zhang , Yufei Tao , Xin Jin

Cloud services are critical to society. However, their reliability is poorly understood. Towards solving the problem, we propose a standard repository for cloud uptime data. We populate this repository with the data we collect containing…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-15 Sacheendra Talluri , Dante Niewenhuis , Xiaoyu Chu , Jakob Kyselica , Mehmet Cetin , Alexander Balgavy , Alexandru Iosup

Online data sources offer tremendous promise to demography and other social sciences, but researchers worry that the group of people who are represented in online datasets can be different from the general population. We show that by…

Applications · Statistics 2019-07-01 Dennis M. Feehan , Curtiss Cobb

Archive collections are nowadays mostly available through search engines interfaces, which allow a user to retrieve documents by issuing queries. The study of these collections may be, however, impaired by some aspects of search engines,…

Computation and Language · Computer Science 2023-02-01 Nicolas Gutehrlé , Antoine Doucet , Adam Jatowt

Introduction: Before embarking on the design of any computer system it is first necessary to assess the magnitude of the problem. In the case of a web search engine this assessment amounts to determining the current size of the web, the…

Information Retrieval · Computer Science 2013-07-05 Andrew Trotman , Jinglan Zhang

The Common Crawl (CC) corpus is the largest open web crawl dataset containing 9.5+ petabytes of data captured since 2008. The dataset is instrumental in training large language models, and as such it has been studied for (un)desirable…

Computation and Language · Computer Science 2026-05-07 Ilya Ilyankou , Meihui Wang , Stefano Cavazzi , James Haworth