English
Related papers

Related papers: First in the Web, but Where are the Pieces?

200 papers

The Linked Data Paradigm is one of the most promising technologies for publishing, sharing, and connecting data on the Web, and offers a new way for data integration and interoperability. However, the proliferation of distributed,…

Databases · Computer Science 2012-05-11 Yannis Stavrakas , George Papastefanatos , Theodore Dalamagas , Vassilis Christophides

In recent years studying the content of the World Wide Web became a very important yet rather difficult task. There is a need for a compression technique that would allow a web graph representation to be put into the memory while…

Data Structures and Algorithms · Computer Science 2013-05-02 Filip Proborszcz

Web tracking is an omnipresent phenomenon in today's web, affecting users in their day-to-day lives. Filter lists and blockers were invented to detect trackers and to protect users. Due to limitations of said tools, researchers developed…

Cryptography and Security · Computer Science 2026-05-06 Wolf Rieder , Philip Raschke , Thomas Cory , Christian René Sechting , Aditya Kumar , Axel Küpper

Long-term Web archives comprise Web documents gathered over longer time periods and can easily reach hundreds of terabytes in size. Semantic annotations such as named entities can facilitate intelligent access to the Web archive data.…

Information Retrieval · Computer Science 2017-02-03 Tarcisio Souza , Elena Demidova , Thomas Risse , Helge Holzmann , Gerhard Gossen , Julian Szymanski

Beyond the information stored in pages of the World Wide Web, novel types of ``meta-information'' are created when they connect to each other. This information is a collective effect of independent users writing and linking pages, hidden…

Disordered Systems and Neural Networks · Physics 2009-11-07 Jean-Pierre Eckmann , Elisha Moses

Structure information extraction refers to the task of extracting structured text fields from web pages, such as extracting a product offer from a shopping page including product title, description, brand and price. It is an important…

Computation and Language · Computer Science 2022-02-02 Qifan Wang , Yi Fang , Anirudh Ravula , Fuli Feng , Xiaojun Quan , Dongfang Liu

The Semantic Web is an extension of the current web in which information is given well-defined meaning. The perspective of Semantic Web is to promote the quality and intelligence of the current web by changing its contents into machine…

Artificial Intelligence · Computer Science 2012-08-06 Hamed Hassanzadeh , MohammadReza Keyvanpour

Page layout analysis is a fundamental step in document processing which enables to segment a page into regions of interest. With highly complex layouts and mixed scripts, scholarly commentaries are text-heavy documents which remain…

Information Retrieval · Computer Science 2022-12-29 Najem-Meyer Sven , Romanello Matteo

Software has long been established as an essential aspect of the scientific process in mathematics and other disciplines. However, reliably referencing software in scientific publications is still challenging for various reasons. A crucial…

Digital Libraries · Computer Science 2017-02-07 Helge Holzmann , Wolfram Sperber , Mila Runnwerth

The goal of a technology-assisted review is to achieve high recall with low human effort. Continuous active learning algorithms have demonstrated good performance in locating the majority of relevant documents in a collection, however their…

Information Retrieval · Computer Science 2018-10-15 Jie Zou , Dan Li , Evangelos Kanoulas

Search engines are a combination of hardware and computer software supplied by a particular company through the website which has been determined. Search engines collect information from the web through bots or web crawlers that crawls the…

Information Retrieval · Computer Science 2014-10-22 Ahmad Josi , Leon Andretti Abdillah , Suryayusra

We are presenting a set of multilingual text analysis tools that can help analysts in any field to explore large document collections quickly in order to determine whether the documents contain information of interest, and to find the…

Computation and Language · Computer Science 2007-05-23 Camelia Ignat , Bruno Pouliquen , Ralf Steinberger , Tomaz Erjavec

Collections of research article data harvested from the web have become common recently since they are important resources for experimenting on tasks such as named entity recognition, text summarization, or keyword generation. In fact,…

Information Retrieval · Computer Science 2022-05-24 Erion Çano , Benjamin Roth

On the worldwide web, not only are webpages connected but source code is too. Software development is becoming more accessible to everyone and the licensing for software remains complicated. We need to know if software licenses are being…

Software Engineering · Computer Science 2018-08-02 Stephen Romansky , Cheng Chen , Baljeet Malhotra , Abram Hindle

In this paper authors analyzed 163412 keywords and results with featured snippets collected from localized Polish Google search engine. A method-ology for retrieving data from Google search engine was proposed in terms of obtaining…

Information Retrieval · Computer Science 2019-12-05 Artur Strzelecki , Paulina Rutecka

Document networks are found in various collections of real-world data, such as citation networks, hyperlinked web pages, and online social networks. A large number of generative models have been proposed because they offer intuitive and…

Physics and Society · Physics 2020-01-22 Takafumi J. Suzuki

Over the past few years, we have built a system that has exposed large volumes of Deep-Web content to Google.com users. The content that our system exposes contributes to more than 1000 search queries per-second and spans over 50 languages…

Databases · Computer Science 2009-09-15 Jayant Madhavan , Loredana Afanasiev , Lyublena Antova , Alon Halevy

Common Crawl is a multi-petabyte longitudinal dataset containing over 100 billion web pages which is widely used as a source of language data for sequence model training and in web science research. Each of its constituent archives is on…

Networking and Internet Architecture · Computer Science 2024-04-16 Henry S. Thompson

Linked Open Datasets about scholarly publications enable the development and integration of sophisticated end-user services; however, richer datasets are still needed. The first goal of this Challenge was to investigate novel approaches to…

Digital Libraries · Computer Science 2014-08-22 Christoph Lange , Angelo Di Iorio

Experience is what makes our life more effective that is why it is necessary to share experience among people. The use of information technologies is the most technological way to work with experience, and the use of the Web is the best way…

Information Retrieval · Computer Science 2018-05-18 Olegs Verhodubs