English
Related papers

Related papers: Using Web Page Titles to Rediscover Lost Web Pages

200 papers

When working with any sort of knowledge base (KB) one has to make sure it is as complete and also as up-to-date as possible. Both tasks are non-trivial as they require recall-oriented efforts to determine which entities and relationships…

Information Retrieval · Computer Science 2020-02-06 Shuo Zhang , Edgar Meij , Krisztian Balog , Ridho Reinanda

Scientific digital libraries play a critical role in the development and dissemination of scientific literature. Despite dedicated search engines, retrieving relevant publications from the ever-growing body of scientific literature remains…

Information Retrieval · Computer Science 2021-06-29 Florian Boudin , Béatrice Daille , Evelyne Jacquey , Jian-Yun Nie

Topic modeling is an unsupervised method for revealing the hidden semantic structure of a corpus. It has been increasingly widely adopted as a tool in the social sciences, including political science, digital humanities and sociological…

Information Retrieval · Computer Science 2022-01-12 Zheng Fang , Yulan He , Rob Procter

The extraction of main content from web pages is an important task for numerous applications, ranging from usability aspects, like reader views for news articles in web browsers, to information retrieval or natural language processing.…

Machine Learning · Computer Science 2020-04-30 Jurek Leonhardt , Avishek Anand , Megha Khosla

Significant parts of cultural heritage are produced on the web during the last decades. While easy accessibility to the current web is a good baseline, optimal access to the past web faces several challenges. This includes dealing with…

Digital Libraries · Computer Science 2017-01-31 Nattiya Kanhabua , Philipp Kemkes , Wolfgang Nejdl , Tu Ngoc Nguyen , Felipe Reis , Nam Khanh Tran

In this paper we introduce the concept of dynamic link pages. A web site/page contains a number of links to other pages. All the links are not equally important. Few links are more frequently visited and few rarely visited. In this…

Information Retrieval · Computer Science 2010-03-29 C Ravindranath Chowdary

PageRank has become a key element in the success of search engines, allowing to rank the most important hits in the top screen of results. One key aspect that distinguishes PageRank from other prestige measures such as in-degree is its…

Information Retrieval · Computer Science 2007-05-23 Santo Fortunato , Marian Boguna , Alessandro Flammini , Filippo Menczer

In this paper we study the prevalence of unique entity identifiers on the Web. These are, e.g., ISBNs (for books), GTINs (for commercial products), DOIs (for documents), email addresses, and others. We show how these identifiers can be…

Databases · Computer Science 2016-07-19 Aliaksandr Talaika , Joanna Biega , Antoine Amarilli , Fabian M. Suchanek

Search query suggestions affect users' interactions with search engines, which then influences the information they encounter. Thus, bias in search query suggestions can lead to exposure to biased search results and can impact opinion…

Information Retrieval · Computer Science 2024-11-01 Fabian Haak , Björn Engelmann , Christin Katharina Kreutz , Philipp Schaer

Boilerplate refers to unwanted and repeated parts of a webpage (such as ads or table of contents) that distracts the user from reading the core content of the webpage, such as a news article. Accurate detection and removal of boilerplate…

Information Retrieval · Computer Science 2020-01-15 Joy Bose

For many companies, competitiveness in e-commerce requires a successful presence on the web. Web sites are used to establish the company's image, to promote and sell goods and to provide customer support. The success of a web site affects…

Machine Learning · Computer Science 2007-05-23 Myra Spiliopoulou , Carsten Pohle

Searching health information on web has become an integral part of today's world, and many people turn to the Web for healthcare information and healthcare assessment. Our pilot study investigates users' preferences for the type of search…

Information Retrieval · Computer Science 2014-10-30 Shanu Sushmita , Si-Chi Chin

Table retrieval, essential for accessing information through tabular data, is less explored compared to text retrieval. The row/column structure and distinct fields of tables (including titles, headers, and cells) present unique challenges.…

Information Retrieval · Computer Science 2025-03-05 Da Li , Keping Bi , Jiafeng Guo , Xueqi Cheng

Thousands of documents are made available to the users via the web on a daily basis. One of the most extensively studied problems in the context of such document streams is burst identification. Given a term t, a burst is generally…

Databases · Computer Science 2012-05-31 Theodoros Lappas , Marcos R. Vieira , Dimitrios Gunopulos , Vassilis J. Tsotras

The objective of this research is to determine if the reference to a country in the title, keywords or abstract of a publication can influence its visibility (measured by the impact factor of the publishing journal) and citability (measured…

Digital Libraries · Computer Science 2018-10-31 Giovanni Abramo , Ciriaco Andrea D'Angelo , Flavia Di Costa

Information Extraction from scientific literature can be challenging due to the highly specialised nature of such text. We describe our entity recognition methods developed as part of the DEAL (Detecting Entities in the Astrophysics…

Computation and Language · Computer Science 2022-11-28 Xiang Dai , Sarvnaz Karimi

Topic models are popular models for analyzing a collection of text documents. The models assert that documents are distributions over latent topics and latent topics are distributions over words. A nested document collection is where…

Information Retrieval · Computer Science 2021-04-05 Jason Wang , Robert E. Weiss

Everyone knows that thousand of words are represented by a single image. As a result image search has become a very popular mechanism for the Web searchers. Image search means, the search results are produced by the search engine should be…

Computer Vision and Pattern Recognition · Computer Science 2014-01-14 Sukanta Sinha , Rana Dattagupta , Debajyoti Mukhopadhyay

Wikipedia categories, a classification scheme built for organizing and describing Wikpedia articles, are being applied in computer science research. This paper adopts a systematic literature review approach, in order to identify different…

Digital Libraries · Computer Science 2020-04-22 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón

Crawler-based search engines are the mostly used search engines among web and Internet users, involve web crawling, storing in database, ranking, indexing and displaying to the user. But it is noteworthy that because of increasing changes…

Information Retrieval · Computer Science 2013-05-14 Ali Tourani , Amir Seyed Danesh