English
Related papers

Related papers: Index wiki database: design and experiments

200 papers

Although the vast majority of knowledge bases KBs are heavily biased towards English, Wikipedias do cover very different topics in different languages. Exploiting this, we introduce a new multilingual dataset (X-WikiRE), framing relation…

Computation and Language · Computer Science 2019-08-16 Mostafa Abdou , Cezar Sas , Rahul Aralikatte , Isabelle Augenstein , Anders Søgaard

Traditional information retrieval systems rely on keywords to index documents and queries. In such systems, documents are retrieved based on the number of shared keywords with the query. This lexical-focused retrieval leads to inaccurate…

Information Retrieval · Computer Science 2013-03-08 Fatiha Boubekeur , Wassila Azzoug

Search engine has become a fundamental component in various web and mobile applications. Retrieving relevant documents from the massive datasets is challenging for a search engine system, especially when faced with verbose or tail queries.…

Information Retrieval · Computer Science 2020-08-11 Kuan Fang , Long Zhao , Zhan Shen , RuiXing Wang , RiKang Zhour , LiWen Fan

This paper presents our experience on building RDF knowledge graphs for an industrial use case in the legal domain. The information contained in legal information systems are often accessed through simple keyword interfaces and presented as…

Databases · Computer Science 2019-11-19 Ademar Crotti Junior , Fabrizio Orlandi , Declan O'Sullivan , Christian Dirschl , Quentin Reul

The Internet-based encyclopaedia Wikipedia has grown to become one of the most visited web-sites on the Internet. However, critics have questioned the quality of entries, and an empirical study has shown Wikipedia to contain errors in a…

Digital Libraries · Computer Science 2011-01-04 Finn Aarup Nielsen

Today's conventional search engines hardly do provide the essential content relevant to the user's search query. This is because the context and semantics of the request made by the user is not analyzed to the full extent. So here the need…

Information Retrieval · Computer Science 2012-07-25 Swathi Rajasurya , Tamizhamudhu Muralidharan , Sandhiya Devi , S. Swamynathan

We propose a stochastic model for the number of different words in a given database which incorporates the dependence on the database size and historical changes. The main feature of our model is the existence of two different classes of…

Physics and Society · Physics 2013-05-16 Martin Gerlach , Eduardo G. Altmann

We present a novel method for efficiently searching top-k neighbors for documents represented in high dimensional space of terms based on the cosine similarity. Mostly, documents are stored as bag-of-words tf-idf representation. One of the…

Information Retrieval · Computer Science 2016-05-24 Gaurav Singh , Benjamin Piwowarski

Full-text search engines are important tools for information retrieval. Term proximity is an important factor in relevance score measurement. In a proximity full-text search, we assume that a relevant document contains query terms near each…

Information Retrieval · Computer Science 2018-11-20 Alexander B. Veretennikov

In this paper we extract the topology of the semantic space in its encyclopedic acception, measuring the semantic flow between the different entries of the largest modern encyclopedia, Wikipedia, and thus creating a directed complex network…

Physics and Society · Physics 2011-03-08 A. P. Masucci , A. Kalampokis , V. M. Eguíluz , E. Hernández-García

This paper aims to review the fiercely discussed question of whether the ranking of Wikipedia articles in search engines is justified by the quality of the articles. After an overview of current research on information quality in Wikipedia,…

Information Retrieval · Computer Science 2011-09-06 Dirk Lewandowski , Ulrike Spree

In this paper we present a novel method for retrieving information in languages other than that of the query. We use this technique in combination with existing traditional Cross Language Information Retrieval (CLIR) techniques to improve…

Information Retrieval · Computer Science 2009-06-17 Mikhail Basilyan

Document retrieval is a key stage of standard Web search engines. Existing dual-encoder dense retrievers obtain representations for questions and documents independently, allowing for only shallow interactions between them. To overcome this…

Computation and Language · Computer Science 2023-05-17 Noah Ziems , Wenhao Yu , Zhihan Zhang , Meng Jiang

Wikidata is a multi-language knowledge base that is being edited and maintained by editors from different language communities. Due to the structured nature of its content, the contributions are in various forms, including manual edit,…

Human-Computer Interaction · Computer Science 2023-11-07 Jeffrey Jun-jie Ma , Charles Chuankai Zhang

This paper highlights the growing importance of information retrieval (IR) engines in the scientific community, addressing the inefficiency of traditional keyword-based search engines due to the rising volume of publications. The proposed…

Information Retrieval · Computer Science 2024-10-24 Mahsa Shamsabadi , Jennifer D'Souza

This paper presents a pipeline designed to transform raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from raw…

Computation and Language · Computer Science 2026-05-18 Mihailo Škorić , Cosimo Palma

Scientific digital libraries play a critical role in the development and dissemination of scientific literature. Despite dedicated search engines, retrieving relevant publications from the ever-growing body of scientific literature remains…

Information Retrieval · Computer Science 2021-06-29 Florian Boudin , Béatrice Daille , Evelyne Jacquey , Jian-Yun Nie

Inverted indexing has traditionally been a cornerstone of modern search systems, leveraging exact term matches to determine relevance between queries and documents. However, this term-based approach often emphasizes surface-level token…

Information Retrieval · Computer Science 2025-09-30 Zan Li , Jiahui Chen , Yuan Chai , Xiaoze Jiang , Xiaohua Qi , Zhiheng Qin , Runbin Zhou , Shun Zuo , Guangchao Hao , Kefeng Wang , Jingshan Lv , Yupeng Huang , Xiao Liang , Han Li

We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. Relating visual and textual items on document pages has…

The task of determining the similarity of text documents has received considerable attention in many areas such as Information Retrieval, Text Mining, Natural Language Processing (NLP) and Computational Linguistics. Transferring data to…

Information Retrieval · Computer Science 2022-11-23 Bakhyt Bakiyev
‹ Prev 1 8 9 10 Next ›