English
Related papers

Related papers: Index wiki database: design and experiments

200 papers

The search for relevant information can be very frustrating for users who, unintentionally, use too general or inappropriate keywords to express their requests. To overcome this situation, query expansion techniques aim at transforming the…

Information Retrieval · Computer Science 2016-05-13 Joan Guisado-Gámez , Arnau Prat-Pérez , Josep Lluís Larriba-Pey

There is an explosive growth of information in the World Wide Web thus posing a challenge to Web users to extract essential knowledge from the Web. Search engines help us to narrow down the search in the form of Search Engine Result Pages…

Information Retrieval · Computer Science 2013-03-26 Srikantaiah K C , Suraj M , Venugopal K R , L M Patnaik

We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy…

Information Retrieval · Computer Science 2018-12-24 Milad Alshomary , Michael Völske , Tristan Licht , Henning Wachsmuth , Benno Stein , Matthias Hagen , Martin Potthast

In this report, we unify two quite distinct approaches to information retrieval: region models and language models. Region models were developed for structured document retrieval. They provide a well-defined behaviour as well as a simple…

Information Retrieval · Computer Science 2012-05-02 Djoerd Hiemstra , Vojkan Mihajlovic

While several self-indexes for highly repetitive texts exist, developing a practical self-index applicable to real world repetitive texts remains a challenge. ESP-index is a grammar-based self-index on the notion of edit-sensitive parsing…

Data Structures and Algorithms · Computer Science 2014-04-29 Yoshimasa Takabatake , Yasuo Tabei , Hiroshi Sakamoto

Mathematical information retrieval (MathIR) applications such as semantic formula search and question answering systems rely on knowledge-bases that link mathematical expressions to their natural language names. For database population,…

Digital Libraries · Computer Science 2021-04-13 Philipp Scharpf , Moritz Schubotz , Bela Gipp

Wikipedia categories, a classification scheme built for organizing and describing Wikpedia articles, are being applied in computer science research. This paper adopts a systematic literature review approach, in order to identify different…

Digital Libraries · Computer Science 2020-04-22 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón

With the advent of the Internet, a new era of digital information exchange has begun. Currently, the Internet encompasses more than five billion online sites and this number is exponentially increasing every day. Fundamentally, Information…

Information Retrieval · Computer Science 2012-04-03 Youssef Bassil , Paul Semaan

The limited size of existing query-focused summarization datasets renders training data-driven summarization models challenging. Meanwhile, the manual construction of a query-focused summarization corpus is costly and time-consuming. In…

Computation and Language · Computer Science 2022-07-25 Haichao Zhu , Li Dong , Furu Wei , Bing Qin , Ting Liu

The performance of processing search queries depends heavily on the stored index size. Accordingly, considerable research efforts have been devoted to the development of efficient compression techniques for inverted indexes. Roughly, index…

Information Retrieval · Computer Science 2011-07-29 M. Feldman , R. Lempel , O. Somekh , K. Vornovitsky

We present results from our quantitative study of statistical and network properties of literary and scientific texts written in two languages: English and Polish. We show that Polish texts are described by the Zipf law with the scaling…

Physics and Society · Physics 2013-12-16 Iwona Grabska-Gradzinska , Andrzej Kulig , Jaroslaw Kwapien , Stanislaw Drozdz

Wikipedia is a rich and invaluable source of information. Its central place on the Web makes it a particularly interesting object of study for scientists. Researchers from different domains used various complex datasets related to Wikipedia…

Information Retrieval · Computer Science 2019-03-21 Nicolas Aspert , Volodymyr Miz , Benjamin Ricaud , Pierre Vandergheynst

Searches for phrases and word sets in large text arrays by means of additional indexes are considered. Their use may reduce the query-processing time by an order of magnitude in comparison with standard inverted files.

Information Retrieval · Computer Science 2018-11-27 A. B. Veretennikov

In this article, I conduct a textual and contextual analysis of the empirical literature on Zipf's law for cities. Building on previous meta-analysis material openly available, I collect full texts and bibliographies of 66 scientific…

Physics and Society · Physics 2022-02-01 Clémentine Cottineau

We study how differences in persuasive language across Wikipedia articles, written in either English and Russian, can uncover each culture's distinct perspective on different subjects. We develop a large language model (LLM) powered system…

Computation and Language · Computer Science 2024-10-01 Bryan Li , Aleksey Panasyuk , Chris Callison-Burch

Knowledge bases are very good sources for knowledge extraction, the ability to create knowledge from structured and unstructured sources and use it to improve automatic processes as query expansion. However, extracting knowledge from…

Information Retrieval · Computer Science 2015-05-07 Joan Guisado-Gámez , Arnau Prat-Pérez

From more than half a century ago indexing scientific articles has been studied intensively to provide a more efficient data retrieval and to conserve researchers invaluable time. In the last two decades with the emergence of the World Wide…

Digital Libraries · Computer Science 2014-01-07 Azam Majooni , Mona Masood , Amir Akhavan

In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of…

Information Retrieval · Computer Science 2017-03-06 Nemanja Spasojevic , Preeti Bhargava , Guoning Hu

Information Synchronization of semi-structured data across languages is challenging. For instance, Wikipedia tables in one language should be synchronized across languages. To address this problem, we introduce a new dataset InfoSyncC and a…

Computation and Language · Computer Science 2023-07-10 Siddharth Khincha , Chelsi Jain , Vivek Gupta , Tushar Kataria , Shuo Zhang

This paper describes a prototype of extended XDB. XDB is an open-source and extensible database architecture developed by National Aeronautics and Space Administration (NASA) to provide integration of heterogeneous and distributed…

Databases · Computer Science 2012-11-27 Wook-Sung Yoo