English
Related papers

Related papers: Analyzing and Visualizing the Semantic Coverage of…

200 papers

Explicit Semantic Analysis (ESA) is a technique used to represent a piece of text as a vector in the space of concepts, such as Articles found in Wikipedia. We propose a methodology to incorporate knowledge of Inter-relatedness between…

Computation and Language · Computer Science 2020-12-02 Naveen Elango , Pawan Prasad K

Many real systems have been modelled in terms of network concepts, and written texts are a particular example of information networks. In recent years, the use of network methods to analyze language has allowed the discovery of several…

Computation and Language · Computer Science 2016-06-28 Henrique F. de Arruda , Luciano da F. Costa , Diego R. Amancio

Wikipedia represents the largest and most popular source of encyclopedic knowledge in the world today, aiming to provide equal access to information worldwide. From a global online survey of 65,031 readers of Wikipedia and their…

Computers and Society · Computer Science 2020-07-22 Isaac Johnson , Florian Lemmerich , Diego Sáez-Trumper , Robert West , Markus Strohmaier , Leila Zia

Understanding texts requires memory: the reader has to keep in mind enough words to create meaning. This calls for a relation between the memory of the reader and the structure of the text. To investigate this interaction, we first identify…

Physics and Society · Physics 2007-05-23 E. Alvarez-Lacalle , B. Dorow , J. -P. Eckmann , E. Moses

Contributing to the writing of history has never been as easy as it is today thanks to Wikipedia, a community-created encyclopedia that aims to document the world's knowledge from a neutral point of view. Though everyone can participate it…

Social and Information Networks · Computer Science 2016-03-04 Claudia Wagner , Eduardo Graells-Garrido , David Garcia , Filippo Menczer

The frequency of a web search keyword generally reflects the degree of public interest in a particular subject matter. Search logs are therefore useful resources for trend analysis. However, access to search logs is typically restricted to…

Social and Information Networks · Computer Science 2015-09-09 Mitsuo Yoshida , Yuki Arase , Takaaki Tsunoda , Mikio Yamamoto

Wikipedia can be considered as an extreme form of a self-managing team, as a means of labour division. One could expect that this bottom-up approach, with the absense of top-down organisational control, would lead to a chaos, but our…

Digital Libraries · Computer Science 2007-05-23 Sander Spek , Eric Postma , H. Jaap van den Herik

The Wikipedia editors' community has been actively pursuing the intent of achieving gender equality. To that end, it is important to explore the historical evolution of underlying gender disparities in Wikipedia articles. This paper…

Computers and Society · Computer Science 2025-01-23 Yahya Yunus , Tianwa Chen , Gianluca Demartini

Since the advent of the web, the amount of data on wen has been increased several million folds. In recent years web data generated is more than data stored for years. One important data format is text. To answer user queries over the…

Information Retrieval · Computer Science 2018-11-19 Chandra Shekhar Yadav

We propose an edit-centric approach to assess Wikipedia article quality as a complementary alternative to current full document-based techniques. Our model consists of a main classifier equipped with an auxiliary generative module which,…

Computation and Language · Computer Science 2019-09-20 Edison Marrese-Taylor , Pablo Loyola , Yutaka Matsuo

We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy…

Information Retrieval · Computer Science 2018-12-24 Milad Alshomary , Michael Völske , Tristan Licht , Henning Wachsmuth , Benno Stein , Matthias Hagen , Martin Potthast

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

Computation and Language · Computer Science 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

Wikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In Wikipedia, content…

Success of planetary-scale online collaborative platforms such as Wikipedia is hinged on active and continued participation of its voluntary contributors. The phenomenal success of Wikipedia as a valued multilingual source of information is…

Social and Information Networks · Computer Science 2021-09-22 Paramita Das , Bhanu Prakash Reddy Guda , Debajit Chakraborty , Soumya Sarkar , Animesh Mukherjee

Working with Web archives raises a number of issues caused by their temporal characteristics. Depending on the age of the content, additional knowledge might be needed to find and understand older texts. Especially facts about entities are…

Computation and Language · Computer Science 2017-02-07 Helge Holzmann , Thomas Risse

Research on vandalism in Wikipedia has been of interest for the last decade. This paper performs a literature review on the subject, with the goal of identifying the main research topics and approaches, methods and techniques used. 67…

Digital Libraries · Computer Science 2016-06-20 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón

We describe a deployed scalable system for organizing published scientific literature into a heterogeneous graph to facilitate algorithmic manipulation and discovery. The resulting literature graph consists of more than 280M nodes,…

A thermodynamic framework is presented to characterize the evolution of efficiency, order, and quality in social content production systems, and this framework is applied to the analysis of Wikipedia. Contributing editors are characterized…

Social and Information Networks · Computer Science 2012-04-18 Huan-Kai Peng , Ying Zhang , Peter Pirolli , Tad Hogg

Unlike static documents, version-controlled documents are edited by one or more authors over a certain period of time. Examples include large scale computer code, papers authored by a team of scientists, and online discussion boards. Such…

Human-Computer Interaction · Computer Science 2012-05-16 Seungyeon Kim , Joshua V. Dillon , Guy Lebanon

We propose a framework for analyzing discourse by combining two interdependent concepts from sociolinguistic theory: face acts and politeness. While politeness has robust existing tools and data, face acts are less resourced. We introduce a…

Computation and Language · Computer Science 2024-08-07 Adil Soubki , Shyne Choi , Owen Rambow