English
Related papers

Related papers: Language-agnostic Topic Classification for Wikiped…

200 papers

Wikipedia serves as a good example of how editors collaborate to form and maintain an article. The relationship between editors, derived from their sequence of editing activity, results in a directed network structure called the revision…

Social and Information Networks · Computer Science 2019-04-18 James R. Ashford , Liam D. Turner , Roger M. Whitaker , Alun Preece , Diane Felmlee , Don Towsley

Automatic text categorization is a complex and useful task for many natural language processing applications. Recent approaches to text categorization focus more on algorithms than on resources involved in this operation. In contrast to…

cmp-lg · Computer Science 2008-02-03 Jose Maria Gomez Hidalgo , Manuel de Buenaga Rodriguez

Cross-document event coreference resolution is a foundational task for NLP applications involving multi-text processing. However, existing corpora for this task are scarce and relatively small, while annotating only modest-size clusters of…

Computation and Language · Computer Science 2021-05-03 Alon Eirew , Arie Cattan , Ido Dagan

This paper presents a novel analysis and visualization of English Wikipedia data. Our specific interest is the analysis of basic statistics, the identification of the semantic structure and age of the categories in this free online…

Information Retrieval · Computer Science 2007-05-23 Todd Holloway , Miran Bozicevic , Katy Börner

With more than 11 times as many pageviews as the next largest edition, English Wikipedia dominates global knowledge access relative to other language editions. Readers are prone to assuming English Wikipedia as a superset of all language…

Human-Computer Interaction · Computer Science 2026-01-21 Zining Wang , Yuxuan Zhang , Dongwook Yoon , Nicholas Vincent , Farhan Samir , Vered Shwartz

Identifying which Wikipedia articles are related to science fiction, fantasy, or their hybrids is challenging because genre boundaries are porous and frequently overlap. Wikipedia nonetheless offers machine-readable structure beyond text,…

Information Retrieval · Computer Science 2026-03-02 Włodzimierz Lewoniewski , Milena Stróżyna , Izabela Czumałowska , Elżbieta Lewańska

Wikipedia is the largest online encyclopedia: its open contribution policy allows everyone to edit and share their knowledge. A challenge of radical openness is that it facilitates introducing biased contents or perspectives in Wikipedia.…

Digital Libraries · Computer Science 2022-11-22 Puyu Yang , Giovanni Colavizza

Wikipedia has high-quality articles on a variety of topics and has been used in diverse research areas. In this study, a method is presented for using Wikipedia's editor information to build recommender systems in various domains that…

Information Retrieval · Computer Science 2023-06-16 Katsuhiko Hayashi

Despite recent progress in computer vision, fine-grained interpretation of satellite images remains challenging because of a lack of labeled training data. To overcome this limitation, we propose using Wikipedia as a previously untapped…

Computer Vision and Pattern Recognition · Computer Science 2018-09-28 Evan Sheehan , Burak Uzkent , Chenlin Meng , Zhongyi Tang , Marshall Burke , David Lobell , Stefano Ermon

Social tagging has become an interesting approach to improve search and navigation over the actual Web, since it aggregates the tags added by different users to the same resource in a collaborative way. This way, it results in a list of…

Information Retrieval · Computer Science 2012-02-27 Arkaitz Zubiaga

This paper presents a new method for automatically detecting words with lexical gender in large-scale language datasets. Currently, the evaluation of gender bias in natural language processing relies on manually compiled lexicons of…

Computation and Language · Computer Science 2022-06-29 Marion Bartl , Susan Leavy

Cross-lingual Entity Linking (XEL), the problem of grounding mentions of entities in a foreign language text into an English knowledge base such as Wikipedia, has seen a lot of research in recent years, with a range of promising techniques.…

Computation and Language · Computer Science 2020-10-08 Xingyu Fu , Weijia Shi , Xiaodong Yu , Zian Zhao , Dan Roth

In this work, we compare two simple methods of tagging scientific publications with labels reflecting their content. As a first source of labels Wikipedia is employed, second label set is constructed from the noun phrases occurring in the…

Computation and Language · Computer Science 2014-11-04 Michał Łopuszyński , Łukasz Bolikowski

Knowledge is useless without structure. While the classification of knowledge has been an enduring philosophical enterprise, it recently found applications in computer science, notably for artificial intelligence. The availability of large…

Physics and Society · Physics 2018-03-06 Maxime Gabella

Wikipedia is a huge opportunity for machine learning, being the largest semi-structured base of knowledge available. Because of this, many works examine its contents, and focus on structuring it in order to make it usable in learning tasks,…

Machine Learning · Computer Science 2020-01-23 Tiphaine Viard , Thomas McLachlan , Hamidreza Ghader , Satoshi Sekine

Recent works on language identification and generation have established tight statistical rates at which these tasks can be achieved. These works typically operate under a strong realizability assumption: that the input data is drawn from…

Machine Learning · Computer Science 2026-04-23 Mikael Møller Høgsgaard , Chirag Pabbaraju

We study the task of generating from Wikipedia articles question-answer pairs that cover content beyond a single sentence. We propose a neural network approach that incorporates coreference knowledge via a novel gating mechanism. Compared…

Computation and Language · Computer Science 2018-05-16 Xinya Du , Claire Cardie

Knowledge bases are very good sources for knowledge extraction, the ability to create knowledge from structured and unstructured sources and use it to improve automatic processes as query expansion. However, extracting knowledge from…

Information Retrieval · Computer Science 2015-05-07 Joan Guisado-Gámez , Arnau Prat-Pérez

Defining psycholinguistic characteristics in written texts is a task gaining increasing attention from researchers. One of the most widely used tools in the current field is Linguistic Inquiry and Word Count (LIWC) that originally was…

Computation and Language · Computer Science 2026-01-29 Elina Sigdel , Anastasia Panfilova

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

Computation and Language · Computer Science 2026-03-05 Dan Saattrup Smart