English
Related papers

Related papers: Pantheon 1.0, a manually verified dataset of globa…

200 papers

The life trajectories of notable people have been studied to pinpoint the times and places of significant events such as birth, death, education, marriage, competition, work, speeches, scientific discoveries, artistic achievements, and…

Computation and Language · Computer Science 2025-06-10 Ying Zhang , Xiaofeng Li , Zhaoyang Liu , Haipeng Zhang

A biography of a person is the detailed description of several life events including his education, work, relationships, and death. Wikipedia, the free web-based encyclopedia, consists of millions of manually curated biographies of eminent…

Digital Libraries · Computer Science 2019-06-28 Heer Ambavi , Ayush Garg , Ayush Garg , Nitiksha , Mridul Sharma , Rohit Sharma , Jayesh Choudhari , Mayank Singh

At least since Priestley's 1765 Chart of Biography, large numbers of individual person records have been used to illustrate aggregate patterns of cultural history. Wikidata, the structured database sister of Wikipedia, currently contains…

Social and Information Networks · Computer Science 2015-06-23 Doron Goldfarb , Dieter Merkl , Maximilian Schich

The dataset focuses on Wikipedia users and contains information about demographic and socioeconomic characteristics of the respondents and their activity on Wikipedia. The data was collected using a questionnaire available online between…

Computers and Society · Computer Science 2023-12-06 Caterina Cruciani , Léo Joubert , Nicolas Jullien , Laurent Mell , Sasha Piccione , Jeanne Vermeirsche

Human activities can be seen as sequences of events, which are crucial to understanding societies. Disproportional event distribution for different demographic groups can manifest and amplify social stereotypes, and potentially jeopardize…

Computation and Language · Computer Science 2022-11-03 Jiao Sun , Nanyun Peng

In this paper we present the Wikipedia Cultural Diversity dataset. For each existing Wikipedia language edition, the dataset contains a classification of the articles that represent its associated cultural context, i.e. all concepts and…

Computers and Society · Computer Science 2019-06-11 Marc Miquel-Ribé , David Laniado

The evaluation of large language models faces significant challenges. Technical benchmarks often lack real-world relevance, while existing human preference evaluations suffer from unrepresentative sampling, superficial assessment depth, and…

Computation and Language · Computer Science 2026-03-06 Nora Petrova , Andrew Gordon , Enzo Blindow

We introduce a new reading comprehension dataset, dubbed MultiWikiQA, which covers 306 languages and has 1,220,757 samples in total. We start with Wikipedia articles, which also provide the context for the dataset samples, and use an LLM to…

Computation and Language · Computer Science 2026-03-05 Dan Saattrup Smart

Wikipedia is a community-created online encyclopedia; arguably, it is the most popular and largest knowledge resource on the Internet. Thus, reliability and neutrality are of high importance for Wikipedia. Previous research [3] reveals…

Computers and Society · Computer Science 2017-02-06 Olga Zagovora

Wikipedia (WP) as a collaborative, dynamical system of humans is an appropriate subject of social studies. Each single action of the members of this society, i.e. editors, is well recorded and accessible. Using the cumulative data of 34…

Physics and Society · Physics 2023-01-05 Taha Yasseri , Róbert Sumi , János Kertész

It is arguable whether history is made by great men and women or vice versa, but undoubtably social connections shape history. Analysing Wikipedia, a global collective memory place, we aim to understand how social links are recorded across…

Social and Information Networks · Computer Science 2012-07-05 Pablo Aragón , Andreas Kaltenbrunner , David Laniado , Yana Volkovich

Interactions among notable individuals -- whether examined individually, in groups, or as networks -- often convey significant messages across cultural, economic, political, scientific, and historical perspectives. By analyzing the times…

Social and Information Networks · Computer Science 2025-10-02 Zhongyang Liu , Ying Zhang , Xiangyi Xiao , Wenting Liu , Yuanting Zha , Haipeng Zhang

In this paper, we present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow…

Computers and Society · Computer Science 2021-04-01 Roldolfo Valentim , Giovanni Comarela , Souneil Park , Diego Saez-Trumper

With over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge. While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the lack of readability…

Computation and Language · Computer Science 2024-06-05 Mykola Trokhymovych , Indira Sen , Martin Gerlach

Wikipedia is a huge global repository of human knowledge, that can be leveraged to investigate interwinements between cultures. With this aim, we apply methods of Markov chains and Google matrix, for the analysis of the hyperlink networks…

Social and Information Networks · Computer Science 2015-03-06 Young-Ho Eom , Pablo Aragón , David Laniado , Andreas Kaltenbrunner , Sebastiano Vigna , Dima L. Shepelyansky

Among the manifold takes on world literature, it is our goal to contribute to the discussion from a digital point of view by analyzing the representation of world literature in Wikipedia with its millions of articles in hundreds of…

Information Retrieval · Computer Science 2017-01-05 Christoph Hube , Frank Fischer , Robert Jäschke , Gerhard Lauer , Mads Rosendahl Thomsen

To cope with the large number of publications, more and more researchers are automatically extracting data of interest using natural language processing methods based on supervised learning. Much data, especially in the natural and…

Computation and Language · Computer Science 2025-03-19 Jan Göpfert , Patrick Kuckertz , Jann M. Weinand , Detlef Stolten

We release 70,509 high-quality social networks extracted from multilingual fiction and nonfiction narratives. We additionally provide metadata for $\sim$30,000 of these texts (73\% nonfiction and 27\% fiction) written between 1800 and 1999…

Computation and Language · Computer Science 2025-04-01 Sil Hamilton , Rebecca M. M. Hicke , David Mimno , Matthew Wilkens

Sequence-to-sequence models have recently gained the state of the art performance in summarization. However, not too many large-scale high-quality datasets are available and almost all the available ones are mainly news articles with…

Computation and Language · Computer Science 2018-10-23 Mahnaz Koupaee , William Yang Wang

Several studies have used Wikipedia (WP) data-set to analyse worldwide human preferences by languages. However, those studies could suffer from bias related to exceptional social circumstances. Any massive event promoting the exceptional…

Physics and Society · Physics 2022-05-17 Julien Assuied , Yérali Gandica
‹ Prev 1 2 3 10 Next ›