English
Related papers

Related papers: Matching and Linking Entries in Historical Swedish…

200 papers

In this paper, we describe the extraction of all the location entries from a prominent Swedish encyclopedia from the early 20th century, the \textit{Nordisk Familjebok} `Nordic Family Book.' We focused on the second edition called…

Computation and Language · Computer Science 2024-06-27 Axel Ahlin , Alfred Myrne , Pierre Nugues

The digitization of old encyclopedias represents an important step to improve access to historically structured knowledge. Often, however, this process does not go beyond an optical character recognition, leaving all the underlying…

Computation and Language · Computer Science 2026-05-05 Albin Andersson , Salam Jonasson , Fredrik Wastring , Pierre Nugues

Diderot's \textit{Encyclop\'edie} is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era. \textit{Wikipedia} has the same ambition with a much greater scope. However, the lack of digital…

Computation and Language · Computer Science 2024-06-06 Pierre Nugues

The \textit{Petit Larousse illustr\'e} is a French dictionary first published in 1905. Its division in two main parts on language and on history and geography corresponds to a major milestone in French lexicography as well as a repository…

Computation and Language · Computer Science 2022-08-02 Pierre Nugues

The aim of the study was to assess and compare established search systems and approaches by using the search goal of identifying the first (as in oldest) nursing-related document, with reference to the first Swedish-affiliated document, in…

Digital Libraries · Computer Science 2023-11-28 Christopher Holmberg

This report summarizes the results of a short-term student research project focused on the usage of Swedish Wikipedia. It is trying to answer the following question: To what extent (and why) do people from non-English language communities…

Physics and Society · Physics 2013-08-09 Berit Schreck , Mirko Kämpf , Jan W. Kantelhardt , Holger Motzkau

Many Swedish benchmarks are translations of US-centric benchmarks and are therefore not suitable for testing knowledge that is particularly relevant, or even specific, to Sweden. We therefore introduce a manually written question-answering…

Computation and Language · Computer Science 2026-03-03 Jenny Kunz

Wikipedia is a huge global repository of human knowledge, that can be leveraged to investigate interwinements between cultures. With this aim, we apply methods of Markov chains and Google matrix, for the analysis of the hyperlink networks…

Social and Information Networks · Computer Science 2015-03-06 Young-Ho Eom , Pablo Aragón , David Laniado , Andreas Kaltenbrunner , Sebastiano Vigna , Dima L. Shepelyansky

We present a large Norwegian lexical resource of categorized medical terms. The resource merges information from large medical databases, and contains over 77,000 unique entries, including automatically mapped terms from a Norwegian medical…

Computation and Language · Computer Science 2020-04-07 Ildikó Pilán , Pål H. Brekke , Lilja Øvrelid

We analyze the information provided by the word embeddings about the grammatical gender in Swedish. We wish that this paper may serve as one of the bridges to connect the methods of computational linguistics and general linguistics. Taking…

Computation and Language · Computer Science 2020-07-29 Marc Allassonnière-Tang , Ali Basirat

This article presents a large-scale effort to create a structured dataset of internal migration in Finland between 1800 and 1920 using digitized church moving records. These records, maintained by Evangelical-Lutheran parishes, document the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Ari Vesalainen , Jenna Kanerva , Aida Nitsch , Kiia Korsu , Ilari Larkiola , Laura Ruotsalainen , Filip Ginter

Research in the social sciences is increasingly based on large and complex data collections, where individual data sets from different domains are linked and integrated to allow advanced analytics. A popular type of data used in such a…

Databases · Computer Science 2018-07-09 Charini Nanayakkara , Peter Christen , Thilina Ranbaduge

Information presented in Wikipedia articles must be attributable to reliable published sources in the form of references. This study examines over 5 million Wikipedia articles to assess the reliability of references in multiple language…

Computers and Society · Computer Science 2023-09-06 Aitolkyn Baigutanova , Diego Saez-Trumper , Miriam Redi , Meeyoung Cha , Pablo Aragón

Background: Clinical natural language processing (NLP) refers to the use of computational methods for extracting, processing, and analyzing unstructured clinical text data, and holds a huge potential to transform healthcare in various…

Most datasets in the field of document analysis utilize highly standardized labels, which, while simplifying specific tasks, often produce outputs that are not directly applicable to humanities research. In contrast, the Nuremberg…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Martin Mayr , Julian Krenz , Katharina Neumeier , Anna Bub , Simon Bürcky , Nina Brolich , Klaus Herbers , Mechthild Habermann , Peter Fleischmann , Andreas Maier , Vincent Christlein

In this paper, we present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow…

Computers and Society · Computer Science 2021-04-01 Roldolfo Valentim , Giovanni Comarela , Souneil Park , Diego Saez-Trumper

In this paper we focus on the beginning of publication of the Large Soviet Encyclopedia, launched in 1925. We present the context of this launching and explain why it was tightly connected to the period of the New Economical Policy. In a…

History and Overview · Mathematics 2016-02-18 Laurent Mazliak

We investigated the evolution and transformation of scientific knowledge in the early modern period, analyzing more than 350 different editions of textbooks used for teaching astronomy in European universities from the late fifteenth…

History and Philosophy of Physics · Physics 2020-04-02 Maryam Zamani , Alejandro Tejedor , Malte Vogl , Florian Krautli , Matteo Valleriani , Holger Kantz

This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel…

Computation and Language · Computer Science 2024-10-08 Tobias Norlund , Tim Isbister , Amaru Cuba Gyllensten , Paul Dos Santos , Danila Petrelli , Ariel Ekgren , Magnus Sahlgren

The rapid growth of scientific publishing has made it increasingly difficult to track how fast-moving areas evolve. Search engines and LLM-based assistants retrieve or summarize papers, but often hide how the corpus was selected, organized,…

Information Retrieval · Computer Science 2026-05-28 Bernardo A. Denkvitts , Nitin Gupta , Biplav Srivastava
‹ Prev 1 2 3 10 Next ›