English
Related papers

Related papers: Tagging Scientific Publications using Wikipedia an…

200 papers

The instances of templates in Wikipedia form an interesting data set of structured information. Here I focus on the cite journal template that is primarily used for citation to articles in scientific journals. These citations can be…

Digital Libraries · Computer Science 2008-06-12 Finn Aarup Nielsen

Scientific information expresses human understanding of nature. This knowledge is largely disseminated in different forms of text, including scientific papers, news articles, and discourse among people on social media. While important for…

Computation and Language · Computer Science 2025-07-01 Dustin Wright

Wikipedia abstract generation aims to distill a Wikipedia abstract from web sources and has met significant success by adopting multi-document summarization techniques. However, previous works generally view the abstract as plain text,…

Computation and Language · Computer Science 2021-06-30 Fangwei Zhu , Shangqing Tu , Jiaxin Shi , Juanzi Li , Lei Hou , Tong Cui

This study investigates how YouTube content creators utilize scientific evidence in videos. Log-linear regression examines the influence of alternative communication channels on video creators in Biotechnology, using data from 81,302 papers…

Digital Libraries · Computer Science 2026-01-22 Pablo Dorta-González , María Isabel Dorta-González

Automatic quality evaluation of Web information is a task with many fields of applications and of great relevance, especially in critical domains like the medical one. We move from the intuition that the quality of content of medical Web…

Information Retrieval · Computer Science 2016-03-08 Vittoria Cozza , Marinella Petrocchi , Angelo Spognardi

To cope with the large number of publications, more and more researchers are automatically extracting data of interest using natural language processing methods based on supervised learning. Much data, especially in the natural and…

Computation and Language · Computer Science 2025-03-19 Jan Göpfert , Patrick Kuckertz , Jann M. Weinand , Detlef Stolten

Information Extraction is a well-researched area of Natural Language Processing with applications in web search and question answering concerned with identifying entities and relationships between them as expressed in a given context,…

Information Retrieval · Computer Science 2020-11-17 Erin Macdonald , Denilson Barbosa

Wikipedia categories, a classification scheme built for organizing and describing Wikpedia articles, are being applied in computer science research. This paper adopts a systematic literature review approach, in order to identify different…

Digital Libraries · Computer Science 2020-04-22 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón

Scientific writing is an iterative process that generates rich revision traces, yet publicly available resources typically expose only final or near-final versions of papers. This limits empirical study of revision behaviour and evaluation…

Computation and Language · Computer Science 2026-03-31 Léane Jourdan , Julien Aubert-Béduchaud , Yannis Chupin , Marah Baccari , Florian Boudin

Wikipedia is playing an increasingly central role on the web,and the policies its contributors follow when sourcing and fact-checking content affect million of readers. Among these core guiding principles, verifiability policies have a…

Computers and Society · Computer Science 2019-03-01 Miriam Redi , Besnik Fetahu , Jonathan Morgan , Dario Taraborelli

The scientific world is changing at a rapid pace, with new technology being developed and new trends being set at an increasing frequency. This paper presents a framework for conducting scientific analyses of academic publications, which is…

Computation and Language · Computer Science 2021-12-28 Trisha Singhal , Junhua Liu , Lucienne T. M. Blessing , Kwan Hui Lim

Scientific jargon can impede researchers when they read materials from other domains. Current methods of jargon identification mainly use corpus-level familiarity indicators (e.g., Simple Wikipedia represents plain language). However,…

Computation and Language · Computer Science 2023-11-17 Yue Guo , Joseph Chee Chang , Maria Antoniak , Erin Bransom , Trevor Cohen , Lucy Lu Wang , Tal August

Large text data sets, such as publications, websites, and other text-based media, inherit two distinct types of features: (1) the text itself, its information conveyed through semantics, and (2) its relationship to other texts through…

Computation and Language · Computer Science 2026-02-05 Tim Kunt , Annika Buchholz , Imene Khebouri , Thorsten Koch , Ida Litzel , Thi Huong Vu

We study the task of generating from Wikipedia articles question-answer pairs that cover content beyond a single sentence. We propose a neural network approach that incorporates coreference knowledge via a novel gating mechanism. Compared…

Computation and Language · Computer Science 2018-05-16 Xinya Du , Claire Cardie

Can the analysis of the semantics of words used in the text of a scientific paper predict its future impact measured by citations? This study details examples of automated text classification that achieved 80% success rate in distinguishing…

Computation and Language · Computer Science 2021-04-28 Neslihan Suzen , Alexander Gorban , Jeremy Levesley , Evgeny Mirkes

Tags assigned by users to shared content can be ambiguous. As a possible solution, we propose semantic tagging as a collaborative process in which a user selects and associates Web resources drawn from a knowledge context. We applied this…

Digital Libraries · Computer Science 2013-04-08 Bernhard Haslhofer , Werner Robitza , Carl Lagoze , Francois Guimbretiere

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

Information Retrieval · Computer Science 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

We introduce a new dataset named WikiVitals which contains a large graph of 48k mutually referred Wikipedia articles classified into 32 categories and connected by 2.3M edges. Our aim is to rigorously evaluate the contributions of three…

Machine Learning · Computer Science 2024-02-12 Pirmin Lemberger , Antoine Saillenfest

Hyperlinks and other relations in Wikipedia are a extraordinary resource which is still not fully understood. In this paper we study the different types of links in Wikipedia, and contrast the use of the full graph with respect to just…

Computation and Language · Computer Science 2015-03-16 Eneko Agirre , Ander Barrena , Aitor Soroa

Existing Natural Language Inference (NLI) datasets, while being instrumental in the advancement of Natural Language Understanding (NLU) research, are not related to scientific text. In this paper, we introduce SciNLI, a large dataset for…

Computation and Language · Computer Science 2022-03-16 Mobashir Sadat , Cornelia Caragea