English
Related papers

Related papers: Automatic Multi-Path Web Story Creation from a Str…

200 papers

We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. Relating visual and textual items on document pages has…

Wikipedia entity pages are a valuable source of information for direct consumption and for knowledge-base construction, update and maintenance. Facts in these entity pages are typically supported by references. Recent studies show that as…

Information Retrieval · Computer Science 2017-03-31 Besnik Fetahu , Katja Markert , Avishek Anand

Narrative visualization is a powerful communicative tool that can take on various formats such as interactive articles, slideshows, and data videos. These formats each have their strengths and weaknesses, but existing authoring tools only…

Human-Computer Interaction · Computer Science 2022-05-23 Matthew Conlen , Jeffrey Heer

Creating a cohesive, high-quality, relevant, media story is a challenge that news media editors face on a daily basis. This challenge is aggravated by the flood of highly relevant information that is constantly pouring onto the newsroom. To…

Multimedia · Computer Science 2021-10-14 Gonçalo Marcelino , David Semedo , André Mourão , Saverio Blasi , Marta Mrak , João Magalhães

Nowadays, editors tend to separate different subtopics of a long Wiki-pedia article into multiple sub-articles. This separation seeks to improve human readability. However, it also has a deleterious effect on many Wikipedia-based tasks that…

Information Retrieval · Computer Science 2019-06-24 Muhao Chen , Changping Meng , Gang Huang , Carlo Zaniolo

The different Wikipedia language editions vary dramatically in how comprehensive they are. As a result, most language editions contain only a small fraction of the sum of information that exists across all Wikipedias. In this paper, we…

Social and Information Networks · Computer Science 2016-04-13 Ellery Wulczyn , Robert West , Leila Zia , Jure Leskovec

Generating scientific manuscripts requires maintaining alignment between narrative reasoning, experimental evidence, and visual artifacts across the document lifecycle. Existing language-model generation pipelines rely on unconstrained text…

The DBpedia project extracts structured information from Wikipedia and makes it available on the web. Information is gathered mainly with the help of infoboxes that contain structured information of the Wikipedia article. A lot of…

Information Retrieval · Computer Science 2012-05-21 Daniel Hienert , Francesco Luciano

We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline with precise character…

Digital Libraries · Computer Science 2025-08-07 Samuel Barham , Chandler May , Benjamin Van Durme

Datasets for data-to-text generation typically focus either on multi-domain, single-sentence generation or on single-domain, long-form generation. In this work, we cast generating Wikipedia sections as a data-to-text generation task and…

Computation and Language · Computer Science 2021-06-03 Mingda Chen , Sam Wiseman , Kevin Gimpel

This paper presents a new way to increase interconnectivity in small Wikipedias (fewer than a 100,000 articles), by automatically linking articles based on interlanguage links. Many small Wikipedias have many articles with very few links,…

Social and Information Networks · Computer Science 2017-01-10 Michael Lotkowski

This paper proposes to tackle open- domain question answering using Wikipedia as the unique knowledge source: the answer to any factoid question is a text span in a Wikipedia article. This task of machine reading at scale combines the…

Computation and Language · Computer Science 2017-05-01 Danqi Chen , Adam Fisch , Jason Weston , Antoine Bordes

Wikipedia articles representing an entity or a topic in different language editions evolve independently within the scope of the language-specific user communities. This can lead to different points of views reflected in the articles, as…

Computation and Language · Computer Science 2017-02-03 Simon Gottschalk , Elena Demidova

Following a particular news story online is an important but difficult task, as the relevant information is often scattered across different domains/sources (e.g., news articles, blogs, comments, tweets), presented in various formats and…

Computation and Language · Computer Science 2018-08-20 Bichen Shi , Thanh-Binh Le , Neil Hurley , Georgiana Ifrim

Wikipedia is one of the most visited websites in the world and is also a frequent subject of scientific research. However, the analytical possibilities of Wikipedia information have not yet been analyzed considering at the same time both a…

Digital Libraries · Computer Science 2022-11-18 Wenceslao Arroyo-Machado , Daniel Torres-Salinas , Rodrigo Costas

This paper presents a novel story construction system to track the evolution of stories in an online fashion. The proposed system uses a novel sliding window solution, named Inching Window, allowing the processing of each new data element…

Social and Information Networks · Computer Science 2020-07-22 Özgür Can , Selma Tekir

Recent research has taken advantage of Wikipedia's multilingualism as a resource for cross-language information retrieval and machine translation, as well as proposed techniques for enriching its cross-language structure. The availability…

Databases · Computer Science 2011-11-01 Thanh Nguyen , Viviane Moreira , Huong Nguyen , Hoa Nguyen , Juliana Freire

Wikipedia is a popular web-based encyclopedia edited freely and collaboratively by its users. In this paper we present an analysis of Wikipedias in several languages as complex networks. The hyperlinks pointing from one Wikipedia article to…

Physics and Society · Physics 2009-11-11 V. Zlatic , M. Bozicevic , H. Stefancic , M. Domazet

Wikipedia has high-quality articles on a variety of topics and has been used in diverse research areas. In this study, a method is presented for using Wikipedia's editor information to build recommender systems in various domains that…

Information Retrieval · Computer Science 2023-06-16 Katsuhiko Hayashi

We describe our experience of implementing a news content organization system at Tencent that discovers events from vast streams of breaking news and evolves news story structures in an online fashion. Our real-world system has distinct…

Information Retrieval · Computer Science 2018-03-02 Bang Liu , Di Niu , Kunfeng Lai , Linglong Kong , Yu Xu