English
Related papers

Related papers: TWikiL -- The Twitter Wikipedia Link Dataset

200 papers

With social media becoming increasingly pop-ular on which lots of news and real-time eventsare reported, developing automated questionanswering systems is critical to the effective-ness of many applications that rely on real-time knowledge.…

Computation and Language · Computer Science 2019-07-16 Wenhan Xiong , Jiawei Wu , Hong Wang , Vivek Kulkarni , Mo Yu , Shiyu Chang , Xiaoxiao Guo , William Yang Wang

Despite recent progress in computer vision, fine-grained interpretation of satellite images remains challenging because of a lack of labeled training data. To overcome this limitation, we propose using Wikipedia as a previously untapped…

Computer Vision and Pattern Recognition · Computer Science 2018-09-28 Evan Sheehan , Burak Uzkent , Chenlin Meng , Zhongyi Tang , Marshall Burke , David Lobell , Stefano Ermon

The production and consumption of information about Bitcoin and other digital-, or 'crypto'-, currencies have grown together with their market capitalisation. However, a systematic investigation of the relationship between online attention…

Physics and Society · Physics 2020-05-22 Abeer ElBahrawy , Laura Alessandretti , Andrea Baronchelli

In this paper we present our web application SeRE designed to explore semantically related concepts. Wikipedia and DBpedia are rich data sources to extract related entities for a given topic, like in- and out-links, broader and narrower…

Computation and Language · Computer Science 2015-04-28 Daniel Hienert , Dennis Wegener , Siegfried Schomisch

Online government petitions represent a new data-rich mode of political participation. This work examines the thus far understudied dynamics of sharing petitions on social media in order to garner signatures and, ultimately, a government…

Social and Information Networks · Computer Science 2021-01-05 Peter Cihon , Taha Yasseri , Scott Hale , Helen Margetts

While a plethora of hypertext links exist on the Web, only a small amount of them are regularly clicked. Starting from this observation, we set out to study large-scale click data from Wikipedia in order to understand what makes a link…

Social and Information Networks · Computer Science 2017-02-21 Dimitar Dimitrov , Philipp Singer , Florian Lemmerich , Markus Strohmaier

By linking to external websites, Wikipedia can act as a gateway to the Web. To date, however, little is known about the amount of traffic generated by Wikipedia's external links. We fill this gap in a detailed analysis of usage logs…

Computers and Society · Computer Science 2021-02-16 Tiziano Piccardi , Miriam Redi , Giovanni Colavizza , Robert West

Video communication has been rapidly increasing over the past decade, with YouTube providing a medium where users can post, discover, share, and react to videos. There has also been an increase in the number of videos citing research…

Digital Libraries · Computer Science 2022-09-07 Abdul Rahman Shaikh , Hamed Alhoori , Maoyuan Sun

Wikipedia relies on an extensive review process to verify that the content of each individual page is unbiased and presents a neutral point of view. Less attention has been paid to possible biases in the hyperlink structure of Wikipedia,…

Social and Information Networks · Computer Science 2022-04-04 Cristina Menghini , Aris Anagnostopoulos , Eli Upfal

Wikipedia is the largest source of free encyclopedic knowledge and one of the most visited sites on the Web. To increase reader understanding of the article, Wikipedia editors add images within the text of the article's body. However,…

Computers and Society · Computer Science 2021-12-06 Daniele Rama , Tiziano Piccardi , Miriam Redi , Rossano Schifanella

We describe NatCat, a large-scale resource for text classification constructed from three data sources: Wikipedia, Stack Exchange, and Reddit. NatCat consists of document-category pairs derived from manual curation that occurs naturally…

Computation and Language · Computer Science 2021-09-21 Zewei Chu , Karl Stratos , Kevin Gimpel

We describe a knowledge graph derived from Twitter data with the goal of discovering relationships between people, links, and topics. The goal is to filter out noise from Twitter and surface an inside-out view that relies on high quality…

Information Retrieval · Computer Science 2019-06-17 Omar Alonso , Vasileios Kandylas , Serge-Eric Tremblay

In the past decade, the DBpedia community has put significant amount of effort on developing technical infrastructure and methods for efficient extraction of structured information from Wikipedia. These efforts have been primarily focused…

Computation and Language · Computer Science 2018-12-27 Milan Dojchinovski , Julio Hernandez , Markus Ackermann , Amit Kirschenbaum , Sebastian Hellmann

We present TweeNLP, a one-stop portal that organizes Twitter's natural language processing (NLP) data and builds a visualization and exploration platform. It curates 19,395 tweets (as of April 2021) from various NLP conferences and general…

Computation and Language · Computer Science 2021-06-22 Viraj Shah , Shruti Singh , Mayank Singh

As a platform, Twitter has been a significant public space for discussion related to the COVID-19 pandemic. Public social media platforms such as Twitter represent important sites of engagement regarding the pandemic and these data can be…

Social and Information Networks · Computer Science 2021-01-29 Hassan Dashtian , Dhiraj Murthy

Extracting useful information from the user history to clearly understand informational needs is a crucial feature of a proactive information retrieval system. Regarding understanding information and relevance, Wikipedia can provide the…

Information Retrieval · Computer Science 2022-10-19 Tabish Ahmed , Sahan Bulathwela

We show that information about social relationships can be used to improve user-level sentiment analysis. The main motivation behind our approach is that users that are somehow "connected" may be more likely to hold similar opinions;…

Computation and Language · Computer Science 2011-09-29 Chenhao Tan , Lillian Lee , Jie Tang , Long Jiang , Ming Zhou , Ping Li

Traditional disease surveillance systems suffer from several disadvantages, including reporting lags and antiquated technology, that have caused a movement towards internet-based disease surveillance systems. Internet systems are…

Information Retrieval · Computer Science 2015-08-26 Geoffrey Fairchild , Lalindra De Silva , Sara Y. Del Valle , Alberto M. Segre

Wikipedia is widely used for finding general information about a wide variety of topics. Its vocation is not to provide local information. For example, it provides plot, cast, and production information about a given movie, but not showing…

Information Retrieval · Computer Science 2016-05-31 Gregory Grefenstette , Karima Rafes

Geopolitics focuses on political power in relation to geographic space. Interactions among world countries have been widely studied at various scales, observing economic exchanges, world history or international politics among others. This…

Social and Information Networks · Computer Science 2017-07-21 Klaus M. Frahm , Samer El Zant , Katia Jaffrès-Runser , Dima L. Shepelyansky