English
Related papers

Related papers: Scraping SERPs for Archival Seeds: It Matters When…

200 papers

Twitter is among the commonest sources of data employed in social media research mainly because of its convenient APIs to collect tweets. However, most researchers do not have access to the expensive Firehose and Twitter Historical Archive,…

Computers and Society · Computer Science 2016-11-28 Daniel Gayo-Avello

In this paper, we propose a web search retrieval approach which automatically detects recency sensitive queries and increases the freshness of the ordinary document ranking by a degree proportional to the probability of the need in recent…

Information Retrieval · Computer Science 2024-02-08 Andrey Styskin , Fedor Romanenko , Fedor Vorobyev , Pavel Serdyukov

Twitter, like many social media and data brokering companies, makes their data available through a search API (application programming interface). In addition to filtering results by date and location, researchers can search for tweets with…

Social and Information Networks · Computer Science 2020-06-23 Emory Hufbauer , Hana Khamfroush

Twitter updates now represent an enormous stream of information originating from a wide variety of formal and informal sources, much of which is relevant to real-world events. In this paper we adapt existing bio-surveillance algorithms to…

Social and Information Networks · Computer Science 2015-04-10 Nicholas Thapen , Donal Simmie , Chris Hankin

This paper investigates the composition of search engine results pages. We define what elements the most popular web search engines use on their results pages (e.g., organic results, advertisements, shortcuts) and to which degree they are…

Information Retrieval · Computer Science 2015-11-19 Nadine Hoechstoetter , Dirk Lewandowski

Millions of news articles from hundreds of thousands of sources around the globe appear in news aggregators every day. Consuming such a volume of news presents an almost insurmountable challenge. For example, a reader searching on…

Computation and Language · Computer Science 2020-06-02 Joshua Bambrick , Minjie Xu , Andy Almonte , Igor Malioutov , Guim Perarnau , Vittorio Selo , Iat Chong Chan

We extend the concept of Named Entities to Named Events - commonly occurring events such as battles and earthquakes. We propose a method for finding specific passages in news articles that contain information about such events and report…

Computation and Language · Computer Science 2013-06-21 Luis Marujo , Wang Ling , Anatole Gershman , Jaime Carbonell , João P. Neto , David Matos

Time is an important relevance signal when searching streams of social media posts. The distribution of document timestamps from the results of an initial query can be leveraged to infer the distribution of relevant documents, which can…

Information Retrieval · Computer Science 2017-07-26 Jinfeng Rao , Hua He , Haotian Zhang , Ferhan Ture , Royal Sequiera , Salman Mohammed , Jimmy Lin

Secondary analysis or the reuse of existing survey data is a common practice among social scientists. Searching for relevant datasets in Digital Libraries is a somehow unfamiliar behaviour for this community. Dataset retrieval, especially…

Digital Libraries · Computer Science 2020-10-13 Zeljko Carevic , Dwaipayan Roy , Philipp Mayr

People post information about different topics which are in their active vocabulary over social media platforms (like Twitter, Facebook, PInterest and Google+). They follow each other and it is more likely that the person who posts…

Social and Information Networks · Computer Science 2022-08-30 Muskan Garg

The news coverage of events often contains not one but multiple incompatible accounts of what happened. We develop a query-based system that extracts compatible sets of events (scenarios) from such data, formulated as one-class clustering.…

Computation and Language · Computer Science 2019-09-17 Su Wang , Greg Durrett , Katrin Erk

It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evolution of cultural…

Physics and Society · Physics 2020-05-28 Eitan Adam Pechenick , Christopher M. Danforth , Peter Sheridan Dodds

A stream of unstructured news can be a valuable source of hidden relations between different entities, such as financial institutions, countries, or persons. We present an approach to continuously collect online news, recognize relevant…

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

Computation and Language · Computer Science 2023-08-25 Emily Silcock , Melissa Dell

People increasingly use microblogging platforms such as Twitter during natural disasters and emergencies. Research studies have revealed the usefulness of the data available on Twitter for several disaster response tasks. However, making…

Social and Information Networks · Computer Science 2018-05-16 Firoj Alam , Ferda Ofli , Muhammad Imran , Michael Aupetit

Retrieving information from social networks is the first and primordial step many data analysis fields such as Natural Language Processing, Sentiment Analysis and Machine Learning. Important data science tasks relay on historical data…

Information Retrieval · Computer Science 2018-03-28 A. Hernandez-Suarez , G. Sanchez-Perez , K. Toscano-Medina , V. Martinez-Hernandez , V. Sanchez , H. Perez-Meana

Document ranking experiments should be repeatable. However, the interaction between multi-threaded indexing and score ties during retrieval may yield non-deterministic rankings, making repeatability not as trivial as one might imagine. In…

Information Retrieval · Computer Science 2019-09-04 Jimmy Lin , Peilin Yang

Earthquakes have a deep impact on wide areas, and emergency rescue operations may benefit from social media information about the scope and extent of the disaster. Therefore, this work presents a text miningbased approach to collect and…

Computation and Language · Computer Science 2022-12-14 Zhe Zheng , Hong-Zheng Shi , Yu-Cheng Zhou , Xin-Zheng Lu , Jia-Rui Lin

Writers such as journalists often use automatic tools to find relevant content to include in their narratives. In this paper, we focus on supporting writers in the news domain to develop event-centric narratives. Given an incomplete…

Computation and Language · Computer Science 2021-07-01 Nikos Voskarides , Edgar Meij , Sabrina Sauer , Maarten de Rijke

Search engines play an important role in the context of modern elections. By curating information in response to user queries, search engines influence how individuals are informed about election-related developments and perceive the media…

Computers and Society · Computer Science 2025-01-10 Mykola Makhortykh , Tobias Rorhbach , Maryna Sydorova , Elizaveta Kuznetsova