English
Related papers

Related papers: Scraping SERPs for Archival Seeds: It Matters When…

200 papers

A growing number of people are changing the way they consume news, replacing the traditional physical newspapers and magazines by their virtual online versions or/and weblogs. The interactivity and immediacy present in online news are…

Computers and Society · Computer Science 2015-04-17 Julio Reis , Fabrıcio Benevenuto , Pedro O. S. Vaz de Melo , Raquel Prates , Haewoon Kwak , Jisun An

This paper critically audits the search endpoint of YouTube's Data API (v3), a common tool for academic research. Through systematic weekly searches over six months using eleven queries, we identify major limitations regarding completeness,…

Information Retrieval · Computer Science 2025-11-25 Bernhard Rieder , Adrian Padilla , Oscar Coromina

Generative Retrieval (GR) is an emerging paradigm in information retrieval that leverages generative models to directly map queries to relevant document identifiers (DocIDs) without the need for traditional query processing or document…

Information Retrieval · Computer Science 2024-06-05 Tzu-Lin Kuo , Tzu-Wei Chiu , Tzung-Sheng Lin , Sheng-Yang Wu , Chao-Wei Huang , Yun-Nung Chen

Portfolio diversification and active risk management are essential parts of financial analysis which became even more crucial (and questioned) during and after the years of the Global Financial Crisis. We propose a novel approach to…

Portfolio Management · Quantitative Finance 2013-10-08 Ladislav Kristoufek

In Twitter, and other microblogging services, the generation of new content by the crowd is often biased towards immediacy: what is happening now. Prompted by the propagation of commentary and information through multiple mediums, users on…

Information Retrieval · Computer Science 2016-02-10 Flávio Martins , João Magalhães , Jamie Callan

Search techniques make use of elementary information such as term frequencies and document lengths in computation of similarity weighting. They can also exploit richer statistics, in particular the number of documents in which any two terms…

Information Retrieval · Computer Science 2020-07-20 Bodo Billerbeck , Justin Zobel , Nicholas Lester , Nick Craswell

This study examines Facebook and YouTube content from over a thousand news outlets in four European languages from 2018 to 2023, using a Bayesian structural time-series model to evaluate the impact of viral posts. Our results show that most…

Social and Information Networks · Computer Science 2024-07-19 Emanuele Sangiorgio , Niccolò Di Marco , Gabriele Etta , Matteo Cinelli , Roy Cerqueti , Walter Quattrociocchi

Over two hundreds health awareness events take place in the United States in order to raise attention and educate the public about diseases. It would be informative and instructive for the organization to know the impact of these events,…

Applications · Statistics 2018-10-10 Zheng Hao , Miao Liu , Xijin Ge

The vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose…

We consider the problem of detecting the source of a rumor which has spread in a network using only observations about which set of nodes are infected with the rumor and with no information as to \emph{when} these nodes became infected. In…

Probability · Mathematics 2015-11-04 Devavrat Shah , Tauhid Zaman

Search engines are a combination of hardware and computer software supplied by a particular company through the website which has been determined. Search engines collect information from the web through bots or web crawlers that crawls the…

Information Retrieval · Computer Science 2014-10-22 Ahmad Josi , Leon Andretti Abdillah , Suryayusra

The Web is ephemeral. Many resources have representations that change over time, and many of those representations are lost forever. A lucky few manage to reappear as archived resources that carry their own URIs. For example, some content…

Search result snippets are crucial in modern search engines, providing users with a quick overview of a website's content. Snippets help users determine the relevance of a document to their information needs, and in certain scenarios even…

Information Retrieval · Computer Science 2024-01-30 Anat Hashavit , Tamar Stern , Hongning Wang , Sarit Kraus

Lifelong experiences and learned knowledge lead to shared expectations about how common situations tend to unfold. Such knowledge of narrative event flow enables people to weave together a story. However, comparable computational tools to…

Computation and Language · Computer Science 2022-07-12 Maarten Sap , Anna Jafarpour , Yejin Choi , Noah A. Smith , James W. Pennebaker , Eric Horvitz

Conspiracy theories, as a type of misinformation, are narratives that explains an event or situation in an irrational or malicious manner. While most previous work examined conspiracy theory in social media short texts, limited attention…

Computation and Language · Computer Science 2023-10-31 Yuanyuan Lei , Ruihong Huang

Unsupervised discovery of stories with correlated news articles in real-time helps people digest massive news streams without expensive human annotations. A common approach of the existing studies for unsupervised online story discovery is…

Information Retrieval · Computer Science 2023-05-05 Susik Yoon , Dongha Lee , Yunyi Zhang , Jiawei Han

As the exploration of digital behavioral data revolutionizes communication research, understanding the nuances of data collection methodologies becomes increasingly pertinent. This study focuses on one prominent data collection approach,…

Computers and Society · Computer Science 2024-12-03 Roberto Ulloa , Frank Mangold , Felix Schmidt , Judith Gilsbach , Sebastian Stier

We present a new machine learning and text information extraction approach to detection of cyber threat events in Twitter that are novel (previously non-extant) and developing (marked by significance with respect to similarity with a…

Information Retrieval · Computer Science 2019-07-19 Avishek Bose , Vahid Behzadan , Carlos Aguirre , William H. Hsu

Google Trends is a tool that allows researchers to analyze the popularity of Google search queries across time and space. In a single request, users can obtain time series for up to 5 queries on a common scale, normalized to the range from…

Social and Information Networks · Computer Science 2021-02-05 Robert West

With the rapid advance of the Internet, search engines (e.g., Google, Bing, Yahoo!) are used by billions of users for each day. The main function of a search engine is to locate the most relevant webpages corresponding to what the user…

Applications · Statistics 2018-03-15 Xinzhi Han , Sen Lei