English
Related papers

Related papers: Scraping SERPs for Archival Seeds: It Matters When…

200 papers

This paper presents an analysis of the publication of datasets collected via Google Dataset Search, specialized in families of RNA viruses, whose terminology was obtained from the National Cancer Institute (NCI) thesaurus developed by the…

Digital Libraries · Computer Science 2021-01-12 Manuel Blázquez-Ochando , Juan-José Prieto-Gutiérrez

Search engines are the most commonly used type of tool for finding relevant information on the Internet. However, today's search engines are far from perfect. Typical search queries are short, often one or two words, and can be ambiguous…

Information Retrieval · Computer Science 2014-07-24 Dilip K. Limbu , Andy M. Connor , Stephen G. MacDonell

Predicting time-to-event outcomes in large databases can be a challenging but important task. One example of this is in predicting the time to a clinical outcome for patients in intensive care units (ICUs), which helps to support critical…

Computation · Statistics 2019-08-06 Yingying Xu , Joon Lee , Joel A. Dubin

Online news can quickly reach and affect millions of people, yet we do not know yet whether there exist potential dynamical regularities that govern their impact on the public. We use data from two major news outlets, BBC and New York…

Physics and Society · Physics 2021-01-25 Matúš Medo , Manuel S. Mariani , Linyuan Lü

Quantifying the captures of a URI over time is useful for researchers to identify the extent to which a Web page has been archived. Memento TimeMaps provide a format to list mementos (URI-Ms) for captures along with brief metadata, like…

Digital Libraries · Computer Science 2019-05-17 Mat Kelly , Lulwah M. Alkwai , Michael L. Nelson , Michele C. Weigle , Herbert Van de Sompel

Web search data are a valuable source of business and economic information. Previous studies have utilized Google Trends web search data for economic forecasting. We expand this work by providing algorithms to combine and aggregate search…

Econometrics · Economics 2018-03-28 Stephen L. France , Yuying Shi

We found that a simple property of clusters in a clustered dataset of news correlate strongly with importance and urgency of news (IUN) as assessed by LLM. We verified our finding across different news datasets, dataset sizes, clustering…

Computation and Language · Computer Science 2024-02-19 Oleg Vasilyev , John Bohannon

In scientific publications, citations allow readers to assess the authenticity of the presented information and verify it in the original context. News articles, however, do not contain citations and only rarely refer readers to further…

Information Retrieval · Computer Science 2019-09-27 Felix Hamborg , Philipp Meschenmoser , Moritz Schubotz , Bela Gipp

Tools such as Google News and Flipboard exist to convey daily news, but what about the past? In this paper, we describe how to combine several existing tools with web archive holdings to perform news analysis and visualization of the…

Digital Libraries · Computer Science 2021-03-23 Shawn M. Jones , Alexander C. Nwala , Martin Klein , Michele C. Weigle , Michael L. Nelson

In the past few decades, there has been an explosion in the amount of available data produced from various sources with different topics. The availability of this enormous data necessitates us to adopt effective computational tools to…

Computation and Language · Computer Science 2022-12-20 Mina Samizadeh

It has been suggested that online search and retrieval contributes to the intellectual isolation of users within their preexisting ideologies, where people's prior views are strengthened and alternative viewpoints are infrequently…

Information Retrieval · Computer Science 2014-05-08 Danai Koutra , Paul Bennett , Eric Horvitz

News articles capture a variety of topics about our society. They reflect not only the socioeconomic activities that happened in our physical world, but also some of the cultures, human interests, and public concerns that exist only in the…

Social and Information Networks · Computer Science 2018-09-11 Yingjie Hu , Xinyue Ye , Shih-Lung Shaw

Online social media platforms are turning into the prime source of news and narratives about worldwide events. However,a systematic summarization-based narrative extraction that can facilitate communicating the main underlying events is…

Social and Information Networks · Computer Science 2020-12-29 Toktam A. Oghaz , Ece C. Mutlu , Jasser Jasser , Niloofar Yousefi , Ivan Garibay

Production of news content is growing at an astonishing rate. To help manage and monitor the sheer amount of text, there is an increasing need to develop efficient methods that can provide insights into emerging content areas, and stratify…

Computation and Language · Computer Science 2020-10-29 M. Tarik Altuncu , Sophia N. Yaliraki , Mauricio Barahona

Systematic reviews (SRs) - the librarian-assisted literature survey of scholarly articles takes time and requires significant human resources. Given the ever-increasing volume of published studies, applying existing computing and…

Information Retrieval · Computer Science 2023-12-18 Kaushik Roy , Vedant Khandelwal , Harshul Surana , Valerie Vera , Amit Sheth , Heather Heckman

Generative models for Information Retrieval, where ranking of documents is viewed as the task of generating a query from a document's language model, were very successful in various IR tasks in the past. However, with the advent of modern…

Computation and Language · Computer Science 2020-10-08 Cicero Nogueira dos Santos , Xiaofei Ma , Ramesh Nallapati , Zhiheng Huang , Bing Xiang

Search engines play a central role in routing political information to citizens. The algorithmic personalization of search results by large search engines like Google implies that different users may be offered systematically different…

General Economics · Economics 2022-09-29 Ulrich Matter , Roland Hodler , Johannes Ladwig

Background. From information theory, surprisal is a measurement of how unexpected an event is. Statistical language models provide a probabilistic approximation of natural languages, and because surprisal is constructed with the probability…

Computation and Language · Computer Science 2022-04-18 James Caddy , Markus Wagner , Christoph Treude , Earl T. Barr , Miltiadis Allamanis

Online encyclopedia such as Wikipedia has become one of the best sources of knowledge. Much effort has been devoted to expanding and enriching the structured data by automatic information extraction from unstructured text in Wikipedia.…

Information Retrieval · Computer Science 2014-06-26 Kezun Zhang , Yanghua Xiao , Hanghang Tong , Haixun Wang , Wei Wang

Traditional information retrieval (IR) ranking models process the full text of documents. Newer models based on Transformers, however, would incur a high computational cost when processing long texts, so typically use only snippets from the…

Information Retrieval · Computer Science 2022-01-24 Gabriella Kazai , Bhaskar Mitra , Anlei Dong , Nick Craswell , Linjun Yang