English
Related papers

Related papers: Scraping SERPs for Archival Seeds: It Matters When…

200 papers

There are several ideas being used today for Web information retrieval, and specifically in Web search engines. The PageRank algorithm is one of those that introduce a content-neutral ranking function over Web pages. This ranking is applied…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-05-23 Giorgos Kollias , Efstratios Gallopoulos , Daniel B. Szyld

GoogleTrendArchive is a comprehensive archive of Google Trending Now data spanning over one year (from November 28, 2024 to January 3, 2026) across 125 countries and 1,358 locations. Unlike Google Trends, which requires specifying search…

Information Retrieval · Computer Science 2026-03-24 Aleksandra Urman , Anikó Hannák , Joachim Baumann

Twitter has become a leading source of real-time world-wide information and a great medium for exploring emerging events, breaking news and general topics which most matter to a broad audience. On the other hand, the explosive rate of…

Social and Information Networks · Computer Science 2017-04-04 Nazanin Dehghani , Masoud Asadpour

The Probability Ranking Principle (PRP) ranks search results based on their expected utility derived solely from document contents, often overlooking the nuances of presentation and user interaction. However, with the evolution of Search…

Information Retrieval · Computer Science 2024-04-02 Kanaad Pathak , Leif Azzopardi , Martin Halvey

Microblogging sites like Twitter and Weibo have emerged as important sourcesof real-time information on ongoing events, including socio-political events, emergency events, and so on. For instance, during emergency events (such as…

Information Retrieval · Computer Science 2017-07-20 Prannay Khosla , Moumita Basu , Kripabandhu Ghosh , Saptarshi Ghosh

Identifying and exploring emerging trends in the news is becoming more essential than ever with many changes occurring worldwide due to the global health crises. However, most of the recent research has focused mainly on detecting trends in…

Computation and Language · Computer Science 2023-01-27 Nhu Khoa Nguyen , Thierry Delahaut , Emanuela Boros , Antoine Doucet , Gaël Lejeune

Search engines today present results that are often oblivious to abrupt shifts in intent. For example, the query `independence day' usually refers to a US holiday, but the intent of this query abruptly changed during the release of a major…

Machine Learning · Computer Science 2010-07-23 Umar Syed , Aleksandrs Slivkins , Nina Mishra

Web search engines have marked everyone's life by transforming how one searches and accesses information. Search engines give special attention to the user interface, especially search engine result pages (SERP). The well-known ''10 blue…

Information Retrieval · Computer Science 2023-01-23 B. Oliveira , C. T. Lopes

The frequency of a web search keyword generally reflects the degree of public interest in a particular subject matter. Search logs are therefore useful resources for trend analysis. However, access to search logs is typically restricted to…

Social and Information Networks · Computer Science 2015-09-09 Mitsuo Yoshida , Yuki Arase , Takaaki Tsunoda , Mikio Yamamoto

This paper highlights the growing importance of information retrieval (IR) engines in the scientific community, addressing the inefficiency of traditional keyword-based search engines due to the rising volume of publications. The proposed…

Information Retrieval · Computer Science 2024-10-24 Mahsa Shamsabadi , Jennifer D'Souza

Sentiment signals derived from sparse news are commonly used in financial analysis and technology monitoring, yet transforming raw article-level observations into reliable temporal series remains a largely unsolved engineering problem.…

Machine Learning · Computer Science 2026-03-26 Stefania Stan , Marzio Lunghi , Vito Vargetto , Claudio Ricci , Rolands Repetto , Brayden Leo , Shao-Hong Gan

In this era of the Internet, the amount of news articles added every minute of everyday is humongous. As a result of this explosive amount of news articles, news retrieval systems are required to process the news articles frequently and…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-04-17 Arockia Anand Raj , T. Mala

Narratives are fundamental to our understanding of the world, providing us with a natural structure for knowledge representation over time. Computational narrative extraction is a subfield of artificial intelligence that makes heavy use of…

Computation and Language · Computer Science 2023-03-14 Brian Keith Norambuena , Tanushree Mitra , Chris North

Text extraction from web pages has many applications, including web crawling optimization and document clustering. Though much has been written about the acquisition of content from live web pages, content acquisition of archived web pages,…

Digital Libraries · Computer Science 2016-02-24 Shawn M. Jones , Harihar Shankar

The Environmental Governance and Data Initiative (EDGI) regularly crawled US federal environmental websites between 2016 and 2020 to capture changes between two presidential administrations. However, because it does not include the previous…

Digital Libraries · Computer Science 2025-06-02 Lesley Frew , Michael L. Nelson , Michele C. Weigle

Common Crawl is a multi-petabyte longitudinal dataset containing over 100 billion web pages which is widely used as a source of language data for sequence model training and in web science research. Each of its constituent archives is on…

Networking and Internet Architecture · Computer Science 2024-04-16 Henry S. Thompson

Generative document retrieval, an emerging paradigm in information retrieval, learns to build connections between documents and identifiers within a single model, garnering significant attention. However, there are still two challenges: (1)…

Information Retrieval · Computer Science 2024-05-14 Yong Guan , Dingxiao Liu , Jinchen Ma , Hao Peng , Xiaozhi Wang , Lei Hou , Ru Li

Search engine retrieval effectiveness studies are usually small-scale, using only limited query samples. Furthermore, queries are selected by the researchers. We address these issues by taking a random representative sample of 1,000…

Information Retrieval · Computer Science 2014-05-12 Dirk Lewandowski

Recent decisions to discontinue access to social media APIs are having detrimental effects on Internet research and the field of computational social science as a whole. This lack of access to data has been dubbed the Post-API era of…

Information Retrieval · Computer Science 2024-11-28 Amrit Poudel , Tim Weninger

We focus on the statistics of word occurrences and of the waiting times between such occurrences in Blogs. Due to the heterogeneity of words' frequencies, the empirical analysis is performed by studying classes of "frequently-equivalent"…

Information Theory · Computer Science 2012-09-25 R. Lambiotte , M. Ausloos , M. Thelwall
‹ Prev 1 3 4 5 6 7 10 Next ›