English
Related papers

Related papers: Exploratory Analysis of a Terabyte Scale Web Corpu…

200 papers

Many questions in computational social science rely on datasets assembled from heterogeneous online sources, a process that is often labor-intensive, costly, and difficult to reproduce. Recent advances in large language models enable…

Computation and Language · Computer Science 2026-01-07 Mengyi Sun

Twitter is one of the most prominent Online Social Networks. It covers a significant part of the online worldwide population~20% and has impressive growth rates. The social graph of Twitter has been the subject of numerous studies since it…

Social and Information Networks · Computer Science 2023-05-29 Despoina Antonakaki , Sotiris Ioannidis , Paraskevi Fragopoulou

Recent machine translation algorithms mainly rely on parallel corpora. However, since the availability of parallel corpora remains limited, only some resource-rich language pairs can benefit from them. We constructed a parallel corpus for…

Computation and Language · Computer Science 2020-03-17 Makoto Morishita , Jun Suzuki , Masaaki Nagata

Wikipedia is the largest online encyclopedia, used by algorithms and web users as a central hub of reliable information on the web. The quality and reliability of Wikipedia content is maintained by a community of volunteer editors. Machine…

Information Retrieval · Computer Science 2021-06-02 KayYen Wong , Miriam Redi , Diego Saez-Trumper

Exploratory analysis of a text corpus is essential for assessing data quality and developing meaningful hypotheses. Text analysis relies on understanding documents through structured attributes spanning various granularities of the…

Human-Computer Interaction · Computer Science 2025-04-24 Will Epperson , Arpit Mathur , Adam Perer , Dominik Moritz

Exploring large-scale text corpora presents a significant challenge in biomedical, finance, and legal domains, where vast amounts of documents are continuously published. Traditional search methods, such as keyword-based search, often…

Computation and Language · Computer Science 2025-06-18 Ashish Chouhan , Saifeldin Mandour , Michael Gertz

The models of the Internet reported in the literature are mainly aimed at reproducing the scale-free structure, the high clustering coefficient and the small world effects found in the real Internet, while other important properties (e.g.…

Physics and Society · Physics 2011-11-10 F. A. Rodrigues , P. R. Villas Boas , G. Travieso , L. da F. Costa

Online tracking has become of increasing concern in recent years, however our understanding of its extent to date has been limited to snapshots from web crawls. Previous at-tempts to measure the tracking ecosystem, have been done using…

Computers and Society · Computer Science 2020-07-07 Arjaldo Karaj , Sam Macbeth , Rémi Berson , Josep M. Pujol

Recent breakthroughs in large models have highlighted the critical significance of data scale, labels and modals. In this paper, we introduce MS MARCO Web Search, the first large-scale information-rich web dataset, featuring millions of…

Each complex network (or class of networks) presents specific topological features which characterize its connectivity and highly influence the dynamics of processes executed on the network. The analysis, discrimination, and synthesis of…

Disordered Systems and Neural Networks · Physics 2009-09-29 Luciano da F. Costa , Francisco A. Rodrigues , Gonzalo Travieso , P. R. Villas Boas

This paper is focused on the computational analysis of collective discourse, a collective behavior seen in non-expert content contributions in online social media. We collect and analyze a wide range of real-world collective discourse…

Social and Information Networks · Computer Science 2012-04-18 Vahed Qazvinian , Dragomir R. Radev

This paper surveys 60 English Machine Reading Comprehension datasets, with a view to providing a convenient resource for other researchers interested in this problem. We categorize the datasets according to their question and answer form…

Computation and Language · Computer Science 2021-10-11 Daria Dzendzik , Carl Vogel , Jennifer Foster

World Wide Web is a huge repository of web pages and links. It provides abundance of information for the Internet users. The growth of web is tremendous as approximately one million pages are added daily. Users' accesses are recorded in web…

Information Retrieval · Computer Science 2010-04-09 V. Chitraa , Dr. Antony Selvdoss Davamani

We propose to use MapReduce to quickly test new retrieval approaches on a cluster of machines by sequentially scanning all documents. We present a small case study in which we use a cluster of 15 low cost ma- chines to search a web crawl of…

Information Retrieval · Computer Science 2012-05-02 Djoerd Hiemstra , Claudia Hauff

Web crawlers visit internet applications, collect data, and learn about new web pages from visited pages. Web crawlers have a long and interesting history. Early web crawlers collected statistics about the web. In addition to collecting…

Official government publications are key sources for understanding the history of societies. Web publishing has fundamentally changed the scale and processes by which governments produce and disseminate information. Significantly, a range…

Digital Libraries · Computer Science 2021-12-07 Benjamin Charles Germain Lee , Trevor Owens

Indexing the Web of Data offers many opportunities, in particular, to find and explore data sources. One major design decision when indexing the Web of Data is to find a suitable index model, i.e., how to index and summarize data. Various…

Databases · Computer Science 2020-06-15 Till Blume , Ansgar Scherp

A powerful means to help users discover new content in the overwhelming amount of information available today is sharing in online communities such as social networks or crowdsourced platforms. This means comes short in the case of what we…

Human-Computer Interaction · Computer Science 2016-02-26 Giuseppe Scavo , Zied Ben Houidi , Stefano Traverso , Renata Teixeira , Marco Mellia

We present a preview of the Syntactic Acceptability Dataset, a resource being designed for both syntax and computational linguistics research. In its current form, the dataset comprises 1,000 English sequences from the syntactic discourse:…

Computation and Language · Computer Science 2025-06-24 Tom S Juzek

Popularized by the Differentiable Search Index, the emerging paradigm of generative retrieval re-frames the classic information retrieval problem into a sequence-to-sequence modeling task, forgoing external indices and encoding an entire…

Information Retrieval · Computer Science 2023-05-22 Ronak Pradeep , Kai Hui , Jai Gupta , Adam D. Lelkes , Honglei Zhuang , Jimmy Lin , Donald Metzler , Vinh Q. Tran
‹ Prev 1 4 5 6 7 8 10 Next ›