English
Related papers

Related papers: Media Cloud: Massive Open Source Collection of Glo…

200 papers

Over the past few years, we have built a system that has exposed large volumes of Deep-Web content to Google.com users. The content that our system exposes contributes to more than 1000 search queries per-second and spans over 50 languages…

Databases · Computer Science 2009-09-15 Jayant Madhavan , Loredana Afanasiev , Lyublena Antova , Alon Halevy

Social network has gained remarkable attention in the last decade. Accessing social network sites such as Twitter, Facebook LinkedIn and Google+ through the internet and the web 2.0 technologies has become more affordable. People are…

Social and Information Networks · Computer Science 2023-06-22 Mariam Adedoyin-Olowe , Mohamed Medhat Gaber , Frederic Stahl

Understanding the flow of information across today's fragmented digital media landscape requires scalable, cross-platform infrastructure. In this paper, we present the Canadian Media Ecosystem Observatory, a national-scale infrastructure…

Cloud computing is a complex infrastructure of software, hardware, processing, and storage that is available as a service. Cloud computing offers immediate access to large numbers of the world's most sophisticated supercomputers and their…

Cryptography and Security · Computer Science 2013-03-07 Sugata Sanyal , Parthasarathy P. Iyer

ClueWeb22, the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information. Its design was influenced by the need for a high quality, large scale web corpus to support a range of academic…

Information Retrieval · Computer Science 2022-12-05 Arnold Overwijk , Chenyan Xiong , Xiao Liu , Cameron VandenBerg , Jamie Callan

Event collections are frequently built by crawling the live web on the basis of seed URIs nominated by human experts. Focused web crawling is a technique where the crawler is guided by reference content pertaining to the event. Given the…

Digital Libraries · Computer Science 2018-04-06 Martin Klein , Lyudmila Balakireva , Herbert Van de Sompel

We are presenting a set of multilingual text analysis tools that can help analysts in any field to explore large document collections quickly in order to determine whether the documents contain information of interest, and to find the…

Computation and Language · Computer Science 2007-05-23 Camelia Ignat , Bruno Pouliquen , Ralf Steinberger , Tomaz Erjavec

The excessive amounts of data generated by devices and Internet-based sources at a regular basis constitute, big data. This data can be processed and analyzed to develop useful applications for specific domains. Several mathematical and…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-12-15 Samiya Khan , Kashish Ara Shakil , Mansaf Alam

GitHub is the largest code hosting platform, with millions of repositories spanning multiple technologies. Despite this, little is known about the actual contents of GitHub's repositories in the wild. This paper presents an initial…

Software Engineering · Computer Science 2026-05-19 Andre Hora , João Eduardo Montandon , Diego Elias Costa

A large amount of data on the WWW remains inaccessible to crawlers of Web search engines because it can only be exposed on demand as users fill out and submit forms. The Hidden web refers to the collection of Web data which can be accessed…

Information Retrieval · Computer Science 2014-07-23 Sonali Gupta , Komal Kumar Bhatia

Web browsers are increasingly used as middleware platforms offering a central access point for service provision. Using backend containerization, RESTful APIs, and distributed computing allows for complex systems to be realized that address…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-10-03 Rudolph Pienaar , Ata Turk , Jorge Bernal-Rusiel , Nicolas Rannou , Daniel Haehn , P. Ellen Grant , Orran Krieger

The complexity and diversity of today's media landscape provides many challenges for researchers studying news producers. These producers use many different strategies to get their message believed by readers through the writing styles they…

Computers and Society · Computer Science 2018-08-17 Benjamin D. Horne , William Dron , Sara Khedr , Sibel Adali

We are presenting a text analysis tool set that allows analysts in various fields to sieve through large collections of multilingual news items quickly and to find information that is of relevance to them. For a given document collection,…

Computation and Language · Computer Science 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Camelia Ignat

Modern technologies are enabling scientists to collect extraordinary amounts of complex and sophisticated data across a huge range of scales like never before. With this onslaught of data, we can allow the focal point to shift towards…

Applications requiring real-time processing of large volumes of data have been the main driver for rethinking the traditional cloud, giving rise to novel cloud models. Distributed cloud (DC) is a model that allows users to dynamically…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-06 Tamara Ranković , Mateja Rilak , Janko Rakonjac , Miloš Simić

The idea of a social cloud has emerged as a resource sharing paradigm in a social network context. Undoubtedly, state-of-the-art social cloud systems demonstrate the potential of the social cloud acting as complementary to other computing…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-03-05 Pramod C. Mane , Kapil Ahuja , Pradeep Singh

Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources.…

Digital Libraries · Computer Science 2017-05-18 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

The Data Web refers to the vast and rapidly increasing quantity of scientific, corporate, government and crowd-sourced data published in the form of Linked Open Data, which encourages the uniform representation of heterogeneous data items…

In this paper we describe the design, and implementation of the Open Science Data Cloud, or OSDC. The goal of the OSDC is to provide petabyte-scale data cloud infrastructure and related services for scientists working with large quantities…

This paper presents the first framework for integrating procedural knowledge, or "know-how", into the Linked Data Cloud. Know-how available on the Web, such as step-by-step instructions, is largely unstructured and isolated from other…

Artificial Intelligence · Computer Science 2016-04-18 Paolo Pareti , Benoit Testu , Ryutaro Ichise , Ewan Klein , Adam Barker