English
Related papers

Related papers: Exploratory Analysis of a Terabyte Scale Web Corpu…

200 papers

In this paper we provide two introductory analyses of CAPs, based exclusively on the analysis of documents found on the Internet. The first analysis allowed us to investigate the world of CAPs, in particular for what concerned their status…

Human-Computer Interaction · Computer Science 2016-09-16 Giovanna Pacini , Franco Bagnoli

In this paper the preliminary results of a literature review on characteristics used to define continuous experiments are presented. In total 14 papers were selected. The results were synthesized into a model that gives an overview of all…

Software Engineering · Computer Science 2019-12-11 Florian Auer , Michael Felderer

The web is the prominent way information is exchanged in the 21st century. However, ensuring web-based information is accessible is complicated, particularly with web applications that rely on JavaScript and other technologies to deliver…

Information Retrieval · Computer Science 2019-08-09 Trevor Bostic , Jeff Stanley , John Higgins , Rachael L. Bradley-Montgomery , Justin F. Brunelle , Daniel Chudnov

Wikipedia is the largest existing knowledge repository that is growing on a genuine crowdsourcing support. While the English Wikipedia is the most extensive and the most researched one with over five million articles, comparatively little…

Digital Libraries · Computer Science 2017-10-20 Kristina Ban , Matjaz Perc , Zoran Levnajic

The availability of large data sets is providing an impetus for driving current artificial intelligent developments. There are, however, challenges for developing solutions with small data sets due to practical and cost-effective deployment…

Machine Learning · Computer Science 2024-05-17 Luca Gherardini , Varun Ravi Varma , Karol Capala , Roger Woods , Jose Sousa

In today's digital landscape, the proliferation of conspiracy theories within the disinformation ecosystem of online platforms represents a growing concern. This paper delves into the complexities of this phenomenon. We conducted a…

Social and Information Networks · Computer Science 2024-05-22 Alessandra Recordare , Guglielmo Cola , Tiziano Fagni , Maurizio Tesconi

In order to deal efficiently with the exponential growth of the Web services landscape in composition life cycle activities, it is necessary to have a clear view of its main features. As for many situations where there is a lot of…

Software Engineering · Computer Science 2013-05-03 Chantal Cherifi , Jean-François Santucci

Large generative language models such as GPT-2 are well-known for their ability to generate text as well as their utility in supervised downstream tasks via fine-tuning. Our work is twofold: firstly we demonstrate via human evaluation that…

Computation and Language · Computer Science 2020-09-01 Dara Bahri , Yi Tay , Che Zheng , Donald Metzler , Cliff Brunk , Andrew Tomkins

As the Distributed Collection Manager's work on building tools to support users maintaining collections of changing web-based resources has progressed, questions about the characteristics of people's collections of web pages have arisen.…

Digital Libraries · Computer Science 2011-01-05 Paul Logasa Bogen , Frank Shipman , Richard Furuta

In this work, we revisit linguistic acceptability in the context of large language models. We introduce CoLAC - Corpus of Linguistic Acceptability in Chinese, the first large-scale acceptability dataset for a non-Indo-European language. It…

Computation and Language · Computer Science 2023-09-29 Hai Hu , Ziyin Zhang , Weifang Huang , Jackie Yan-Ki Lai , Aini Li , Yina Patterson , Jiahui Huang , Peng Zhang , Chien-Jer Charles Lin , Rui Wang

There is an increasing interest in ensuring machine learning (ML) frameworks behave in a socially responsible manner and are deemed trustworthy. Although considerable progress has been made in the field of Trustworthy ML (TwML) in the…

Social and Information Networks · Computer Science 2022-06-22 Noemi Derzsy , Subhabrata Majumdar , Rajat Malik

There is a practically unlimited amount of natural language data available. Still, recent work in text comprehension has focused on datasets which are small relative to current computing possibilities. This article is making a case for the…

Computation and Language · Computer Science 2016-10-05 Ondrej Bajgar , Rudolf Kadlec , Jan Kleindienst

Recently, neural models have been leveraged to significantly improve the performance of information extraction from semi-structured websites. However, a barrier for continued progress is the small number of datasets large enough to train…

Computation and Language · Computer Science 2023-06-16 Aidan San , Yuan Zhuang , Jan Bakus , Colin Lockard , David Ciemiewicz , Sandeep Atluri , Yangfeng Ji , Kevin Small , Heba Elfardy

Many recent news reports have claimed that content generated by large language models (LLMs) is taking over the web. However, these claims are typically not based on a representative sample of the web and the methodology underlying them is…

Networking and Internet Architecture · Computer Science 2026-05-04 Sichang Steven He , Calvin Ardi , Ramesh Govindan , Harsha V. Madhyastha

Dashboards remain ubiquitous artifacts for presenting or reasoning with data across different domains. Yet, there has been little work that provides a quantifiable, systematic, and descriptive overview of dashboard designs at scale. We…

Human-Computer Interaction · Computer Science 2023-10-18 Joanna Purich , Arjun Srinivasan , Michael Correll , Leilani Battle , Vidya Setlur , Anamaria Crisan

The promise of "free and open" multi-terabyte datasets often collides with harsh realities. While these datasets may be technically accessible, practical barriers -- from processing complexity to hidden costs -- create a system that…

Computers and Society · Computer Science 2025-06-17 Marc Bara

The Web today has millions of datasets, and the number of datasets continues to grow at a rapid pace. These datasets are not standalone entities; rather, they are intricately connected through complex relationships. Semantic relationships…

Information Retrieval · Computer Science 2024-08-28 Kate Lin , Tarfah Alrashed , Natasha Noy

This paper contributes a new large-scale dataset for weakly supervised cross-media retrieval, named Twitter100k. Current datasets, such as Wikipedia, NUS Wide and Flickr30k, have two major limitations. First, these datasets are lacking in…

Computer Vision and Pattern Recognition · Computer Science 2017-03-21 Yuting Hu , Liang Zheng , Yi Yang , Yongfeng Huang

Existing machine reading comprehension (MRC) models do not scale effectively to real-world applications like web-level information retrieval and question answering (QA). We argue that this stems from the nature of MRC datasets: most of…

Computation and Language · Computer Science 2020-04-17 Xingdi Yuan , Jie Fu , Marc-Alexandre Cote , Yi Tay , Christopher Pal , Adam Trischler

Processing large complex networks recently attracted considerable interest. Complex graphs are useful in a wide range of applications from technological networks to biological systems like the human brain. Sometimes these networks are…

Data Structures and Algorithms · Computer Science 2019-12-03 Christian Schulz