English
Related papers

Related papers: Should we trust web-scraped data?

200 papers

Sampling is a fundamental problem in computer science and statistics. However, for a given task and stream, it is often not possible to choose good sampling probabilities in advance. We derive a general framework for adaptively changing the…

Machine Learning · Statistics 2022-06-16 Daniel Ting

A drive by download is a download that occurs without users action or knowledge. It usually triggers an exploit of vulnerability in a browser to downloads an unknown file. The malicious program in the downloaded file installs itself on the…

Cryptography and Security · Computer Science 2020-04-10 Saeed Ibrahim , Nawwaf Al Herami , Ebrahim Al Naqbi , Monther Aldwairi

In recent years, the study of complex networks has received a lot of attention. Real systems have gained importance in scientific publications, despite of an important drawback: the difficulty of retrieving and manage such great quantity of…

Computers and Society · Computer Science 2007-10-29 Massimiliano Zanin

Adaptive collection of data is commonplace in applications throughout science and engineering. From the point of view of statistical inference however, adaptive data collection induces memory and correlation in the samples, and poses…

Methodology · Statistics 2020-05-07 Yash Deshpande , Adel Javanmard , Mohammad Mehrabi

Subject selection plays a critical role in experimental studies, especially ones with human subjects. Anecdotal evidence suggests that many such studies, done at or near university campus settings suffer from selection bias, i.e., the…

Machine Learning · Computer Science 2020-12-21 Tahereh Arabghalizi , Alexandros Labrinidis

Websites are regarded as domains of limitless information which anyone and everyone can access. The new trend of technology put us to change the way we are doing our business. The Internet now is fastly becoming a new place for business and…

Information Retrieval · Computer Science 2021-09-03 Ikechukwu Onyenwe , Ebele Onyedinma , Chidinma Nwafor , Obinna Agbata

In theory, a major advantage to the big data approach in studying online communities is that it should be possible to collect a representative random sample from a broadly defined population. However, in practice, data collection processes…

Social and Information Networks · Computer Science 2021-02-02 Muhammad Umer Gurchani

The effect of bias on hypothesis formation is characterized for an automated data-driven projection pursuit neural network to extract and select features for binary classification of data streams. This intelligent exploratory process…

Machine Learning · Computer Science 2022-01-05 John Patterson , Chris Avery , Tyler Grear , Donald J. Jacobs

Bias in online information has recently become a pressing issue, with search engines, social networks and recommendation services being accused of exhibiting some form of bias. In this vision paper, we make the case for a systematic…

Web images come in hand with valuable contextual information. Although this information has long been mined for various uses such as image annotation, clustering of images, inference of image semantic content, etc., insufficient attention…

Multimedia · Computer Science 2020-05-21 F. Fauzi , H. J. Long , M. Belkhatir

Safely deploying machine learning models to the real world is often a challenging process. Models trained with data obtained from a specific geographic location tend to fail when queried with data obtained elsewhere, agents trained in a…

Machine Learning · Computer Science 2021-11-02 Marco Federici , Ryota Tomioka , Patrick Forré

We study the statistical properties of the sampled scale-free networks, deeply related to the proper identification of various real-world networks. We exploit three methods of sampling and investigate the topological properties such as…

Disordered Systems and Neural Networks · Physics 2009-11-24 Sang Hoon Lee , Pan-Jun Kim , Hawoong Jeong

This article aims at summarizing the existing methods for sampling social networking services and proposing a faster confidence interval for related sampling methods. It also includes comparisons of common network sampling techniques.

Applications · Statistics 2013-02-21 Baiyang Wang

In many situations, sample data is obtained from a noisy or imperfect source. In order to address such corruptions, this paper introduces the concept of a sampling corrector. Such algorithms use structure that the distribution is purported…

Data Structures and Algorithms · Computer Science 2018-04-03 Clément Canonne , Themis Gouleakis , Ronitt Rubinfeld

Information distributed through the Web keeps growing faster day by day, and for this reason, several techniques for extracting Web data have been suggested during last years. Often, extraction tasks are performed through so called…

Artificial Intelligence · Computer Science 2013-06-06 Emilio Ferrara , Robert Baumgartner

In real-world applications, observations are often constrained to a small fraction of a system. Such spatial subsampling can be caused by the inaccessibility or the sheer size of the system, and cannot be overcome by longer sampling.…

Data Analysis, Statistics and Probability · Physics 2017-06-02 Anna Levina , Viola Priesemann

Internet is one of the main sources of information for millions of people. One can find information related to practically all matters on internet. Moreover if we want to retrieve information about some particular topic we may find…

Information Retrieval · Computer Science 2012-10-01 Deepika Sharma , Deepak Garg

The Internet is used by billions of users every day because it offers fast and free communication tools and platforms. Nevertheless, with this significant increase in usage, huge amounts of spam are generated every second, which wastes…

Cryptography and Security · Computer Science 2023-12-05 Omar Husni Odeh , Anas Arram , Murad Njoum

This publication describes the motivation and generation of $Q_{bias}$, a large dataset of Google and Bing search queries, a scraping tool and dataset for biased news articles, as well as language models for the investigation of bias in…

Information Retrieval · Computer Science 2023-11-30 Fabian Haak , Philipp Schaer

Online controlled experiments, also known as A/B testing, are the digital equivalent of randomized controlled trials for estimating the impact of marketing campaigns on website visitors. Stratified sampling is a traditional technique for…

‹ Prev 1 3 4 5 6 7 10 Next ›