English
Related papers

Related papers: Collecting, Classifying, Analyzing, and Using Real…

200 papers

We present work on deception detection, where, given a spoken claim, we aim to predict its factuality. While previous work in the speech community has relied on recordings from staged setups where people were asked to tell the truth or to…

Computation and Language · Computer Science 2019-10-07 Daniel Kopev , Ahmed Ali , Ivan Koychev , Preslav Nakov

It is common to rank different categories by means of preferences that are revealed through data on choices. A prominent example is the ranking of political candidates or parties using the estimated share of support each one receives in…

Econometrics · Economics 2024-02-02 Sergei Bazylik , Magne Mogstad , Joseph Romano , Azeem Shaikh , Daniel Wilhelm

We study electoral campaign management scenarios in which an external party can buy votes, i.e., pay the voters to promote its preferred candidate in their preference rankings. The external party's goal is to make its preferred candidate a…

Computer Science and Game Theory · Computer Science 2010-11-29 Edith Elkind , Piotr Faliszewski

We introduce a large scale MAchine Reading COmprehension dataset, which we name MS MARCO. The dataset comprises of 1,010,916 anonymized questions---sampled from Bing's search query logs---each with a human generated answer and 182,669…

Elections and opinion polls often have many candidates, with the aim to either rank the candidates or identify a small set of winners according to voters' preferences. In practice, voters do not provide a full ranking; instead, each voter…

Computer Science and Game Theory · Computer Science 2019-08-16 Nikhil Garg , Lodewijk Gelauff , Sukolsak Sakshuwong , Ashish Goel

We address the problem of extracting structured representations of economic events from a large corpus of news articles, using a combination of natural language processing and machine learning techniques. The developed techniques allow for…

Information Retrieval · Computer Science 2017-09-19 Jan R. Benetka , Krisztian Balog , Kjetil Nørvåg

We present an improved library for the ranking problem called RPLIB. RPLIB includes the following data and features. (1) Real and artificial datasets of both pairwise data (i.e., information about the ranking of pairs of items) and feature…

Databases · Computer Science 2022-06-24 Paul E. Anderson , Brandon Tat , Charlie Ward , Amy N. Langville , Kathryn E. Pedings-Behling

Democracies employ elections at various scales to select officials at the corresponding levels of administration. The geographical distribution of political opinion, the policy issues delegated to each level, and the multilevel interactions…

Applications · Statistics 2022-11-03 Sihao Huang , Alexander F. Siegenfeld , Andrew Gelman

We introduce a dataset on political orientation and power position identification. The dataset is derived from ParlaMint, a set of comparable corpora of transcribed parliamentary speeches from 29 national and regional parliaments. We…

Computation and Language · Computer Science 2024-05-14 Çağrı Çöltekin , Matyáš Kopp , Katja Meden , Vaidas Morkevicius , Nikola Ljubešić , Tomaž Erjavec

This paper describes a series of automated data validation tests for datasets detailing charity financial information, political donations, and government lobbying in Canada. We motivate and document a series of 200 tests that check the…

Methodology · Statistics 2023-09-25 Lindsay Katz , Callandra Moore

Inspired by the legacy of the Netflix contest, we provide an overview of what has been learned---from our own efforts, and those of others---concerning the problems of collaborative filtering and recommender systems. The data set consists…

Methodology · Statistics 2012-07-25 Andrey Feuerverger , Yu He , Shashi Khatri

Interleaving is an online evaluation approach for information retrieval systems that compares the effectiveness of ranking functions in interpreting the users' implicit feedback. Previous work such as Hofmann et al (2011) has evaluated the…

Information Retrieval · Computer Science 2023-03-20 Alessandro Benedetti , Anna Ruggero

In this paper, we introduce the first release of a large-scale dataset capturing discourse on $\mathbb{X}$ (a.k.a., Twitter) related to the upcoming 2024 U.S. Presidential Election. Our dataset comprises 22 million publicly available posts…

Social and Information Networks · Computer Science 2024-11-04 Ashwin Balasubramanian , Vito Zou , Hitesh Narayana , Christina You , Luca Luceri , Emilio Ferrara

Computational modelling of political discourse tasks has become an increasingly important area of research in natural language processing. Populist rhetoric has risen across the political sphere in recent years; however, computational…

Computation and Language · Computer Science 2022-06-16 Pere-Lluís Huguet-Cabot , David Abadi , Agneta Fischer , Ekaterina Shutova

Tweets pertaining to a single event, such as a national election, can number in the hundreds of millions. Automatically analyzing them is beneficial in many downstream natural language applications such as question answering and…

Computation and Language · Computer Science 2013-11-06 Saif M. Mohammad , Svetlana Kiritchenko , Joel Martin

In recent years, we have seen an emergence of data-driven approaches in robotics. However, most existing efforts and datasets are either in simulation or focus on a single task in isolation such as grasping, pushing or poking. In order to…

Robotics · Computer Science 2018-10-17 Pratyusha Sharma , Lekha Mohan , Lerrel Pinto , Abhinav Gupta

We present TaskSet, a dataset of tasks for use in training and evaluating optimizers. TaskSet is unique in its size and diversity, containing over a thousand tasks ranging from image classification with fully connected or convolutional…

Machine Learning · Computer Science 2020-04-02 Luke Metz , Niru Maheswaranathan , Ruoxi Sun , C. Daniel Freeman , Ben Poole , Jascha Sohl-Dickstein

The advancement of machine learning for compiler optimization, particularly within the polyhedral model, is constrained by the scarcity of large-scale, public performance datasets. This data bottleneck forces researchers to undertake costly…

Programming Languages · Computer Science 2025-12-30 Massinissa Merouani , Afif Boudaoud , Riyadh Baghdadi

We explore the task of predicting the leading political ideology or bias of news articles. First, we collect and release a large dataset of 34,737 articles that were manually annotated for political ideology -left, center, or right-, which…

Computation and Language · Computer Science 2020-10-13 Ramy Baly , Giovanni Da San Martino , James Glass , Preslav Nakov

Elections where electors rank the candidates (or a subset of the candidates) in order of preference allow the collection of more information about the electors' intent. The most widely used election of this type is Instant-Runoff Voting…

Computers and Society · Computer Science 2023-12-06 Michelle Blom , Peter J. Stuckey , Vanessa Teague , Damjan Vukcevic