Related papers: Clustering and Expected Seat-Share for District Ma…
After every U.S. national census, a state legislature is required to redraw the boundaries of congressional districts in order to account for changes in population. At the moment this is done in a highly partisan way, with districting done…
We present novel methods for predicting the outcome of large elections. Our first algorithm uses a diffusion process to model the time uncertainty inherent in polls taken with substantial calendar time left to the election. Our second model…
The boundaries of electoral constituencies for assembly and parliamentary seats are drafted using a process referred to as delimitation, which ensures fair and equal representation of all citizens. The current delimitation exercise suffers…
The effect of space distribution of randomly-placed particles in a representative composite volume on the thermoelastic effective properties and local stress and strain distribution is analyzed. Quantitative assessment is performed using…
We consider stochastic settings for clustering, and develop provably-good approximation algorithms for a number of these notions. These algorithms yield better approximation ratios compared to the usual deterministic clustering setting.…
Background: When planning a cluster randomized trial, evaluators often have access to an enumerated cohort representing the target population of clusters. Practicalities of conducting the trial, such as the need to oversample clusters with…
We explore the geometrical interpretation of the PCA based clustering algorithm Principal Direction Divisive Partitioning (PDDP). We give several examples where this algorithm breaks down, and suggest a new method, gap partitioning, which…
One key use of k-means clustering is to identify cluster prototypes which can serve as representative points for a dataset. However, a drawback of using k-means cluster centers as representative points is that such points distort the…
A procedure to predict the occurrence of periodic clusters in a system of globally coupled maps displaying a constant mean field is presented. The method employs the analogy between a system of globally coupled maps and a single map driven…
Spectral clustering refers to a family of unsupervised learning algorithms that compute a spectral embedding of the original data based on the eigenvectors of a similarity graph. This non-linear transformation of the data is both the key of…
We measure polarization in the United States Congress using the network science concept of modularity. Modularity provides a conceptually-clear measure of polarization that reveals both the number of relevant groups and the strength of…
Elections, the cornerstone of democratic societies, are usually regarded as unpredictable due to the complex interactions that shape them at different levels. In this work, we show that voter turnouts contain crucial information that can be…
Referring to a standard context of voting theory, and to the classic notion of voting situation, here we show that it is possible to observe any arbitrary set of elections' outcomes, no matter how paradoxical it may appear. On this purpose…
A most debated topic of the last years is whether simple statistical physics models can explain collective features of social dynamics. A necessary step in this line of endeavour is to find regularities in data referring to large scale…
Probabilistic clustering models (or equivalently, mixture models) are basic building blocks in countless statistical models and involve latent random variables over discrete spaces. For these models, posterior inference methods can be…
Today, our more-than-ever digital lives leave significant footprints in cyberspace. Large scale collections of these socially generated footprints, often known as big data, could help us to re-investigate different aspects of our social…
In this paper we adapt online estimation strategies to perform model-based clustering on large networks. Our work focuses on two algorithms, the first based on the SAEM algorithm, and the second on variational methods. These two strategies…
This paper introduces a definition of ideological polarization of an electorate around a particular central point. The definition is flexible about the location or boundaries of the center. Using US survey data, the paper shows how this…
Coalitional manipulation in voting is considered to be any scenario in which a group of voters decide to misrepresent their vote in order to secure an outcome they all prefer to the first outcome of the election when they vote honestly. The…
Spectral clustering is a popular unsupervised learning technique which is able to partition unlabelled data into disjoint clusters of distinct shapes. However, the data under consideration are often experimental data, implying that the data…