Related papers: Bernoulli Runs: Using "Book Cricket" to Evaluate C…
Evaluating the overall ability of players in the National Hockey League (NHL) is a difficult task. Existing methods such as the famous "plus/minus" statistic have many shortcomings. Standard linear regression methods work well when player…
The success of a football team depends on various individual skills and performances of the selected players as well as how cohesively they perform. We propose a two-stage process for selecting optimal playing eleven of a football team from…
Deducing whodunit proves challenging for LLM agents. In this paper, we implement a text-based multi-agent version of the classic board game Clue as a rule-based testbed for evaluating multi-step deductive reasoning, with six agents drawn…
Research community evaluations in information retrieval, such as NIST's Text REtrieval Conference (TREC), build reusable test collections by pooling document rankings submitted by many teams. Naturally, the quality of the resulting test…
A checkers-like model game with a simplified set of rules is studied through extensive simulations of agents with different expertise and strategies. The introduction of complementary strategies, in a quite general way, provides a tool to…
This paper proposes a new way of evaluating the accuracy and validity of probabilistic forecasts that change over time (such as an in-game win probability model, or an election forecast). Under this approach, each model to be evaluated is…
When facing a heavily-favored opponent, an underdog must be willing to assume greater-than-average risk. In statistical language, one would say that an underdog must be willing to adopt a strategy whose outcome has a larger-than-average…
Frequently in sporting competitions it is desirable to compare teams based on records of varying schedule strength. Methods have been developed for sports where the result outcomes are win, draw, or loss. In this paper those ideas are…
Rating systems play an important role in competitive sports and games. They provide a measure of player skill, which incentivizes competitive performances and enables balanced match-ups. In this paper, we present a novel Bayesian rating…
Code documentation is useful, but writing it is time-consuming. Different techniques for generating code summaries have emerged, but comparing them is difficult because human evaluation is expensive and automatic metrics are unreliable. In…
We define a notion of the criticality of a player for simple monotone games based on cooperation with other players, either to form a winning coalition or to break a winning one, with an essential role for all the players involved. We…
In many classification tasks, there is no definitive ground truth, only human judgments that may disagree. We address two challenges that arise in such settings: (1) how to use human raters to score classifiers, and (2) how to use them for…
We prove that the lonely runner conjecture holds for eight runners. Our proof relies on a computer verification and on recent results that allow bounding the size of a minimal counterexample. We note that our approach also applies to the…
We present a regularized logistic regression model for evaluating player contributions in hockey. The traditional metric for this purpose is the plus-minus statistic, which allocates a single unit of credit (for or against) to each player…
In the sport of cricket, the side that wins the toss and has the first choice to bat or bowl can have an unfair or a critical advantage. The issue has been discussed by International Cricket Council committees, as well as several cricket…
When recruiting job candidates, employers rarely observe their underlying skill level directly. Instead, they must administer a series of interviews and/or collate other noisy signals in order to estimate the worker's skill. Traditional…
In this study we illustrate a statistical approach to questioned document examination. Specifically, we consider the construction of three classifiers that predict the writer of a sample document based on categorical data. To evaluate these…
Basketball is often referred to as "a game of runs." We investigate the appropriateness of this claim using data from the full NBA 2016-17 season, comparing actual longest runs of scoring events to what long run theory predicts under the…
There have been more hitting streaks in Major League Baseball than we would expect. All batting lines of MLB hitters from 1957-2006 were randomly permuted 10,000 times and the number of hitting streaks of each length from 2 to 100 was…
This paper presents a solution to the Knights and Spies Problem: In a room there are n people, each labelled with a unique number between 1 and n. A person may either be a knight or a spy. Knights always tell the truth, while spies may…