English
Related papers

Related papers: Reducing the Incentive to Tank: The Ex Post Gold P…

200 papers

Recently, Press and Dyson have proposed a new class of probabilistic and conditional strategies for the two-player iterated Prisoner's Dilemma, so-called zero-determinant strategies. A player adopting zero-determinant strategies is able to…

Computer Science and Game Theory · Computer Science 2014-02-17 Liming Pan , Dong Hao , Zhihai Rong , Tao Zhou

The lottery ticket hypothesis suggests that sparse, sub-networks of a given neural network, if initialized properly, can be trained to reach comparable or even better performance to that of the original network. Prior works in lottery…

Machine Learning · Computer Science 2021-02-01 Neha Mukund Kalibhat , Yogesh Balaji , Soheil Feizi

In tasks aiming for long-term returns, planning becomes essential. We study generative modeling for planning with datasets repurposed from offline reinforcement learning. Specifically, we identify temporal consistency in the absence of…

Machine Learning · Computer Science 2025-08-19 Deqian Kong , Dehong Xu , Minglu Zhao , Bo Pang , Jianwen Xie , Andrew Lizarraga , Yuhao Huang , Sirui Xie , Ying Nian Wu

Pairwise LLM-as-a-judge evaluation asks the judge to identify the \emph{better} of two candidate answers. We study a one-line modification that asks for the \emph{worse} answer instead and recovers the preference by elimination, a procedure…

Computation and Language · Computer Science 2026-05-13 Mingyang Song , Mao Zheng , Xuan Luo

Teams frequently compete on multiple fronts: political parties contest districts for majority control, contractors field specialized units to win procurement contracts, and squads play match by match for titles. Although the prize accrues…

Theoretical Economics · Economics 2026-05-26 Zhonghong Kuang , Jingfeng Lu , Yiyao Zhu

When users lack specific knowledge of various system parameters, their uncertainty may lead them to make undesirable deviations in their decision making. To alleviate this, an informed system operator may elect to signal information to…

Computer Science and Game Theory · Computer Science 2023-03-31 Bryce L. Ferguson , Philip N. Brown , Jason R. Marden

In many repeated auction settings, participants care not only about how frequently they win but also how their winnings are distributed over time. This problem arises in various practical domains where avoiding congested demand is crucial,…

Computer Science and Game Theory · Computer Science 2025-06-13 Giannis Fikioris , Robert Kleinberg , Yoav Kolumbus , Raunak Kumar , Yishay Mansour , Éva Tardos

Tournament organisers supposedly design rules such that a team cannot be strictly better off by exerting a lower effort. However, the European qualification tournaments for recent FIFA soccer World Cups are known to violate this…

Physics and Society · Physics 2020-05-28 László Csató

As agents operate over long horizons, their memory stores grow continuously, making retrieval critical to accessing relevant information. Many agent queries require reasoning-intensive retrieval, where the connection between query and…

Information Retrieval · Computer Science 2026-03-24 Sreeja Apparaju , Nilesh Gupta

Policy optimization methods are popular reinforcement learning algorithms in practice. Recent works have built theoretical foundation for them by proving $\sqrt{T}$ regret bounds even when the losses are adversarial. Such bounds are tight…

Machine Learning · Computer Science 2023-02-21 Christoph Dann , Chen-Yu Wei , Julian Zimmert

Fantasy football leagues involve strategic player trades to optimize team performance. However, identifying optimal trades is complex due to varying player projections, positional needs, and league-specific scoring. Existing approaches…

Neural and Evolutionary Computing · Computer Science 2025-11-25 Evan Parshall , Junaid Ali , Michael Zimmerman

Inference-time computation offers a powerful axis for scaling the performance of language models. However, naively increasing computation in techniques like Best-of-N sampling can lead to performance degradation due to reward hacking.…

Artificial Intelligence · Computer Science 2025-04-09 Audrey Huang , Adam Block , Qinghua Liu , Nan Jiang , Akshay Krishnamurthy , Dylan J. Foster

Aligning generative recommender systems to user preferences via post-training is critical for closing the gap between next-item prediction and actual recommendation quality. Existing post-training methods are ill-suited for production-scale…

Machine Learning · Computer Science 2026-03-12 Keertana Chidambaram , Sanath Kumar Krishnamurthy , Qiuling Xu , Ko-Jen Hsiao , Moumita Bhattacharya

Strategic decision-making in uncertain and adversarial environments is crucial for the security of modern systems and infrastructures. A salient feature of many optimal decision-making policies is a level of unpredictability, or randomness,…

Computer Science and Game Theory · Computer Science 2024-05-03 Keith Paarporn , Rahul Chandan , Dan Kovenock , Mahnoosh Alizadeh , Jason R. Marden

In learning-to-rank (LTR), optimizing only the relevance (or the expected ranking utility) can cause representational harm to certain categories of items. Moreover, if there is implicit bias in the relevance scores, LTR models may fail to…

Machine Learning · Computer Science 2023-08-28 Sruthi Gorantla , Eshaan Bhansali , Amit Deshpande , Anand Louis

State-of-the-art in network science of teams offers effective recommendation methods to answer questions like who is the best replacement, what is the best team expansion strategy, but lacks intuitive ways to explain why the optimization…

Social and Information Networks · Computer Science 2018-09-25 Qinghai Zhou , Liangyue Li , Nan Cao , Norbou Buchler , Hanghang Tong

This paper considers an opportunistic scheduling problem over a renewal system. A controller observes a random event at the beginning of each renewal frame and then chooses an action in response to the event, which affects the duration of…

Optimization and Control · Mathematics 2019-06-10 Xiaohan Wei , Michael J. Neely

The goal of imitation learning is to mimic expert behavior from demonstrations, without access to an explicit reward signal. A popular class of approach infers the (unknown) reward function via inverse reinforcement learning (IRL) followed…

Machine Learning · Computer Science 2022-04-19 Carl Qi , Pieter Abbeel , Aditya Grover

Eco-driving emerges as a cost-effective and efficient strategy to mitigate greenhouse gas emissions in urban transportation networks. Acknowledging the persuasive influence of incentives in shaping driver behavior, this paper presents the…

Systems and Control · Electrical Eng. & Systems 2024-05-20 M. Umar B. Niazi , Jung-Hoon Cho , Munther A. Dahleh , Roy Dong , Cathy Wu

The training process of ranking models involves two key data selection decisions: a sampling strategy, and a labeling strategy. Modern ranking systems, especially those for performing semantic search, typically use a ``hard negative''…

Information Retrieval · Computer Science 2025-05-28 Andrew Parry , Debasis Ganguly , Sean MacAvaney
‹ Prev 1 4 5 6 7 8 10 Next ›