中文
相关论文

相关论文: Incomplete Reputation Information and Punishment i…

200 篇论文

In classification of incomplete pattern, the missing values can either play a crucial role in the class determination, or have only little influence (or eventually none) on the classification results according to the context. We propose a…

人工智能 · 计算机科学 2016-02-09 Zhun-Ga Liu , Quan Pan , Jean Dezert , Arnaud Martin

We consider schemes for obtaining truthful reports on a common but hidden signal from large groups of rational, self-interested agents. One example are online feedback mechanisms, where users provide observations about the quality of a…

计算机科学与博弈论 · 计算机科学 2014-01-16 Radu Jurca , Boi Faltings

We present reputation-based mechanisms for building reliable task computing systems over the Internet. The most characteristic examples of such systems are the volunteer computing and the crowdsourcing platforms. In both examples end users…

分布式、并行与集群计算 · 计算机科学 2018-03-20 Evgenia Christoforou , Antonio Fernandez Anta , Chryssis Georgiou , Miguel A. Mosteiro , Angel Sanchez

Our recent minimal model of cooperation (P. Gawronski et al, Physica A 388 (2009) 3581) is modified as to allow for time-dependent altruism. This evolution is based on reputation of other agents, which in turn depends on history. We show…

物理与社会 · 物理学 2012-08-29 A. Jarynowski , P. Gawronski , K. Kulakowski

Contextual bandits are canonical models for sequential decision-making under uncertainty in environments with time-varying components. In this setting, the expected reward of each bandit arm consists of the inner product of an unknown…

机器学习 · 统计学 2022-05-27 Hongju Park , Mohamad Kazem Shirani Faradonbeh

Communication is essential for successful interaction. In human-robot interaction, implicit communication holds the potential to enhance robots' understanding of human needs, emotions, and intentions. This paper introduces a method to…

机器人学 · 计算机科学 2026-03-10 Haoyang Jiang , Elizabeth A. Croft , Michael G. Burke

We study a novel variant of the parameterized bandits problem in which the learner can observe additional auxiliary feedback that is correlated with the observed reward. The auxiliary feedback is readily available in many real-life…

机器学习 · 计算机科学 2023-11-07 Arun Verma , Zhongxiang Dai , Yao Shu , Bryan Kian Hsiang Low

Decision makers who receive many signals are subject to imperfect recall. This is especially important when learning from feeds that aggregate messages from many senders on social media platforms. In this paper, we study a stylized model of…

社会与信息网络 · 计算机科学 2022-06-22 Jad Sassine , M. Amin Rahimian , Dean Eckles

Human cooperation persists among strangers in large, well-mixed populations despite theoretical predictions of difficulties, leaving a fundamental evolutionary puzzle. While upstream (pay-it-forward: helping others because you were helped)…

种群与进化 · 定量生物学 2026-01-14 Tatsuya Sasaki , Satoshi Uchida , Isamu Okada , Hitoshi Yamamoto , Yutaka Nakai

Cooperation is widespread in human societies, but its maintenance at the group level remains puzzling if individuals benefit from not cooperating. Explanations of the maintenance of cooperation generally assume that cooperative and…

社会与信息网络 · 计算机科学 2013-04-16 Todd J Bodnar , Marcel Salathé

Recommendation from implicit feedback is a highly challenging task due to the lack of reliable negative feedback data. Existing methods address this challenge by treating all the un-observed data as negative (dislike) but downweight the…

信息检索 · 计算机科学 2021-08-03 Can Wang , Jiawei Chen , Sheng Zhou , Qihao Shi , Yan Feng , Chun Chen

Cooperation between self-interested individuals is a widespread phenomenon in the natural world, but remains elusive in interactions between artificially intelligent agents. Instead, naive reinforcement learning algorithms typically…

多智能体系统 · 计算机科学 2025-01-16 John L. Zhou , Weizhe Hong , Jonathan C. Kao

In the paradigm of mobile Ad hoc networks (MANET), forwarding packets originating from other nodes requires cooperation among nodes. However, as each node may not want to waste its energy, cooperative behavior can not be guaranteed.…

网络与互联网体系结构 · 计算机科学 2020-12-04 Sara Berri , Vineeth Varma , Samson Lasaulce , Mohammed Said Radjef , Jamal Daafouz

Q-learning provides a standard reinforcement learning framework for studying cooperation by specifying how agents update action values from repeated local interactions outcomes. Although previous work has shown that reputation can promote…

物理与社会 · 物理学 2026-02-03 Chunpeng Du , Zongyang Li , Yali Zhang , Yikang Lu , Attila Szolnoki

We consider online learning problems under a partial observability model capturing situations where the information conveyed to the learner is between full information and bandit feedback. In the simplest variant, we assume that in addition…

机器学习 · 计算机科学 2026-04-28 Tomas Kocak , Gergely Neu , Michal Valko , Remi Munos

In complex real-world tasks such as robotic manipulation and autonomous driving, collecting expert demonstrations is often more straightforward than specifying precise learning objectives and task descriptions. Learning from expert data can…

机器人学 · 计算机科学 2025-05-05 Daulet Baimukashev , Gokhan Alcan , Kevin Sebastian Luck , Ville Kyrki

Reinforcement learning (RL) systems typically optimize scalar reward functions that assume precise and reliable evaluation of outcomes. However, real-world objectives--especially those derived from human preferences--are often uncertain,…

机器学习 · 计算机科学 2026-04-30 Disha Singha

Learning from implicit feedback is a fundamental problem in modern recommender systems, where only positive interactions are observed and explicit negative signals are unavailable. In such settings, negative sampling plays a critical role…

信息检索 · 计算机科学 2026-02-24 Chen Chen , Haobo Lin , Yuanbo Xu

We study cooperative stochastic multi-armed bandits with vector-valued rewards under adversarial corruption and limited verification. In each of $T$ rounds, each of $N$ agents selects an arm, the environment generates a clean reward vector,…

机器学习 · 计算机科学 2026-02-23 Ming Shi