Related papers: Rejection Mechanism in 2D Bounded Confidence Provi…
High-confidence errors in large language models are often treated as fragile failures. We study an alternative: some errors may be false fixed points, locally stable, internally coherent, and confidently wrong. This separates robustness…
Machine learning models are vulnerable to adversarial examples: minor perturbations to input samples intended to deliberately cause misclassification. While an obvious security threat, adversarial examples yield as well insights about the…
Confidence is an essential ingredient of success in a wide range of domains ranging from job performance and mental health, to sports, business, and combat. Some authors have suggested that not just confidence but overconfidence-believing…
The emergence of structure in cooperative relation is studied in a game theoretical model. It is proved that specific types of reciprocity norm lead individuals to split into two groups. The condition for the evolutionary stability of the…
In this paper, we study the effects of introducing contrarians in a model of Opinion Dynamics where the agents have internal continuous opinions, but exchange information only about a binary choice that is a function of their continuous…
The concept of a bounded confidence level is incorporated in a nonconservative kinetic exchange model of opinion dynamics model where opinions have continuous values $\in [-1,1]$. The characteristics of the unrestricted model, which has one…
We leverage diffusion models to study the robustness-performance tradeoff of robust classifiers. Our approach introduces a simple, pretrained diffusion method to generate low-norm counterfactual examples (CEs): semantically altered data…
We consider a multi-agent system where agents aim to achieve a consensus despite interactions with malicious agents that communicate misleading information. Physical channels supporting communication in cyberphysical systems offer…
Scepticism towards childhood vaccines and genetically modified food has grown despite scientific evidence of their safety. Beliefs about scientific issues are difficult to change because they are entrenched within many related moral…
We propose a stochastic model of opinion exchange in networks. A finite set of agents is organized in a fixed network structure. There is a binary state of the world and each agent receives a private signal on the state. We model beliefs as…
Counterfactual instances are a powerful tool to obtain valuable insights into automated decision processes, describing the necessary minimal changes in the input space to alter the prediction towards a desired target. Most previous…
The evolutionary effect of recombination depends crucially on the epistatic interactions between linked loci. A paradigmatic case where recombination is known to be strongly disadvantageous is a two-locus fitness landscape dis- playing…
Multi-modal foundation models align images, text, and other modalities in a shared embedding space but remain vulnerable to adversarial illusions [35], where imperceptible perturbations disrupt cross-modal alignment and mislead downstream…
Background: Confirmation bias is the tendency to acquire or evaluate new information in a way that is consistent with one's preexisting beliefs. It is omnipresent in psychology, economics, and even scientific practices. Prior theoretical…
Despite their playful purpose social media changed the way users access information, debate, and form their opinions. Recent studies, indeed, showed that users online tend to promote their favored narratives and thus to form polarized…
We present a toy model of opinion spreading in a society which combines a self-reinforcing mechanism with diffusion. The relative strength of these two mechanisms - called the affectability of the system - is a free parameter of the model.…
Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly understood. We uncover a compact causal mechanism for persuasion-induced factual errors. A…
In this work we study a modified version of the two-dimensional Sznajd sociophysics model. In particular, we consider the effects of agents' reputations in the persuasion rules. In other words, a high-reputation group with a common opinion…
Counterfactual explanations describe how to modify a feature vector in order to flip the outcome of a trained classifier. Obtaining robust counterfactual explanations is essential to provide valid algorithmic recourse and meaningful…
We have developed a continuous model of indirect reciprocity and thereby investigated effects of mutation in assessment rules. Within this continuous framework, the difference between the resident and mutant norms is treated as a small…