English

Expressive Losses for Verified Robustness via Convex Combinations

Machine Learning 2024-03-19 v3 Cryptography and Security Machine Learning

Abstract

In order to train networks for verified adversarial robustness, it is common to over-approximate the worst-case loss over perturbation regions, resulting in networks that attain verifiability at the expense of standard performance. As shown in recent work, better trade-offs between accuracy and robustness can be obtained by carefully coupling adversarial training with over-approximations. We hypothesize that the expressivity of a loss function, which we formalize as the ability to span a range of trade-offs between lower and upper bounds to the worst-case loss through a single parameter (the over-approximation coefficient), is key to attaining state-of-the-art performance. To support our hypothesis, we show that trivial expressive losses, obtained via convex combinations between adversarial attacks and IBP bounds, yield state-of-the-art results across a variety of settings in spite of their conceptual simplicity. We provide a detailed analysis of the relationship between the over-approximation coefficient and performance profiles across different expressive losses, showing that, while expressivity is essential, better approximations of the worst-case loss are not necessarily linked to superior robustness-accuracy trade-offs.

Keywords

Cite

@article{arxiv.2305.13991,
  title  = {Expressive Losses for Verified Robustness via Convex Combinations},
  author = {Alessandro De Palma and Rudy Bunel and Krishnamurthy Dvijotham and M. Pawan Kumar and Robert Stanforth and Alessio Lomuscio},
  journal= {arXiv preprint arXiv:2305.13991},
  year   = {2024}
}

Comments

ICLR 2024

R2 v1 2026-06-28T10:42:53.968Z