
A curated list of useful resources that cover Offensive AI.
A curated list of useful resources that cover Offensive AI.
Exploiting the vulnerabilities of AI models.
Adversarial Machine Learning is responsible for assessing their weaknesses and providing countermeasures.
It is organized into four types of attacks: extraction, inversion, poisoning and evasion.

It tries to steal the parameters and hyperparameters of a model by making requests that maximize the extraction of information.

Depending on the knowledge of the adversary's model, white-box and black-box attacks can be performed.
In the simplest white-box case (when the adversary has full knowledge of the model, e.g., a sigmoid function), one can create a system of linear equations that can be easily solved.
In the generic case, where there is insufficient knowledge of the model, the substitute model is used. This model is trained with the requests made to the original model in order to imitate the same functionality as the original one.

Training a substitute model is equivalent (in many cases) to training a model from scratch.
Very computationally intensive.
The adversary has limitations on the number of requests before being detected.
Rounding of output values.
Use of differential privacy.
Use of ensembles.
Use of specific defenses
They are intended to reverse the information flow of a machine-learning model.

They enable an adversary to know the model that was not explicitly intended to be shared.
They allow us to know the training data or information as statistical properties of the model.
Three types are possible:
Membership Inference Attack (MIA): An adversary attempts to determine whether a sample was employed as part of the training.
Property Inference Attack (PIA): An adversary aims to extract statistical properties that were not explicitly encoded as features during the training phase.
Reconstruction: An adversary tries to reconstruct one or more samples from the training set and/or their corresponding labels. Also called inversion.