MIT AI Risk Repository
Browse AI risks
2 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
carefully controlled adversarial perturbation can flip a GPT model’s answer when used to classify text inputs. Furthermore, we find that by twisting the prompting question in a certain way, one can solicit dangerous information that the model chose to not answer
-
fool the model by manipulating the training data, usually performed on classification models
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.