MIT AI Risk Repository · Risk Sub-Category · 62.15.09

Fine-tuning related (Degrading safety training due to benign fine-tuning)

Category: Model Development

Description

"When downstream providers of AI systems fine-tune AI models to be more suitable for their needs, the resulting AI model can be more likely to produce undesired or harmful outputs (as compared to the non-fine-tuned model), even if the fine-tuning was done with harmless and commonly used data [154]."

From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Subdomain
7.0
Causal entity
Human

How other frameworks describe this risk

Other entries from Gipiškis2024