MIT AI Risk Repository · Risk Sub-Category · 62.15.09
Fine-tuning related (Degrading safety training due to benign fine-tuning)
Category: Model Development
Description
"When downstream providers of AI systems fine-tune AI models to be more suitable for their needs, the resulting AI model can be more likely to produce undesired or harmful outputs (as compared to the non-fine-tuned model), even if the fine-tuning was done with harmless and commonly used data [154]."
From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.0
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Post-deployment