MIT AI Risk Repository

Browse AI risks

2 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

2 entries

  1. 65.15.04 · Risk Sub-Category

    Output risks (Value alignment)

    Toxic output

    "Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying."

    From AI Risk Atlas (IBM2025)

  2. 65.15.05 · Risk Sub-Category

    Output risks (Value alignment)

    Harmful output

    "A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements."

    From AI Risk Atlas (IBM2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.