AIPolicyTracker

AI incident ·

UK Ofqual's Algorithm Disproportionately Provided Lower Grades Than Teachers' Assessments

8 news reports Snapshot 7 Sep 2026

In brief

An AI system built and deployed by Uk Office Of Qualifications And Examinations Regulation allegedly harmed Underprivileged Pupils, Students and 5 others.

Risk domain
Discrimination and Toxicity Unfair discrimination and misrepresentation
Occurred
Coverage
8 reportsAug 2020 - Jun 2022

What happened

UK Office of Qualifications and Examinations Regulation (Ofqual)'s grade-standardization algorithm providing predicted grades for A level and GCSE qualifications in the UK, Wales, Northern Ireland, and Scotland was reportedly giving grades lower than teachers' assessments, and disproportionately for state schools.

Laws that address this harm

Policy angle: Classified under Discrimination and Toxicity (Unfair discrimination and misrepresentation) in the MIT AI Risk Repository taxonomy; 5 recorded instruments address this use case.

Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).

News reports (8)

Titles link to the original publisher; report text is not reproduced here.

  1. Controversial exams algorithm to set 97% of GCSE results
    theguardian.com · Donna Ferguson, Michael Savage
  2. Ofqual ignored exams warning a month ago amid ministers' pressure
    theguardian.com · Richard Adams, Jessica Elgot, Heather Stewart
  3. Ofqual chief to face MPs over exams fiasco and botched algorithm grading
    theguardian.com · Heather Stewart, Sally Weale, Kate Proctor
  4. The Algorithmic Imprint
    dl.acm.org · Upol Ehsan, Ranjit Singh, Jacob Metcalf

Who was involved

Alleged harmed party
Underprivileged Pupils, Students, Pupils In State Schools, Gcse Pupils, Epistemic Integrity, Educational Communities, A Level Pupils

Classification (MIT AI Risk Repository taxonomy)

Causal entity
AI
Intent
Intentional
Timing
Post-deployment
Harm level
—
Sectors
—
Countries
—

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 1.1.

  • Discrimination

    "Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other pro...

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Harms of Representation and Other Biases

    "A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Risks from bias and underrepresentation

    "The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in high- stakes domains su...

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Bias

    "General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or other aspects of hu...

    International AI Safety Report 2025 (Bengio2025)

  • Biased Training Data

    "Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demographic differences...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Toxicity and Bias Tendencies

    "Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data."

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Bias

    "The training datasets of LLMs may contain biased information that leads LLMs to generate outputs with social biases"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Broken systems

    "These are the most mentioned cases. They refer to situations where the algorithm or the training data lead to unreliable outputs. These systems frequently assign disproportionate weight to some variables, like race or g...

    Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

Incidents in the same risk subdomain

All incidents in this subdomain

Source record: incident #374 on the AI Incident Database · all 8 reports