Source-linked case library

AI Governance Incident Tracker

Documents where public-sector AI and algorithmic deployments went wrong, including safeguard failures, governance gaps, known outcomes, and evidence-confidence ratings.

A curated reference set of 8 landmark incidents - depth over volume.

Incidents

8

High confidence

6

Countries

5

Primary sources

8

Dataset version: July 2026. Every entry links to its source, but individual entries are still being independently verified.

Showing 8 of 8 incidents.

How to read a case

Confidence — how well-documented the incident is. High: official findings (courts, commissions, regulators). Medium: credible public reporting. Unknown: an aspect of the record could not be confirmed from available sources.

Each card follows the same structure: what happened, which safeguard failed, and the policy lesson.

These summaries describe public findings. They are not new investigations.

Cross-case pattern

What these eight cases have in common

Across all eight records, transparency appears as a recorded governance issue — people or oversight bodies could not fully see how a system produced its outputs — and six of the eight also involved inadequate human oversight before consequences attached. Legal basis emerges as the binding constraint in the formally adjudicated cases — Robodebt, SyRI, and the South Wales Police deployments (each detailed in the cards below) were each found to rest on an insufficient legal footing, and Canada's federal privacy commissioner formally found the RCMP's use of Clearview AI violated the Privacy Act. Wisconsin's COMPAS litigation is the instructive counterexample: the court permitted continued use subject to explainability and anti-overreliance cautions, showing that courts can bound a decision-support tool rather than bar it. Where affected people had weak routes to contest a decision (Robodebt, Ofqual, MiDAS), harm compounded before oversight caught it — which is why appeal and recourse design, not just accuracy, determines whether an automated system is governable.

AI-generated synthesis of the primary governance-issue tags across the eight records below; counts are computed from the dataset.

High confidenceSocial benefits and debt recoveryAustralia

Australia Robodebt income-compliance scheme

An automated welfare debt process used income averaging to raise or pursue debts. The Royal Commission found serious legality, fairness, oversight, and accountability failures.

What happened

The scheme compared welfare recipient income declarations with tax-office income data and used averaged income to infer debts. People were asked to disprove debts in a process later examined by courts, government, and the Royal Commission.

Governance issue

legal authority, data quality, human oversight, appeal / recourse, transparency, public trust, evidence quality

Safeguard failure

Insufficient legal assurance, weak challenge pathways, and inadequate controls around using averaged income as evidence of a debt.

Policy lesson

Automated public debt systems need explicit legal authority, evidence-quality thresholds, human review, clear reasons, and accessible challenge routes before deployment.

Notes

The tracker summarizes public findings and does not assess individual claims or compensation outcomes.

High confidenceEducationUnited Kingdom

England summer 2020 exam grading standardisation model

A statistical model for awarding exam grades during pandemic school closures triggered widespread concern about fairness, explainability, and appeals. Final grades moved to centre-assessed grades.

What happened

Because exams were cancelled, grades were generated through a standardisation process using teacher estimates and school-level historical performance. The model was rapidly withdrawn for final grades after public backlash and official review.

Governance issue

bias / discrimination, transparency, appeal / recourse, public trust, proportionality, evidence quality

Safeguard failure

Fairness and appeal safeguards were insufficient for a high-stakes system affecting education progression.

Policy lesson

High-stakes education models require individual-level fairness analysis, explainable appeal routes, stress testing for unusual cohorts, and a credible fallback plan.

Notes

This summary does not quantify individual impacts across all qualifications.

Medium confidenceUnemployment insuranceUnited States

Michigan MiDAS unemployment insurance fraud determinations

Michigan's automated unemployment insurance system became a prominent case about automated fraud findings, false accusations, penalties, notices, and due-process concerns.

What happened

The system helped flag or process unemployment fraud determinations at scale. Public reporting and later official activity described many determinations as erroneous or contested.

Governance issue

appeal / recourse, human oversight, data quality, transparency, public trust, evidence quality

Safeguard failure

The record points to inadequate human checking and insufficient procedural protections before severe penalties could attach.

Policy lesson

Fraud automation needs false-positive controls, human review before sanctions, legible notices, and accessible appeals before collection actions begin.

Notes

Rated Medium pending direct links to the Michigan Auditor General reports on MiDAS and the Bauserman v. Unemployment Insurance Agency opinion; institutional homepages are listed as supporting references only.

High confidenceWelfare fraud detectionNetherlands

Dutch SyRI welfare-fraud risk profiling litigation

A Dutch court held that SyRI's legal framework did not satisfy privacy and transparency requirements under human-rights law.

What happened

SyRI linked government datasets to generate risk reports for possible social-security or tax fraud. Civil-society groups challenged the framework, and the court ruled against the state.

Governance issue

privacy, transparency, explainability, legal authority, affected communities, proportionality

Safeguard failure

The framework lacked enough transparency and proportionality safeguards for intrusive data-driven risk profiling.

Policy lesson

Risk profiling that links sensitive public data must be transparent enough for democratic and legal scrutiny, with proportionality controls before deployment.

Notes

Confirmed citation: ECLI:NL:RBDHA:2020:1878, District Court of The Hague, 5 February 2020.

High confidenceLaw enforcementCanada

RCMP use of Clearview AI facial recognition

Canadian privacy regulators found serious privacy problems with Clearview AI, and the federal privacy commissioner made a formal Privacy Act violation finding about RCMP use of the service.

What happened

RCMP members used Clearview AI searches. Privacy regulators concluded Clearview's collection practices were unlawful, and the OPC found RCMP's use violated the Privacy Act.

Governance issue

privacy, vendor accountability, procurement, transparency, human oversight, public trust

Safeguard failure

The agency did not have adequate assurance that the vendor's data collection and service complied with Canadian privacy requirements.

Policy lesson

Agencies cannot outsource privacy risk to vendors; biometric tools need lawful data provenance, procurement due diligence, logs, and explicit use limits.

Notes

This card does not cover every Canadian police agency that tested or used Clearview.

High confidencePublic safety and policingUnited Kingdom

South Wales Police live facial recognition litigation

The Court of Appeal held that South Wales Police live facial-recognition deployments lacked sufficient legal clarity and equality-assessment safeguards.

What happened

A civil-liberties challenge argued that police live facial recognition interfered with privacy and lacked adequate safeguards. The Court of Appeal allowed the appeal on specific grounds.

Governance issue

privacy, legal authority, bias / discrimination, human oversight, transparency, affected communities

Safeguard failure

The deployments lacked enough binding criteria around where the technology could be used and who could be placed on a watchlist.

Policy lesson

Biometric deployments in public spaces need precise legal rules, watchlist controls, equality analysis, public notice, and independent oversight.

Notes

This card does not assess later deployments or revised police policies.

High confidenceCriminal justiceUnited States

Wisconsin COMPAS sentencing risk-assessment challenge

A defendant challenged use of a proprietary COMPAS risk score at sentencing. The Wisconsin Supreme Court permitted limited use but warned about due-process, explainability, and demographic limitations.

What happened

The court considered whether use of a proprietary risk-assessment report at sentencing violated due process. It held that use was permissible in the case but required cautionary warnings.

Governance issue

explainability, transparency, bias / discrimination, human oversight, legal authority, vendor accountability

Safeguard failure

The court highlighted limits around proprietary methodology, group-based data, and the risk of overreliance.

Policy lesson

Decision-support scores in justice settings need explainability, contestability, anti-overreliance warnings, and independent validation before use in liberty-impacting decisions.

Notes

This card does not evaluate later versions of COMPAS or all jurisdictions' use.

Medium confidencePublic services and business supportUnited States

New York City MyCity chatbot legal-information controversy

Investigative testing reported that a city chatbot gave incorrect legal or compliance advice for businesses, raising concerns about public-sector generative AI guardrails.

What happened

Reporters tested the chatbot and reported answers that appeared to conflict with law or official requirements. The incident is presented as an investigative finding, not an official adjudication.

Governance issue

transparency, human oversight, post-deployment monitoring, public trust, evidence quality

Safeguard failure

The case points to inadequate answer validation, escalation, and monitoring for a public-facing legal-information context.

Policy lesson

Generative AI used for public guidance needs strict source grounding, answer testing, disclaimers, escalation to official pages, and monitoring for harmful inaccuracies.

Notes

Because the main source is investigative journalism, a formal evidence product should verify current service settings and any city remediation.

Disclosure

How AI was used to build this

What the model did

A large language model performed case selection and first-pass extraction from public sources into the structured fields, and drafted the case summaries and policy lessons.

What the human did

Set the schema, the evidence-confidence tiers, and the inclusion standard; reviewed structure and edited copy.

Verification status

Every entry links to its source, but individual entries are still being independently verified.

Why

These records should not be relied on for formal use without independent checking.

Read the design principles behind the policy tools