Documents where public-sector AI and algorithmic deployments went wrong, including safeguard failures, governance gaps, known outcomes, and evidence-confidence ratings.
A curated reference set of 8 landmark incidents - depth over volume.
Dataset version: July 2026. Every entry links to its source, but individual entries are still being independently verified.
Showing 8 of 8 incidents.
How to read a case
Confidence — how well-documented the incident is. High: official findings (courts, commissions, regulators). Medium: credible public reporting. Unknown: an aspect of the record could not be confirmed from available sources.
Each card follows the same structure: what happened, which safeguard failed, and the policy lesson.
These summaries describe public findings. They are not new investigations.
Cross-case pattern
What these eight cases have in common
Across all eight records, transparency appears as a recorded governance issue — people or oversight bodies could not fully see how a system produced its outputs — and six of the eight also involved inadequate human oversight before consequences attached. Legal basis emerges as the binding constraint in the formally adjudicated cases — Robodebt, SyRI, and the South Wales Police deployments (each detailed in the cards below) were each found to rest on an insufficient legal footing, and Canada's federal privacy commissioner formally found the RCMP's use of Clearview AI violated the Privacy Act. Wisconsin's COMPAS litigation is the instructive counterexample: the court permitted continued use subject to explainability and anti-overreliance cautions, showing that courts can bound a decision-support tool rather than bar it. Where affected people had weak routes to contest a decision (Robodebt, Ofqual, MiDAS), harm compounded before oversight caught it — which is why appeal and recourse design, not just accuracy, determines whether an automated system is governable.
AI-generated synthesis of the primary governance-issue tags across the eight records below; counts are computed from the dataset.
High confidenceSocial benefits and debt recoveryAustralia
Australia Robodebt income-compliance scheme
An automated welfare debt process used income averaging to raise or pursue debts. The Royal Commission found serious legality, fairness, oversight, and accountability failures.
What happened
The scheme compared welfare recipient income declarations with tax-office income data and used averaged income to infer debts. People were asked to disprove debts in a process later examined by courts, government, and the Royal Commission.
Governance issue
legal authority, data quality, human oversight, appeal / recourse, transparency, public trust, evidence quality
Safeguard failure
Insufficient legal assurance, weak challenge pathways, and inadequate controls around using averaged income as evidence of a debt.
Policy lesson
Automated public debt systems need explicit legal authority, evidence-quality thresholds, human review, clear reasons, and accessible challenge routes before deployment.
Notes
The tracker summarizes public findings and does not assess individual claims or compensation outcomes.
England summer 2020 exam grading standardisation model
A statistical model for awarding exam grades during pandemic school closures triggered widespread concern about fairness, explainability, and appeals. Final grades moved to centre-assessed grades.
What happened
Because exams were cancelled, grades were generated through a standardisation process using teacher estimates and school-level historical performance. The model was rapidly withdrawn for final grades after public backlash and official review.
Michigan's automated unemployment insurance system became a prominent case about automated fraud findings, false accusations, penalties, notices, and due-process concerns.
What happened
The system helped flag or process unemployment fraud determinations at scale. Public reporting and later official activity described many determinations as erroneous or contested.
Governance issue
appeal / recourse, human oversight, data quality, transparency, public trust, evidence quality
Safeguard failure
The record points to inadequate human checking and insufficient procedural protections before severe penalties could attach.
Policy lesson
Fraud automation needs false-positive controls, human review before sanctions, legible notices, and accessible appeals before collection actions begin.
Notes
Rated Medium pending direct links to the Michigan Auditor General reports on MiDAS and the Bauserman v. Unemployment Insurance Agency opinion; institutional homepages are listed as supporting references only.
A Dutch court held that SyRI's legal framework did not satisfy privacy and transparency requirements under human-rights law.
What happened
SyRI linked government datasets to generate risk reports for possible social-security or tax fraud. Civil-society groups challenged the framework, and the court ruled against the state.
The framework lacked enough transparency and proportionality safeguards for intrusive data-driven risk profiling.
Policy lesson
Risk profiling that links sensitive public data must be transparent enough for democratic and legal scrutiny, with proportionality controls before deployment.
Notes
Confirmed citation: ECLI:NL:RBDHA:2020:1878, District Court of The Hague, 5 February 2020.
Canadian privacy regulators found serious privacy problems with Clearview AI, and the federal privacy commissioner made a formal Privacy Act violation finding about RCMP use of the service.
What happened
RCMP members used Clearview AI searches. Privacy regulators concluded Clearview's collection practices were unlawful, and the OPC found RCMP's use violated the Privacy Act.
Governance issue
privacy, vendor accountability, procurement, transparency, human oversight, public trust
Safeguard failure
The agency did not have adequate assurance that the vendor's data collection and service complied with Canadian privacy requirements.
Policy lesson
Agencies cannot outsource privacy risk to vendors; biometric tools need lawful data provenance, procurement due diligence, logs, and explicit use limits.
Notes
This card does not cover every Canadian police agency that tested or used Clearview.
High confidencePublic safety and policingUnited Kingdom
South Wales Police live facial recognition litigation
The Court of Appeal held that South Wales Police live facial-recognition deployments lacked sufficient legal clarity and equality-assessment safeguards.
What happened
A civil-liberties challenge argued that police live facial recognition interfered with privacy and lacked adequate safeguards. The Court of Appeal allowed the appeal on specific grounds.
A defendant challenged use of a proprietary COMPAS risk score at sentencing. The Wisconsin Supreme Court permitted limited use but warned about due-process, explainability, and demographic limitations.
What happened
The court considered whether use of a proprietary risk-assessment report at sentencing violated due process. It held that use was permissible in the case but required cautionary warnings.
The court highlighted limits around proprietary methodology, group-based data, and the risk of overreliance.
Policy lesson
Decision-support scores in justice settings need explainability, contestability, anti-overreliance warnings, and independent validation before use in liberty-impacting decisions.
Notes
This card does not evaluate later versions of COMPAS or all jurisdictions' use.
Medium confidencePublic services and business supportUnited States
New York City MyCity chatbot legal-information controversy
Investigative testing reported that a city chatbot gave incorrect legal or compliance advice for businesses, raising concerns about public-sector generative AI guardrails.
What happened
Reporters tested the chatbot and reported answers that appeared to conflict with law or official requirements. The incident is presented as an investigative finding, not an official adjudication.
Governance issue
transparency, human oversight, post-deployment monitoring, public trust, evidence quality
Safeguard failure
The case points to inadequate answer validation, escalation, and monitoring for a public-facing legal-information context.
Policy lesson
Generative AI used for public guidance needs strict source grounding, answer testing, disclaimers, escalation to official pages, and monitoring for harmful inaccuracies.
Notes
Because the main source is investigative journalism, a formal evidence product should verify current service settings and any city remediation.
A large language model performed case selection and first-pass extraction from public sources into the structured fields, and drafted the case summaries and policy lessons.
What the human did
Set the schema, the evidence-confidence tiers, and the inclusion standard; reviewed structure and edited copy.
Verification status
Every entry links to its source, but individual entries are still being independently verified.
Why
These records should not be relied on for formal use without independent checking.