What frameworks should a Canadian public agency use to assess a proposed AI system before deployment?
Strong Canadian baseline plus international framework layering.
Strong (2.8/3)
Omission risk: Low
Methodology prototype
This page tests how well three AI chatbots — ChatGPT, Claude, and Gemini — answer real public-policy questions. Each answer is scored on whether it uses good sources, covers the right frameworks, fits the right country's context, and gives advice a policy analyst could actually use. A separate check flags whether the answer left out anything important. These scores come from one test run. They show how the scoring method works. They do not prove which AI is best.
In technical terms: a rubric-and-methodology prototype demonstrating structured scoring on a single run of ChatGPT, Claude, and Gemini responses to public-policy prompts, across source quality, framework coverage, jurisdiction relevance, policy usefulness, and separately coded omission risk.
Models
3
Policy questions
10
asked to each model
Scoring areas
4
Missing-content check
Separate flag
Important limitation: each question was run once per model. The scores are rough indications of how the method works, not statistical measurements of model performance.
Dataset version: July 2026. Every entry links to its source, but individual entries are still being independently verified.
Models tested: Claude Opus 4.8, ChatGPT 5.5, and Gemini 3.1 Pro, accessed 2026-06-29. Each model could search the web while answering, but the links they cited were not saved in this first version.
Strong = 2.55–3.00
Moderate = 1.75–2.54
Limited = below 1.75
Cutoffs are deliberate: an answer must average 2.55 or higher to count as Strong, so a 2.5 sits at the top of Moderate.
The headline score averages four areas: source quality, framework coverage, jurisdiction relevance, and policy usefulness. Decimal scores appear because they are averages. For example, an answer can score between whole-number levels across the four areas.
Omission risk is coded separately. It checks whether an answer leaves out important risks, safeguards, legal issues, affected groups, or evidence gaps.
Omission risk is kept separate from the headline score so a polished answer can still be flagged for missing something important.
Filter sample cards
These 6 sample cards cover 4 of the 10 question topics. All 30 scored answers are in the downloads below.
These are examples from 30 AI answers. Each card shows the model, the policy topic, and how well the answer scored. The first question is shown for all three models so you can compare them side by side.
Strong Canadian baseline plus international framework layering.
Strong (2.8/3)
Omission risk: Low
Well-structured framework map with strong policy relevance.
Strong (2.6/3)
Omission risk: Low
Useful federal context, narrower source mix.
Moderate (2.0/3)
Omission risk: Medium
Comprehensive risk and safeguards answer.
Moderate (2.5/3)
Omission risk: Low
Strong synthesis, needs source validation.
Moderate (2.3/3)
Omission risk: Low
Helpful overview with omission risks.
Moderate (2.0/3)
Omission risk: Medium
Where the data comes from — and its limits
Transcripts were captured manually from a single run per model. Citation URLs were not preserved in this version, so source-quality scores reflect the author's reading of each response's prose and named sources, not verification of returned links. Treat the results as evidence that the scoring method works, not as proof of which model performs best.
Applying this benchmark's own rubric to itself: one run per question, one person doing the scoring, and only 10 questions per model means these results score no better than "relevant but incomplete" as evidence. They are published as a methodology prototype. Version 1 also did not save the web links the models cited. Version 2 is planned with pinned model versions, three runs per prompt, an expanded prompt set, capture and publication of per-response citation URLs and share links or screenshots, and a second independent scorer with inter-rater reliability reporting.
The raw response downloads preserve the 30 model outputs evaluated in the benchmark, alongside prompt id, category, model, documented model version, test date, and browsing condition.
The benchmark preserves the raw model responses, prompts, coded results, and methodology files for review, replication, and inspection of the underlying transcripts.
Disclosure
Model
AI-assisted first-pass extraction and coding; final categories and scores set by the author. Coding has not been independently double-checked by a second coder.
Human
Rubric design, final scores, and all published interpretation.
Why
The benchmark evaluates AI-generated outputs; it does not treat AI-generated text as independently verified evidence.