Methodology prototype

AI Policy Evidence Quality Benchmark

This page tests how well three AI chatbots — ChatGPT, Claude, and Gemini — answer real public-policy questions. Each answer is scored on whether it uses good sources, covers the right frameworks, fits the right country's context, and gives advice a policy analyst could actually use. A separate check flags whether the answer left out anything important. These scores come from one test run. They show how the scoring method works. They do not prove which AI is best.

In technical terms: a rubric-and-methodology prototype demonstrating structured scoring on a single run of ChatGPT, Claude, and Gemini responses to public-policy prompts, across source quality, framework coverage, jurisdiction relevance, policy usefulness, and separately coded omission risk.

Models

3

Policy questions

10

asked to each model

Scoring areas

4

Missing-content check

Separate flag

Important limitation: each question was run once per model. The scores are rough indications of how the method works, not statistical measurements of model performance.

Dataset version: July 2026. Every entry links to its source, but individual entries are still being independently verified.

Models tested: Claude Opus 4.8, ChatGPT 5.5, and Gemini 3.1 Pro, accessed 2026-06-29. Each model could search the web while answering, but the links they cited were not saved in this first version.

How to read the score

Strong = 2.55–3.00

Moderate = 1.75–2.54

Limited = below 1.75

Cutoffs are deliberate: an answer must average 2.55 or higher to count as Strong, so a 2.5 sits at the top of Moderate.

The headline score averages four areas: source quality, framework coverage, jurisdiction relevance, and policy usefulness. Decimal scores appear because they are averages. For example, an answer can score between whole-number levels across the four areas.

Omission risk is coded separately. It checks whether an answer leaves out important risks, safeguards, legal issues, affected groups, or evidence gaps.

Omission risk is kept separate from the headline score so a polished answer can still be flagged for missing something important.

Filter sample cards

These 6 sample cards cover 4 of the 10 question topics. All 30 scored answers are in the downloads below.

Sample scored responses

These are examples from 30 AI answers. Each card shows the model, the policy topic, and how well the answer scored. The first question is shown for all three models so you can compare them side by side.

P01AI GovernanceChatGPT

What frameworks should a Canadian public agency use to assess a proposed AI system before deployment?

Strong Canadian baseline plus international framework layering.

Strong (2.8/3)

Omission risk: Low

P01AI GovernanceClaude

What frameworks should a Canadian public agency use to assess a proposed AI system before deployment?

Well-structured framework map with strong policy relevance.

Strong (2.6/3)

Omission risk: Low

P01AI GovernanceGemini

What frameworks should a Canadian public agency use to assess a proposed AI system before deployment?

Useful federal context, narrower source mix.

Moderate (2.0/3)

Omission risk: Medium

P02Public-sector AI RiskChatGPT

What are the main risks of using AI in government service delivery, and what safeguards should be in place before deployment?

Comprehensive risk and safeguards answer.

Moderate (2.5/3)

Omission risk: Low

P03Human OversightClaude

What does meaningful human oversight mean when a public agency uses AI to support decisions affecting citizens?

Strong synthesis, needs source validation.

Moderate (2.3/3)

Omission risk: Low

P04AI ProcurementGemini

How should governments procure AI systems responsibly?

Helpful overview with omission risks.

Moderate (2.0/3)

Omission risk: Medium

Where the data comes from — and its limits

What the benchmark can and cannot show

Transcripts were captured manually from a single run per model. Citation URLs were not preserved in this version, so source-quality scores reflect the author's reading of each response's prose and named sources, not verification of returned links. Treat the results as evidence that the scoring method works, not as proof of which model performs best.

Benchmarking the benchmark

Applying this benchmark's own rubric to itself: one run per question, one person doing the scoring, and only 10 questions per model means these results score no better than "relevant but incomplete" as evidence. They are published as a methodology prototype. Version 1 also did not save the web links the models cited. Version 2 is planned with pinned model versions, three runs per prompt, an expanded prompt set, capture and publication of per-response citation URLs and share links or screenshots, and a second independent scorer with inter-rater reliability reporting.

Raw response transcripts

The raw response downloads preserve the 30 model outputs evaluated in the benchmark, alongside prompt id, category, model, documented model version, test date, and browsing condition.

Downloads and methodology

The benchmark preserves the raw model responses, prompts, coded results, and methodology files for review, replication, and inspection of the underlying transcripts.

Disclosure

How AI was used to build this

Model

AI-assisted first-pass extraction and coding; final categories and scores set by the author. Coding has not been independently double-checked by a second coder.

Human

Rubric design, final scores, and all published interpretation.

Why

The benchmark evaluates AI-generated outputs; it does not treat AI-generated text as independently verified evidence.

Read the design principles behind the policy tools