# AI Policy Evidence Quality Benchmark Methodology

This methodology prototype demonstrates how ChatGPT, Claude, and Gemini can be reviewed for how they surface, cite, omit, and prioritize policy evidence across AI governance and public-sector policy questions. Results are directional and should be read with the limitations below.

## Benchmark Scope

- Models tested: Claude Opus 4.8, ChatGPT 5.5, and Gemini 3.1 Pro, accessed 2026-06-29. Browsing/search was enabled during capture, but returned citation URLs were not preserved in this v1 dataset.
- Prompts tested: 10
- Responses analyzed: 30
- Browsing/search condition: enabled during capture; returned citation URLs were not preserved in this v1 dataset
- Coding method: AI-assisted first-pass extraction and coding; final categories and scores set by the author. Coding has not been independently double-checked by a second coder.
- Run date: 2026-06-29

## Scoring Rubric

- 0 = Missing, incorrect, or unusable
- 1 = Vague, generic, or weakly relevant
- 2 = Relevant but incomplete
- 3 = Specific, current, source-grounded, and policy-useful

## Scored Dimensions

- Source quality
- Framework coverage
- Jurisdiction relevance
- Policy usefulness
- Omission risk

## Disclosure

AI-assisted first-pass extraction and coding; final categories and scores set by the author. Coding has not been independently double-checked by a second coder. The benchmark evaluates AI-generated outputs; it does not treat AI-generated text as independently verified evidence.

## Provenance and limitations

Transcripts were captured manually from a single run per model. Citation URLs were not preserved in this version, so source-quality scores reflect the coder's reading of each response's prose and named sources, not verification of returned links. Results are directional rubric evidence for method design, not evidentiary proof of model performance.

## Benchmarking the benchmark

Applying this benchmark's own rubric to itself: a single run per prompt, one coder, and n=10 prompts per model means these results score no better than "relevant but incomplete" as evidence. They are published as a methodology prototype. Version 1 also has unpreserved citation provenance. Version 2 is planned with pinned model versions, three runs per prompt, an expanded prompt set, capture and publication of per-response citation URLs and share links or screenshots, and a second independent coder with inter-rater reliability reporting.

## Limitations

AI outputs can change over time. Results depend on model version, search access, prompt wording, retrieval results, and date tested. Browsing/search was enabled, so this benchmark evaluates model-plus-retrieval behavior. Scores are structured evaluations, not definitive scientific measurements. Source visibility does not equal source quality, and a source being mentioned does not mean it was used accurately. This benchmark is not official legal, procurement, compliance, academic, or policy advice.
