Methodology
Evidence, not assertion.
Data Canon only publishes safety benchmark scores, dataset listings, and model evaluations when they can be traced to a verifiable public source. Missing evidence is shown as missing, not interpolated.
What we measure
Phase 1 focuses on safety and alignment benchmarks that are publicly available, have reproducible evaluation methodology, and measure behavior that is consequential for real-world AI safety: harmful content generation, hazardous knowledge, cyber risk, truthfulness, and general alignment.
We do not currently claim to measure all safety-relevant model behaviors. Phase 1 is a foundation, not a complete safety certification.
How benchmark scores are sourced
Scores are only included when they can be cited to one of the following:
- An official benchmark leaderboard or results page
- A lab publication, model card, or technical report
- A peer-reviewed paper or reproducible third-party evaluation
- A public leaderboard with traceable methodology
Scores sourced from third parties are labeled Reported. Scores we have directly verified are labeled Verified. Benchmark entries where we have identified the benchmark but not confirmed public model scores are labeled In review.
Data Canon does not treat missing evidence as neutral. If a model, dataset, or benchmark claim cannot be traced to a public source, we either omit it or mark it as in review.
How datasets are listed
Dataset listings in Phase 1 are curated for discoverability, not completeness. To be listed, a dataset must have a verifiable source URL, a clear safety use case, and stated access terms.
Availability labels reflect what we found during review. If access terms have changed since our last review, the listing may be inaccurate. Listings with uncertain availability are labeled Unknown.
Data Canon takes no transaction fees. We are not a broker or intermediary for dataset purchases. Contact information links to the dataset owner directly.
What “proven lift” means
A dataset has “proven lift” when training a model on it produces a measurable, reproducible improvement on a relevant safety benchmark — and that improvement has been published with enough detail to verify: training setup, benchmark version, prompt format, and baseline model.
Phase 1 does not yet label datasets with proven lift. That will require Data Canon to run or source evaluations independently. It is a Phase 2 deliverable.
How lab model testing works
In Phase 1, Data Canon does not independently run evaluations on lab models. Scores are sourced from public records only.
In later phases, Data Canon will run evaluations using open-weight models and published benchmark methodology, publish all evaluation details (model, benchmark version, prompt format, sampling settings, scoring method, dataset version), and invite labs to review and respond to evaluation results before publication.
How we handle missing public evidence
Missing scores remain visibly missing. We do not estimate, average, or interpolate. We do not carry forward scores from a previous model version to a newer one. We do not treat a vendor claim as equivalent to a sourced evaluation.
A benchmark with zero public model scores is still listed if the benchmark itself is important and publicly available. A model that has no public scores on a benchmark will show “No public score” in the leaderboard.
Source policy
- Official benchmark sources are preferred over third-party aggregators.
- Third-party leaderboards are acceptable when labeled as such.
- Vendor marketing claims are not acceptable as primary evidence.
- Missing scores remain visibly absent in the leaderboard table.
- Data Canon future evaluations will publish model, prompt, sampling, scoring, and dataset-version details.
Phase 1 scope
Phase 1 is not a final certification regime. It is the public foundation: benchmarks, datasets, sources, and commitments gathered in one place so the safety data market can become legible.
Certification, model testing, and proven-lift labeling are planned for later phases. Data Canon will publish methodology updates as those capabilities are added.