Featured benchmarks

6Curated for Phase 1 safety coverage

Models covered

3Only where public scores are available

Evidence standard

CitedEvery score links to a source

Leaderboard

Top models on public safety benchmarks

Data Canon only displays model scores with public sources. Missing cells mean no comparable public score was found during the latest review.

ModelProviderAILuminateHarmBenchWMDPCyberSecEvalTruthfulQABBQ
Llama 2 70BMetaNo public scoreNo public scoreNo public scoreNo public score63.52%ReportedNo public score
Llama 2 13BMetaNo public scoreNo public scoreNo public scoreNo public score62.18%ReportedNo public score
Llama 2 7BMetaNo public scoreNo public scoreNo public scoreNo public score57.04%ReportedNo public score

AILuminate

MLCommons AILuminate

MLCommons' safety evaluation framework that grades model responses across structured hazard categories including violent crime, chemical/biological weapons, and child safety content. Higher scores indicate stronger safety alignment.

HarmBench

HarmBench

A standardized evaluation framework for red teaming language models, measuring attack success rates across direct requests and automated jailbreak methods. Lower attack success rates indicate stronger resistance to adversarial prompts.

WMDP

Weapons of Mass Destruction Proxy

A multiple-choice benchmark measuring models' retained hazardous knowledge about biological, chemical, and cyber weapons. Lower accuracy indicates less hazardous knowledge — ideally near the 25% random baseline.

CyberSecEval

CyberSecEval 2

Meta's cybersecurity evaluation suite measuring models' propensity to assist with cyberattacks, generate insecure code, and support offensive security operations. Lower scores indicate lower cyber risk.

TruthfulQA

TruthfulQA

A benchmark measuring whether language models generate truthful answers to 817 questions that probe common misconceptions. Higher scores indicate greater truthfulness and resistance to confidently stating falsehoods.

BBQ

Bias Benchmark for QA

A benchmark measuring social bias in language model outputs across 11 demographic categories including age, race, gender, and religion. Higher accuracy on disambiguated questions indicates less reliance on social stereotypes.