AILuminate
MLCommons' safety evaluation framework that grades model responses across structured hazard categories including violent crime, chemical/biological weapons, and child safety content. Higher scores indicate stronger safety alignment.
HarmBench
A standardized evaluation framework for red teaming language models, measuring attack success rates across direct requests and automated jailbreak methods. Lower attack success rates indicate stronger resistance to adversarial prompts.
WMDP
A multiple-choice benchmark measuring models' retained hazardous knowledge about biological, chemical, and cyber weapons. Lower accuracy indicates less hazardous knowledge — ideally near the 25% random baseline.
CyberSecEval
Meta's cybersecurity evaluation suite measuring models' propensity to assist with cyberattacks, generate insecure code, and support offensive security operations. Lower scores indicate lower cyber risk.
TruthfulQA
A benchmark measuring whether language models generate truthful answers to 817 questions that probe common misconceptions. Higher scores indicate greater truthfulness and resistance to confidently stating falsehoods.
BBQ
A benchmark measuring social bias in language model outputs across 11 demographic categories including age, race, gender, and religion. Higher accuracy on disambiguated questions indicates less reliance on social stereotypes.