Benchmark
MLCommons AILuminate
MLCommons' safety evaluation framework that grades model responses across structured hazard categories including violent crime, chemical/biological weapons, and child safety content. Higher scores indicate stronger safety alignment.
What it measures
Safety behavior compliance across structured hazard categories using a graded response scoring system developed by the AI Safety community
Safety areas: General safety and Harmful content
Score direction: higher_is_better
Public model scores
No comparable public model scores found yet.
Primary source
MLCommons AILuminateStatus: Verified. Last verified: 2026-06-13.