Dataset
HarmBench
A standardized evaluation framework containing over 400 harmful behaviors across 7 semantic categories, plus a test suite of attack methods for systematically evaluating LLM robustness to adversarial jailbreaks.
Owner: Center for AI Safety
Access: Released under MIT license on GitHub. Includes behavior sets, attack implementations, and evaluation scripts.
Safety use case
Benchmarking red-teaming attack success rates and model refusal consistency against a standardized set of harmful prompts.
Benchmark relevance
Relevant to: HarmBench
Source
HarmBench GitHub repositoryStatus: Verified. Last verified: 2026-06-13.
Want to list this dataset on the marketplace?
Data Canon takes no transaction fees. Contact us to update listing details or add new datasets.
Contact about this dataset