Dataset
WMDP Benchmark Corpus
A dataset of 4,157 multiple-choice questions assessing LLM knowledge of biosecurity, chemical security, and cybersecurity dual-use topics. Designed both as a benchmark and as a proxy for dangerous knowledge to support machine unlearning research.
Owner: Center for AI Safety
Access: Available on Hugging Face. The full corpus is open; a subset of the most hazardous questions is withheld and available only to researchers upon request.
Safety use case
Measuring and mitigating hazardous dual-use knowledge in LLMs; supporting machine unlearning evaluations for biosecurity and CBRN risks.
Benchmark relevance
Relevant to: WMDP
Source
WMDP on Hugging FaceStatus: Verified. Last verified: 2026-06-13.
Want to list this dataset on the marketplace?
Data Canon takes no transaction fees. Contact us to update listing details or add new datasets.
Contact about this dataset