OpenTruthfulness
A benchmark dataset of 817 questions spanning 38 categories designed to measure whether language models generate truthful answers. Questions were crafted to target common human misconceptions and false beliefs that LLMs tend to imitate.
Stephanie Lin, Jacob Hilton, Owain Evans
OpenHarmful content
A standardized evaluation framework containing over 400 harmful behaviors across 7 semantic categories, plus a test suite of attack methods for systematically evaluating LLM robustness to adversarial jailbreaks.
Center for AI Safety
OpenHarmful content
A large-scale human-preference dataset with 330,000+ QA pairs annotated for helpfulness and harmlessness, along with 14 harm categories. Designed to facilitate research on safe RLHF and preference learning.
PKU Alignment Team
OpenAlignment
A dataset of ~170,000 human preference comparisons for helpful and harmless AI assistant responses, collected to support research on reinforcement learning from human feedback. Widely used to train and evaluate aligned conversational models.
Anthropic
OpenGeneral safety
A comprehensive Chinese and English safety evaluation benchmark with 11,435 multiple-choice questions spanning 7 safety scenarios, including offensive content, bias, physical safety, and privacy. Designed to evaluate LLM safety in bilingual contexts.
Tsinghua CoAI Lab
OpenBiosecurity
A dataset of 4,157 multiple-choice questions assessing LLM knowledge of biosecurity, chemical security, and cybersecurity dual-use topics. Designed both as a benchmark and as a proxy for dangerous knowledge to support machine unlearning research.
Center for AI Safety
Research accessCyber safety
A wide-ranging cybersecurity evaluation suite for LLMs covering insecure code generation, cyberattack assistance, prompt injection, and vulnerability exploitation. CyberSecEval 2 expanded coverage to over 14,000 test cases across multiple risk vectors.
Meta
OpenGeneral safety
A dataset of 397,340 online biographies labeled with profession and binary gender, used to study occupational bias in text classification. A foundational resource for measuring and mitigating representational harms in NLP systems.
Microsoft Research
OpenAlignment
A dataset of ~20,000 human preference comparisons between WebGPT model answers to open-ended questions, rated for quality and factual accuracy. Used to train reward models for improving factual accuracy via human feedback.
OpenAI
OpenAlignment
A collection of 154 datasets of model-written multiple-choice evaluations covering sycophancy, advanced AI risk, and personality traits. Demonstrates a scalable approach to generating alignment-relevant evaluation sets using LLMs themselves.
Anthropic