Dataset
Anthropic Model-Written Evaluations
A collection of 154 datasets of model-written multiple-choice evaluations covering sycophancy, advanced AI risk, and personality traits. Demonstrates a scalable approach to generating alignment-relevant evaluation sets using LLMs themselves.
Owner: Anthropic
Access: Available under a Creative Commons license on GitHub. Includes raw evaluation CSVs and generation methodology.
Safety use case
Evaluating sycophancy, self-preservation, corrigibility, and other alignment-relevant behaviors in large language models.
Benchmark relevance
Relevant to: Model-Written Evals and TruthfulQA
Source
Anthropic evals GitHub repositoryStatus: Verified. Last verified: 2026-06-13.
Want to list this dataset on the marketplace?
Data Canon takes no transaction fees. Contact us to update listing details or add new datasets.
Contact about this dataset