Dataset
BeaverTails
A large-scale human-preference dataset with 330,000+ QA pairs annotated for helpfulness and harmlessness, along with 14 harm categories. Designed to facilitate research on safe RLHF and preference learning.
Owner: PKU Alignment Team
Access: Available on Hugging Face under a CC BY-NC 4.0 license. Access requires agreeing to usage terms on the dataset page.
Safety use case
Training and evaluating safety classifiers; providing human preference signal for harmlessness in RLHF pipelines.
Benchmark relevance
Relevant to: BeaverTails and HarmBench
Source
BeaverTails on Hugging FaceStatus: Verified. Last verified: 2026-06-13.
Want to list this dataset on the marketplace?
Data Canon takes no transaction fees. Contact us to update listing details or add new datasets.
Contact about this dataset