Dataset
Anthropic HH-RLHF
A dataset of ~170,000 human preference comparisons for helpful and harmless AI assistant responses, collected to support research on reinforcement learning from human feedback. Widely used to train and evaluate aligned conversational models.
Owner: Anthropic
Access: Freely available on Hugging Face under the MIT license with no access restrictions.
Safety use case
Training reward models for RLHF; studying the helpfulness-harmlessness tradeoff in LLM alignment.
Benchmark relevance
Relevant to: HH-RLHF
Source
Anthropic HH-RLHF on Hugging FaceStatus: Verified. Last verified: 2026-06-13.
Want to list this dataset on the marketplace?
Data Canon takes no transaction fees. Contact us to update listing details or add new datasets.
Contact about this dataset