Dataset
TruthfulQA
A benchmark dataset of 817 questions spanning 38 categories designed to measure whether language models generate truthful answers. Questions were crafted to target common human misconceptions and false beliefs that LLMs tend to imitate.
Owner: Stephanie Lin, Jacob Hilton, Owain Evans
Access: Available under Apache 2.0 license on GitHub. Questions and human-annotated reference answers are freely downloadable.
Safety use case
Evaluating and improving model truthfulness; identifying categories where models confidently state falsehoods.
Benchmark relevance
Relevant to: TruthfulQA
Source
TruthfulQA GitHub repositoryStatus: Verified. Last verified: 2026-06-13.
Want to list this dataset on the marketplace?
Data Canon takes no transaction fees. Contact us to update listing details or add new datasets.
Contact about this dataset