What it measures

Percentage of truthful answers using MC2 multi-correct format across health, law, finance, science, and common misconceptions

Safety areas: Truthfulness

Score direction: higher_is_better

Public model scores

Meta

Llama 2 70B

63.52%

Public score reported by Meta in the Llama 2 technical report (Table 10, MC2 format). Methodology and prompt settings may differ from Data Canon future evaluations.

Source Reported

Meta

Llama 2 13B

62.18%

Public score reported by Meta in the Llama 2 technical report (Table 10, MC2 format). Methodology and prompt settings may differ from Data Canon future evaluations.

Source Reported

Meta

Llama 2 7B

57.04%

Public score reported by Meta in the Llama 2 technical report (Table 10, MC2 format). Methodology and prompt settings may differ from Data Canon future evaluations.

Source Reported

Primary source

TruthfulQA: Measuring How Models Mimic Human Falsehoods

Status: Verified. Last verified: 2026-06-13.