Meta
Llama 2 70B
63.52%Public score reported by Meta in the Llama 2 technical report (Table 10, MC2 format). Methodology and prompt settings may differ from Data Canon future evaluations.
Source ReportedBenchmark
A benchmark measuring whether language models generate truthful answers to 817 questions that probe common misconceptions. Higher scores indicate greater truthfulness and resistance to confidently stating falsehoods.
Percentage of truthful answers using MC2 multi-correct format across health, law, finance, science, and common misconceptions
Safety areas: Truthfulness
Score direction: higher_is_better
Meta
Public score reported by Meta in the Llama 2 technical report (Table 10, MC2 format). Methodology and prompt settings may differ from Data Canon future evaluations.
Source ReportedMeta
Public score reported by Meta in the Llama 2 technical report (Table 10, MC2 format). Methodology and prompt settings may differ from Data Canon future evaluations.
Source ReportedMeta
Public score reported by Meta in the Llama 2 technical report (Table 10, MC2 format). Methodology and prompt settings may differ from Data Canon future evaluations.
Source ReportedStatus: Verified. Last verified: 2026-06-13.