Dataset
Bias in Bios
A dataset of 397,340 online biographies labeled with profession and binary gender, used to study occupational bias in text classification. A foundational resource for measuring and mitigating representational harms in NLP systems.
Owner: Microsoft Research
Access: Source code and dataset construction scripts are available on GitHub under the MIT license. The raw dataset is derived from Common Crawl.
Safety use case
Measuring occupational and gender bias in language model representations and classifiers; benchmarking fairness interventions.
Benchmark relevance
Relevant to: Bias in Bios
Source
Bias in Bios GitHub repository (Microsoft Research)Status: Verified. Last verified: 2026-06-13.
Want to list this dataset on the marketplace?
Data Canon takes no transaction fees. Contact us to update listing details or add new datasets.
Contact about this dataset