About and methodology
Where the data comes from and how the verdict is calculated.
How the verdict is calculated
Every name is scored against up to 55 countries using its recorded male/female split in that country, weighted by how frequently the name occurs there. Countries with more data pull the overall verdict harder than countries with only a handful of records.
United States figures are blended with Social Security Administration birth statistics, which are far larger than the base dataset for US names.
A name is marked country-dependent when it leans strongly male in at least one well-represented country and strongly female in another — this is the case that trips up most guesswork.
Chinese given names that don't appear in the main dataset are scored character by character using a frequency table of common given-name characters.
Data sources
Name-to-gender data comes from the first-name dictionary by Jörg Michael ("gender.c", nam_dict.txt, version 1.2, 2008-11-30), published in the magazine c't and licensed under the GNU Free Documentation License, version 1.2 or later, with no Invariant Sections and no Cover Texts — our derived data is published under the same licence, full text at /licence —, from US Social Security Administration baby name statistics (public domain), and from the wainshine/Chinese-Names-Corpus on GitHub (Apache License 2.0).
Licence
The derived name dataset is distributed under the GNU Free Documentation License (GFDL). See the full attribution for details.