Data sources and licences
Everything here comes from four public sources. This page lists the exact files, what each one contributes, and the licence it carries.
How each of these is turned into the figures on the site is set out on the methodology page; the about page covers why the site exists at all.
US Social Security Administration — national baby names
names.zip — one file
per year, 1880 to 2025, each line Name,Sex,Count. Names with fewer
than five births in a year and sex are omitted.
Licence: US government work, public domain. Loaded here: 2,181,032 rows.
US Social Security Administration — names by state
namesbystate.zip
— one file per state from 1910, each line State,Sex,Year,Name,Count, with the same
five-birth floor applied per state.
Licence: public domain. Loaded here: 6,696,687 rows across 51 states and DC.
US Social Security Administration — actuarial life table
Period life table (2023)— the number of survivors at each exact age from a starting cohort of 100,000, by sex. This is what turns birth counts into a living-population estimate.
Licence: public domain.
Office for National Statistics — baby names, England and Wales
Boys and girls datasets, annual releases 2019–2025. Counts below three are suppressed. Earlier releases are published in a legacy spreadsheet format this build does not read, which is why the England and Wales layer starts at 2019.
Licence: Open Government Licence v3.0. Contains public sector information licensed under the Open Government Licence v3.0.Loaded here: 95,820 rows.
Wiktionary
Given-name entries and their origin categories, used for the origin section on name pages and for classifying Indian-origin names. Etymology text is quoted, never paraphrased into something new.
Licence: CC BY-SA 4.0. Etymology summarised from Wiktionary, CC BY-SA 4.0.
Wikidata
Given-name items, their language-of-name statements and native labels; and, for the notable-people table on each name page, humans who carry the name as a first name and have an English Wikipedia article. Those are ordered by how many Wikipedia language editions cover them — a measure of reach, not of importance — and nobody appears without a Wikidata identifier behind them.
Licence: CC0.
This site's derived data
The computed metrics — living estimates, median ages, location quotients, trends — are released under CC BY 4.0. Each name page links to a JSON endpoint carrying its own numbers — browse the name index, the year pages or the state pages to find one. Attribution to ninan.org is all that is asked.