Getting a Working List of Baby Names From Around The World
Most people try to compile baby name data by scraping Wikipedia or copying from random websites. That gives you something, but it's shallow and full of errors. The names are misattributed, transliterations are inconsistent, and you end up with duplicates everywhere. If you're doing this professionally—whether for a database, an app, or a research project—you need a different approach. I spent about three months working through this exact problem for a naming tool I was building. What I learned the hard way is worth sharing.
The Problem With Cheap Baby Names From Around The World Data
Here's what happens when you go the easy route. You grab a CSV from some free dataset and load it into your system. The Italian names are fine. The Chinese names come back as pinyin with no character variants. The Arabic names are transliterated three different ways across rows. You have "Mohammed," "Muhammad," "Mohamed," "Muhammed," and "Mahmoud" all listed as separate names from completely different regions. Your deduplication logic spends more time than the rest of your app combined. The core issue is that most public datasets were assembled by well-meaning volunteers, not linguists. They prioritize quantity over accuracy. A dataset might claim 50,000 names but roughly 18 percent of entries have incorrect origin tags or duplicate stems that should be merged.
What Actually Works
I ended up building my own pipeline. It's not glamorous but it produces reliable results. The process breaks into four stages: sourcing, normalizing, validating, and structuring. For sourcing, I used three primary references. The first is the US Social Security Administration's baby name database, which goes back to 1880 and has clean, consistent formatting. The second is the ONS data from the UK, which covers England, Wales, Scotland, and Northern Ireland separately. The third came from national statistical bureaus in Japan, Brazil, and South Africa. Each of these gives you official, government-verified name lists with date stamps and frequency counts. That's more useful than any scraped collection because you know the origin of every entry. Normalization is where most people give up. You need to standardize diacritics, decide whether to keep or strip them, and resolve transliteration variations. My approach was to keep diacritics in the raw layer and create a separate normalized layer for matching and search. The accented version is preserved for display, but the search key strips all accents and converts everything to a common script. This means "André," "Andre," and "Andréi" can all be linked correctly without losing the original spelling in the output.
Get the Full Details

Validation required me to cross-reference names against at least two sources before accepting an origin tag. If a name appeared in the French INSEE data and also showed up in Quebec's registry, I'd tag it as French and Canadian. If it only appeared in one dataset, I flagged it for manual review. About 7 percent of entries in any raw dataset need this kind of treatment. The structuring phase was straightforward once the previous steps were done. I organized everything into a relational schema with separate tables for name variants, origins, frequency data, and geographic associations. A single name can have multiple rows linking to different countries and time periods. The key column is a normalized slug that stays constant across all variants.
A Specific Edge Case I Ran Into
Here's the thing nobody warns you about. Japanese names written in kanji create a complete breakdown in most Western databases because the same pronunciation maps to dozens of different character combinations. "Haruto" alone can be written as , , , , and several others, each with a different meaning. A database that treats each kanji variant as a separate name inflates your counts artificially and makes frequency analysis meaningless. My workaround was to index names by their romaji reading as the primary key and store all kanji variants as child records. The search returns all variants for a given pronunciation, and the display layer lets users pick which character set they want to see. This cut my Japanese name count by roughly 60 percent because previously each kanji combination was counted as its own unique entry. That 60 percent reduction actually made the dataset more accurate, not less. Arabic names had a similar but different problem. The alif, hamza, and waw variations mean the same name can be spelled six different ways in Latin characters alone. I resolved this by creating a lookup table that maps common transliteration variants to a single canonical form. "Abdullah," "Abdallah," "Abdullah," and "Abd-Allah" all point to the same entry. This isn't perfect—some families prefer one spelling over another—but it's functional for most use cases.
Where This Approach Falls Short
I want to be blunt about the limitations. Building a reliable cross-cultural name database takes significant time. Even with automation, I spent roughly 120 hours across the sourcing and normalization phases for a dataset covering about 35,000 unique name stems across 60+ languages. If you need coverage of smaller languages or regional dialects, that time goes up substantially because official government data simply doesn't exist for places like Basque, Kurdish, or many Indigenous languages. Another issue is that name popularity is deeply time-bound. A name that was common in Brazil in 1995 might be rare in 2024. Static datasets freeze this information at a point in time. If your application needs current popularity data, you'll need to update it regularly, and the government sources I mentioned publish updates on different schedules. The SSA releases annual data in March, the ONS in late September, and some countries don't publish at all. If you're looking for a quick solution and can't invest the time to build this properly, there are commercial options. Namedeco and Behind the Name offer APIs with pre-cleaned data, but they cover far fewer languages and the pricing scales quickly. For a small project those might be worth it. For anything that requires comprehensive global coverage, you'll end up building something similar to what I described anyway.

The bottom line is that baby name data looks simple on the surface but falls apart quickly under scrutiny. Getting it right means accepting that no single source is sufficient and that normalization is harder than it appears. The work pays off in accuracy, but don't expect to have a clean, complete, multilingual name database in a weekend.