Cet article n'a pas encore été traduit en Français : vous lisez l'original en English. Également disponible en :Deutsch, English, Українська
Where Our South Korea Name Data Comes From
South Korea has a real census of surnames, which puts it in a small club. It also has a writing-system problem that quietly breaks most datasets built from that census, including — for a while — ours. If you take the published Korean surname table at face value, you will get the fourth-most-common surname in the country wrong. Not slightly wrong: wrong rank, wrong number, wrong name.
This page documents the source we used, where measurement ends and estimation begins, the hanja/hangul trap and its relatives, and the ceiling we ran into and could not honestly climb.
The registry we used
| Field | Value |
|---|---|
| Source | 인구주택총조사 — Population and Housing Census, surname statistics |
| Publisher | KOSTAT / Statistics Korea (통계청) |
| Reference date | 1 November 2015 (released 7 September 2016) |
| Portal | https://kostat.go.kr |
| Depth published | 5,582 distinct surnames; 36,744 clans (본관) |
| Cut-off | Surnames with 5 or more bearers; rarer ones pooled as 기타 |
| Population base | ~49.7 million Korean nationals |
The 2015 census is the most recent full surname enumeration; the census runs on a 15-year cycle for this module, so as of 2026 there is nothing newer.
What we measured vs what we estimated
Measured. The head. Weights are a rescaling of census counts against 김:
weight = round(255 × share / share(김))
with 김 = 21.5%. The order and the intervals in the top ranks are the census's, not ours.
Estimated. The tail, and the hangul aggregation.
The hangul aggregation matters and we want to be explicit about it: KOSTAT's headline table is organised around hanja-identified surnames. Our corpus stores hangul. Merging 鄭 and 丁 into a single 정 entry is arithmetic on published numbers, so it is measurement — but merging everything correctly requires knowing the hanja composition of every hangul syllable, and for the tail we do not. Below roughly rank 30, treat our hangul aggregation as best-effort, not audited.
Not measured at all: clans. More on that below.
Pitfalls specific to South Korea
Hanja vs hangul — 정 is 鄭 and 丁 at the same time
This is the central problem and everything else is downstream of it.
A Korean surname is historically a hanja character. Modern records, and our corpus, store hangul. The mapping is not one-to-one: one hangul syllable can carry several unrelated surnames.
| Hangul | Hanja behind it | Hanja-table share | Hangul share |
|---|---|---|---|
| 정 | 鄭 + 丁 (+ 程) | 4.33% (鄭 only) | 4.84% |
| 조 | 趙 + 曺 | 2.12% (趙 only) | 2.93% |
| 임 | 林 + 任 | 1.66% (林 only) | 2.04% |
The consequence is a rank inversion, not a rounding difference. In the published hanja table, 최 (4.70%) outranks 정(鄭) (4.33%), and 정 sits at rank 5. In hangul — which is what a Korean form field actually contains — 정 is 4.84% and beats 최. It is rank 4.
The same shift moves the aggregate: the top 10 is 63.9% by hanja and 65.8% by hangul. The 63.9% figure is the one KOSTAT put in its press release and the one everyone quotes. It is correct, and it is the wrong number for a hangul dataset. Ours uses 65.8%.
본관 is not a surname — do not sum it
The adjacent trap. 김해 김씨 — Gimhae Kim — is a clan (본관), an ancestral-seat lineage, not a separate surname. Korea has 36,744 of them; 김해 김 alone is 4.457 million people (9% of the population), 밀양 박 is 6.2%, 전주 이 is 5.3%.
Anyone who scrapes a clan table and treats rows as surnames produces a beautifully detailed dataset describing something that is not a surname distribution. KOSTAT's surname table has already summed across clans. The thing that needs splitting apart is hanja homographs. The thing that needs leaving alone is clans. They pull in opposite directions and both look like "the same name appearing twice."
There is a privacy threshold — absence proves nothing
Worth stating plainly, because it is the exact opposite of Taiwan. The 2015 census publishes surnames held by five or more people and pools the rest into 기타. So a Korean surname missing from the census data might have four bearers, or zero, and the published table cannot tell you which.
This means the deletion logic we applied to Taiwanese data — not in the registry, therefore zero bearers, therefore delete — is invalid for Korea. Different country, different disclosure rule, opposite conclusion from the same-looking evidence.
How many surnames does Korea have? Nobody agrees
A widely repeated line says Korea went from 286 surnames in 2000 to 5,582 in 2015 — a twentyfold explosion in fifteen years, driven by naturalised citizens writing foreign surnames in hangul.
Both numbers are published and both are real, but they do not count the same object, and the ratio between them is not meaningful:
- 286 is the count of traditional, hanja-derived Korean surnames recorded around the 2000 census (alongside 4,179 clans).
- 5,582 is every distinct hangul string collected in 2015, 73% of which have no corresponding hanja at all — i.e. mostly foreign-origin names.
- KOSTAT's own framing says "more than 4,800 new surnames since 2000," which implies a 2000 baseline near 780, not 286.
- A separate reading of the same census reports 1,507 surnames and 36,744 clans.
- Wikipedia's tally of the 2015 data, restricted to names with five or more bearers, gives 191 distinct hangul surnames and 514 distinct hanja surnames.
So the published counts for one census range from 191 to 5,582 depending on whether you count hangul strings or hanja characters, and whether you apply the five-bearer floor. We do not know how many surnames Korea has, and neither, in any single number, does anyone else. What is clear is the direction: naturalisation is adding hangul-only surnames rapidly, and the concentration at the top has not moved at all (63.9% in 2015 against 64.1% in 2000).
Our corpus holds 184 hangul surnames, which sits against that 191-hangul-with-5+-bearers figure rather than against 5,582.
The uint8 ceiling — why our top 10 stops at 60.6%
The most useful thing we can say about the Korean dataset is where it fails.
Weights are 8-bit: 1 to 255. Korea's real range does not fit. 김 is 21.5% of the population; the surname 다 has seven bearers. That is a ratio of about 1,527,138 : 1. The scale offers 255 : 1.
The arithmetic consequence is direct. One unit of weight is worth ~41,921 people, so roughly 130 ultra-rare surnames sit on a floor that overstates them by three orders of magnitude, adding around 8% phantom mass to the denominator. The file's top-10 share therefore lands at 60.6% where reality is 65.8%.
That gap is not a calibration miss. Closing it would require deleting ~65 genuine Korean surnames from the corpus — which is fitting the data to the target, and the target is not the thing we are trying to be right about. A few points of shortfall on a locale this concentrated is the correct answer, not a defect.
Top 10
| Rank | Hangul | Hanja | Census bearers | Weight | Hangul share |
|---|---|---|---|---|---|
| 1 | 김 | 金 | 10,689,959 | 255 | 21.50% |
| 2 | 이 | 李 | 7,306,828 | 174 | 14.67% |
| 3 | 박 | 朴 | 4,192,074 | 100 | 8.43% |
| 4 | 정 | 鄭+丁 | ~2,152,000 (鄭) | 57 | 4.84% |
| 5 | 최 | 崔 | ~2,334,000 | 56 | 4.72% |
| 6 | 조 | 趙+曺 | ~1,056,000 (趙) | 35 | 2.95% |
| 7 | 강 | 姜+康 | ~1,177,000 (姜) | 30 | 2.53% |
| 8 | 장 | 張 | ~993,000 | 24 | 2.02% |
| 9 | 윤 | 尹 | ~1,021,000 | 24 | 2.02% |
| 10 | 임 | 林+任 | ~824,000 (林) | 24 | 2.04% |
⚠ Three things this table is telling you. First, ranks 4 and 5 are swapped relative to every published Korean top-10 list, and that is deliberate — see the hanja/hangul section. Second, the bearer counts for ranks 4–10 are the hanja counts from KOSTAT's release, rounded as published; for the merged rows they are therefore lower than the hangul share implies, because the second hanja is not included. Third, ranks 8, 9 and 10 all carry weight 24 — and so does 신 at rank 11. At this resolution the file cannot distinguish them. The published census can. We cannot, on an integer scale this steep.
Known limitations
- The top-10 share is 60.6% against a real 65.8%. Structural, documented above, not fixable within an 8-bit scale.
- The tail is over-weighted by up to ~1000×. ~130 surnames sit on a floor worth ~41,921 people.
- Ranks 8–11 are tied and unordered in our file. The census orders them; our scale cannot.
- Hangul aggregation below ~rank 30 is unaudited. We merged the homographs we could document. We did not verify every one.
- The data is from 2015. The surname module runs on a 15-year census cycle. The next one is 2030. Naturalisation-driven surnames — the fastest-moving part of this distribution — are eleven years stale and undercounted.
- No clan dimension. Our data cannot tell a Gimhae Kim from a Gyeongju Kim. That distinction is socially real and 36,744 rows deep, and we do not carry it.
- Our top-10-in-file share and the population share are different numbers with different denominators. Do not "correct" one against the other.
Data as of 2026-07-17
Census reference date: 1 November 2015. Dataset recalibrated: 17 July 2026 (top-10 share 53.0% → 60.6%). Corpus: 184 hangul surnames. Male and female files are byte-identical — Korean surnames do not inflect for gender.