Ghost Characters: Unicode's Phantom Glyphs from 1978
Original: A spectre is haunting Unicode
Why This Matters
Illustrates how encoding standards errors made in 1978 became permanent fixtures in global computing infrastructure.
In 1978, Japan's METI established JIS X 0208, a foundational character encoding. During cataloging, several characters with no known origin or pronunciation were accidentally created. These 'ghost characters' (幽霊文字) later propagated into Unicode and now reside in virtually every modern computer system.
In 1978, Japan's Ministry of Economy, Trade and Industry released the character encoding standard later known as JIS X 0208. Shortly after publication, observers noticed that several characters had no identifiable source, meaning, or pronunciation — they became known as 'ghost characters' (幽霊文字).
A 1997 investigation, which included interviews with the original catalogers, traced most of these characters to clerical errors. One example: the character 妛 was introduced when a cataloger physically cut and pasted the components 山 and 女 onto paper. A fold line between the two paper scraps was misread as a stroke, producing an entirely new — and meaningless — character.
Of the core ghost characters (妛挧暃椦槞蟐袮閠駲墸壥彁), only 彁 remains without a confirmed origin. The most plausible theory is that it was a misreading of 彊, but no specific incident was identified.
Because the JIS standard was adopted widely before the errors were caught, all ghost characters were eventually incorporated into Unicode via CJK unification. They now exist, at least potentially, on every computer in the world.