How 1978 cataloging errors put ghost characters into Unicode

How 1978 cataloging errors put ghost characters into Unicode

In 1978, Japan's Ministry of Economy, Trade and Industry established the encoding that would later become known as JIS X 0208, which still serves as an important reference for all Japanese encodings. Soon after release, people noticed that several of the newly added characters had no obvious source: nobody could tell what they meant or how they should be pronounced. These became known as the ghost characters (幽霊文字).

The ghost characters sat unexplained for nearly two decades until 1997, when an investigation set out to trace their origins. Every character in the JIS standard was supposed to carry a record of its source, but even where that record existed it was rarely more specific than the name of the document it came from. One of the more common sources cited was the 'Overview of National Administrative Districts' (国土行政区画総覧), a comprehensive list of Japanese place names; rather than a compact atlas, its latest edition runs to seven volumes of roughly nine hundred pages each, which makes tracking a single character to its origin, with no page reference to go on, a serious undertaking.

By interviewing the catalogers who had worked on the original standard, the 1997 investigators established that some of the ghost characters were invented by accident during the cataloging process itself. The character 妛 is the clearest example: it was meant to record the sequence '山 over 女' (mountain over woman), which occurs in a real place name. Because the two components could not yet be printed as a single character, 山 and 女 were printed separately, cut out, pasted onto a sheet of paper, and then copied; when the copy was read, the seam where the two pieces of paper met looked like a stroke and was mistakenly added to the character. The correct character for the sequence, 𡚴, was not added to either JIS or Unicode until much later.

One character resisted explanation entirely: 彁, which the investigation found had neither a clear source nor any historical precedent. The most likely explanation is that it originated as a misreading of the character 彊, though no specific incident behind it was ever uncovered. One of the piece's cited references notes a real instance of that confusion: 彁 turning up, by mistake, in a digitized Taisho-era newspaper, because of a faded printing of 彊.

Once the JIS standards were generally adopted, these accidental characters carried straight into Unicode, which today underlies text on effectively every modern computer. Unicode has since accumulated its own, separate set of ghost characters, introduced through CJK unification, the process of merging equivalent Chinese, Japanese and Korean characters into shared code points, independent of the JIS episode. The piece's own assessment is that the mistakes were caught too late to undo: at this rate, it says, the ghost characters will presumably be with humanity forever. It closes with a short list of related material: a Japanese terminology reference cited as the most thorough source on the 1997 investigation, an Asahi Shimbun Digital piece on 彁's newspaper appearance, a Nico Nico Douga wiki that treats each ghost character as the name of a youkai, and Xu Bing's hand-printed book 天书 (A Book from the Sky), composed entirely of invented Chinese characters.

Key facts

  • In 1978, Japan's Ministry of Economy, Trade and Industry established the JIS X 0208 encoding, which included several characters with no discoverable source, meaning, or pronunciation, later called the ghost characters (幽霊文字).
  • A 1997 investigation, based on interviews with the original catalogers, traced some of the ghost characters to cataloging accidents; one commonly cited source document, the 'Overview of National Administrative Districts,' runs to seven volumes of roughly nine hundred pages each with no page references, which made tracing individual characters difficult.
  • The character 妛 was created when 山 and 女, printed separately and pasted together to depict a place name, were then copied, and the seam between the pasted pieces was mistaken for a stroke and added to the character by mistake.
  • Only one character, 彁, was left with neither a clear source nor a historical precedent; the leading theory is that it began as a misreading of 彊, though no specific incident was ever found.
  • The JIS ghost characters carried over into Unicode once the JIS standards were generally adopted; Unicode separately has its own distinct set of ghost characters introduced through CJK unification.

Why it matters

The story is a case study in how a technical standard can ossify around its own mistakes. When Japan's Ministry of Economy, Trade and Industry assembled the JIS X 0208 encoding in 1978, a handful of characters entered the standard through cataloging slips rather than through any real document, and nobody could say what they meant or how to pronounce them. JIS X 0208 still serves as an important reference for Japanese encodings, and its characters carried over into Unicode once the JIS standards were generally adopted, so the invented characters cannot simply be removed without risking any text that already uses them. The piece's own assessment is blunt: the mistakes were caught too late to undo, and by its own estimate they are likely to remain part of computing indefinitely. Unicode has since accumulated a separate set of ghost characters of its own, introduced during CJK unification, so the JIS X 0208 case is not an isolated one.

Who it affects

Anyone who touches Japanese text encoding inherits the ghost characters: JIS X 0208 still serves as an important reference for Japanese encodings, and its ghost characters live on inside Unicode, which nearly every modern system relies on. That includes font makers, who still have to render glyphs like 妛 and 彁 even though neither corresponds to a real, historically attested character, and anyone handling digitized Japanese text, where a ghost character can surface by accident: one of the piece's cited references describes 彁 turning up in a digitized Taisho-era newspaper because of a faded printing of the similar-looking 彊. More broadly, the piece is a pointed example for anyone who builds or maintains a character encoding, a database schema, or any other system meant to be a definitive catalog: even a source as thorough as the 'Overview of National Administrative Districts,' the seven-volume, roughly nine-hundred-page-per-volume list of Japanese place names cited here as a common source for the ghost characters, turned out too large to make tracking down a single character's origin easy without a page reference to go on.

How to use it

There's no product or service to adopt here, but the piece doubles as a quick reference for the two best-documented ghost characters: 妛, produced by a cut-and-paste-then-copy transcription of 'mountain over woman' (山 over 女) for a place name, and 彁, most likely a misreading of 彊, with no specific incident on record. Readers who want to dig further are pointed to a Japanese terminology reference cited as the most thorough online source on the 1997 investigation, an Asahi Shimbun Digital piece on 彁's Taisho-era newspaper mix-up, and, as a lighter tangent, a Nico Nico Douga wiki that treats each ghost character as the name of a youkai, plus Xu Bing's hand-printed book 天书 (A Book from the Sky), made entirely of invented Chinese characters.

How solid is it

The strongest evidence here is second-hand but well-sourced: the piece leans on a 1997 investigation that reportedly interviewed the catalogers who worked on JIS X 0208, which is the kind of primary-source digging that explains why 妛's origin can be reconstructed step by step. The unresolved case, 彁, is presented with appropriate hedging: the piece calls the theory that it began as a misreading of 彊 the 'most likely explanation' rather than a settled fact, and says plainly that no specific incident was uncovered. The claims rest on a short reference list rather than inline citations, including a Japanese terminology wiki described as 'the most thorough online source, with citations from the 1997 investigation,' and an Asahi Shimbun Digital piece. The article itself carries no visible byline or publication date in the captured text, and the Hacker News submitter is not its author. On Hacker News, the piece drew 210 points and 73 comments, a strong showing that reflects reader interest rather than independent fact-checking.

Risks and caveats

The piece does not say how many characters in JIS X 0208 actually lack a clear source, only that 'several' did, and it details the etymology of just two of them by name, 妛 and 彁; the origins of the rest of that 'several' are not individually described. Likewise, no count is given for Unicode's own separate set of ghost characters introduced during CJK unification, so there is no way from this piece alone to gauge how large that second set is. No institution or individual is named as having run the 1997 investigation beyond unnamed 'investigators,' which limits how far the claims can be checked independently. And because the affected characters are, by definition, ones nobody can reliably read, pronounce, or place in real usage, this is a piece of encoding history and typographic trivia rather than something with a practical fix.

“The errors went undiscovered just long enough to be set in stone.”

— the author