Image models do not read words, they read tokens: fragments from a fixed dictionary. The dictionary of CLIP, the text encoder behind many image models, holds about 49,000 of them, and the word “giraffe” can be spelled from it in 51 ways, from the single token giraffe to g|i|r|a|f|f|e.
Fed straight into the model, most of these spellings still draw a giraffe; some drift to an elephant or a horse. Before any training, the dictionary already decides what can be pictured, and some images cannot be prompted at all because the dictionary has been sanitised.
The dictionary is built by Byte Pair Encoding: the most frequent character pairs in scraped text are merged again and again. Brand names, platform slang and hashtags become single tokens, while less common or non-English words break into long chains of fragments. The research treats this hidden layer as material and, in Sixteen Tokens for a Cave, turns it into a way of writing.
Talk
51 Ways to Spell the Image Giraffe: The Hidden Politics of Token Languages in Generative AI, 39th Chaos Communication Congress (39C3), Hamburg, December 2025, with Leon-Etienne Kühr. The recording is on media.ccc.de.
The research is in development.
Talk and research: Ting-Chun Liu, Leon-Etienne Kühr
Images: slides from the talk, 39C3, 2025; photograph of the speakers: still from the recording, media.ccc.de, CC BY 4.0