CLIP, the model that links pictures and text in many image generators, reads language through a dictionary of about 49,000 tokens. Starting from a stock photograph of a sea cave, a genetic algorithm bred sixteen tokens over two hundred generations (population 2048), keeping whatever CLIP scored closer to the picture.
email, defeats, wallart, cave, impacting, pi, peoplesvote, blames, mcclure, 👣, croati, desk, zazzle, numerous, mediter, eleng
Fragments like croati and mediter never finish their words, yet image models that took no part in the search draw a cave from them. The token IDs are passed to each model directly, skipping the step where a written prompt would be cut into tokens. The work reads such sequences as poetry written for a machine: we are not the addressee, we only overhear.
How close does a description come to the photograph?
CLIP similarity between the target photograph and five descriptions:
(lazy) human, “a photo of a cave and the ocean”: 0.272
the sixteen tokens: 0.433
ChatGPT: 0.258
JoyCaption Beta One: 0.271
Gemini: 0.273
Because this score is what the algorithm maximised, the transfer to the other models is the part that counts.
Status
Preliminary work, first shown as “Prompt Reconstruction” in the talk 51 Ways to Spell the Image Giraffe at 39C3, Hamburg, 2025. The work is in development.
Artists and developers: Ting-Chun Liu, Leon-Etienne Kühr
Target photograph: Stefan Kunze / Unsplash
Four-image grid, row by row: FLUX.1-dev, Stable Diffusion 3.5 medium, Stable Diffusion XL, Stable Diffusion 1.5