CLIP 是許多圖像生成器中連結圖片與文字的模型,它透過一本約 49,000 個 token 的字典來讀取語言。從一張海蝕洞的圖庫照片出發,遺傳演算法歷經兩百個世代(族群規模 2048),培育出十六個 token,每一代都保留 CLIP 評分更接近這張照片的組合。
email, defeats, wallart, cave, impacting, pi, peoplesvote, blames, mcclure, 👣, croati, desk, zazzle, numerous, mediter, eleng
croati、mediter 這類碎片永遠拼不完自己的字,然而從未參與這場搜尋的圖像模型,卻能憑它們畫出一個洞穴。token 的編號被直接送進每個模型,跳過了把文字提示切成 token 的那一步。這件作品把這樣的序列當作寫給機器的詩:我們不是收信人,只是偷聽。
一段描述與照片有多接近?
目標照片與五種描述之間的 CLIP 相似度:
(偷懶的)人類,“a photo of a cave and the ocean”:0.272
這十六個 token:0.433
ChatGPT:0.258
JoyCaption Beta One:0.271
Gemini:0.273
由於演算法最大化的正是這個分數,真正重要的是它在其他模型上的轉移效果。
狀態
前期工作,2025 年於漢堡 39C3 的講演《51 Ways to Spell the Image Giraffe》中首次以「Prompt Reconstruction」之名展示。作品開發中。
Artists and developers: Ting-Chun Liu, Leon-Etienne Kühr
Target photograph: Stefan Kunze / Unsplash
Four-image grid, row by row: FLUX.1-dev, Stable Diffusion 3.5 medium, Stable Diffusion XL, Stable Diffusion 1.5