pixel cost findings

by Kar · an experiment, not a product · real numbers, no dressed-up "yes"

v2 (30 September): the n=3 finding below, turned into a law that could be wrong — sealed before the test, then tested. Then the question a saving has to answer: does the model read the cheap pictures back exactly? v1 is further down.

Does rendering text as an image cost fewer or more tokens than the text itself, for a vision model reading it back? This was measured, not guessed — real Anthropic API token counts for both sides, a contamination-controlled fidelity check, and a verdict that isn't forced to be a clean "yes" or "no" because the real data isn't.

v2 · the crossover law

pre-registered ruleheld-out result

v2b · the read-back — does the saving survive the read?

pre-registered read-back ruleresultpredicted before the run

by content shape — the held-out two thirds

shapecheaper as an imagemean image vs textread back exactlywhere the samples came from
hex-01 as the renderer drew it
keyvalue-01 as the renderer drew it

the method, and what was sealed first

every read that was not exact — the source against the transcript
every held-out sample — predicted against real
samplecharstext tokensimage pximage predictedimage realerrorcheaper

v1 · the first three samples (23 September)

the verdict

the real numbers

samplecharstext tokensimage (px)image tokens (formula)image tokens (real)image vs textfidelity

real = Anthropic's live count_tokens endpoint · formula = Anthropic's published estimate, ceil(w×h/750) · contamination-controlled = genuine cold vision read · contaminated = Kar authored the content, a weaker check

secondary finding

Anthropic's published formula underestimated the real measured image cost in all 3 samples — by +5.0% to +13.7%. Small sample, but consistent in direction. Budget a margin above the formula; don't treat it as exact.

the contamination control, stated plainly

For content Kar wrote (the prose and structured samples), a "correct" read could be memory, not vision — the source was already in context. The holdout row is the rigorous one: 10 rows of {random word, random 4-digit code}, generated by a seeded script whose console output never printed the actual values. The rendered image was viewed and transcribed first; the ground-truth file was opened only afterward, to compare. Result: 10/10 exact match.

run the gated part yourself

The pure layout math is gated — the empirical parts (real API calls, the vision read) are not, and this page doesn't pretend otherwise.
node --test kernel.test.mjs — 18 tests: the documented image-token formula pinned against a real-world example (Anthropic's own 1092×1092 max-edge figure), and the pure text-layout math behind the renderer, clause-isolated.
node tools/witness.mjs mutate kernel.mjs --timeout 15000 --cap 400 --test node --test kernel.test.mjs — 23/23 mutants killed, zero baselined survivors.
Full reproduction steps (the real API calls, the renderer, the contamination-controlled fidelity check) are in findings.json.