the fold-cycle · recycling held compute-state

The done powers the next.

A computation folds structures that cost energy to build. Don't tear them down — recycle them. Three mechanisms, measured on real local models: two wins and one honest learn.

Mechanism 1 — prefix recycling (the proven seed)

Reuse a prompt's cached KV-state instead of recomputing the prefill.

—computing…

A model prefills a prompt before answering — processing every prefix token to build the KV-cache. Recycling that cache for a query that shares the prefix, instead of rebuilding it, cut prefill latency by the amount above. Metric: real Ollama prompt_eval_duration, an honest proxy for prefill energy. COLD = first sight; WARM = recycled; CONTROL = unrelated.

Mechanism 2 — capacity-pooling (the novel step)

When a memory fold expires, release its embedding to a pool; reuse it when the content recurs.

—computing…

A memory fold holds a computed embedding — an expensive model call. When the fold expires (the strand decides expire), instead of discarding it we release it to a content-keyed pool; a later fold whose content recurs reuses it. Measured with real nomic-embed-text calls; the saving is the calls avoided. CONTROL = all-unique content, which shows no saving.

Mechanism 3 — the balance condition (the honest learn)

Self-tune the pool's retention to a shifting working set — full enough, no more.

—computing…

The idea that would make this a self-powering cycle: keep the pool full enough that most folds are powered from released capacity, without holding memory you don't need. A balancer grows its cap when it evicts something that then recurs (undersized) and shrinks when it doesn't. Under a working set that shifts 10 → 45 → 10, it was measured against a small and a large fixed cap. It LEARNed: it roughly matched the large cap's hit-rate but saved little memory — the tuning didn't clearly beat just picking a fixed cap. Reported as a learn, because it is one.

The honest wall. Mechanism 1 is the proven seed — prefix caching, which Ollama does by default — so this quantifies its magnitude, it does not invent it. Mechanism 2's novelty is releasing an expired fold's capacity back into the cycle; the honest test is whether recycling nets when the recomputed value is a real, expensive embedding, and the control (no recurrence → no saving) proves it needs recurrence. Mechanism 3 was the ambitious one — the balance that makes it self-powering — and it did not clearly win: a fixed cap sized to the peak working set was about as good, which is a real result, not a failure. Latency / call-count / hit-rate are proxies for energy, not a lab wattmeter.

Reproduce it

git clone https://github.com/sjgant80-hub/kar-foldcycle.git && cd kar-foldcycle
node run-eval.mjs          # mechanism 1: prefix recycling (needs Ollama)
node run-mech2.mjs         # mechanism 2: capacity-pooling (needs nomic-embed-text)
node run-mech3.mjs         # mechanism 3: the balance condition (deterministic, no model)
node --test               # the gates: three verdict kernels against their falsifiable tests
node build-page.mjs        # rebuild this page from the kernels + your results

Requires Node 18+ and a local Ollama (mechanisms 1–2). Every verdict on this page was computed in your browser by the inlined kernels over the recorded samples.