FallForge Gate — proof-of-play for small language models

A tuned SLM is worth exactly what it can prove. This harness measures a candidate model against a baseline on a named use-case — deterministic scorers, real latencies, a verdict that can and does say LOSES — and seals the numbers into a tamper-evident receipt. Every claim on a receipt is a measurement, never an assertion.

The shipped receipt — re-verified in your browser, right now

This page carries the gate's own kernel. On load it fetches the repo's receipt.json — a real run of two real local models — and re-verifies the hash in front of you.

loading receipt…

Verify any receipt

Score a run yourself (advanced) — eval set + outputs, scored by the same kernel

Paste an eval set and one outputs array per side. The kernel scores, compares, and hands you the verdict — all client-side, nothing leaves this page.

How the gate works

Honest limits (v1)