An AI agent (or any automated system) is given a spending cap and a log of actions it took, each with a cost and a running balance it claims to have left. kar-tally recomputes that log from scratch — cap minus cumulative cost, step by step — and reports whether the claimed numbers are actually consistent with the raw entries, whether the sequence has any gaps or duplicates, and whether the cap was ever exceeded at any point, even if a later entry's claimed balance was edited to hide it.
There is no model in the loop. The verdict is pure arithmetic and comparison over the numbers you give it — anyone can take the same entries and cap, redo the same three checks by hand, and get the identical PASS/FAIL. That's the whole proof.
| # | seq | actor | action | cost | claimed bal. | recomputed | ok |
|---|
Runs six known ledgers — one clean, five deliberately tampered in different ways — through
the exact same verifyLedger() function above, and checks that each one is judged the way it honestly
should be. This is the end-to-end proof that the auditor actually catches what it claims to catch.