deterministic · signed · no LLM judge

Prove your agent
obeyed its rules.

AI agents are non-deterministic — which is exactly why you can't put them in production. agent-proof is the deterministic gate + tamper-evident audit trail that proves an agent obeyed its policy and didn't regress, with a receipt that can say FAILS. No model grading a model. Check one in your browser, right now.

Why agents die in the pilot

70%
name non-determinism their #1 blocker

Engineering leaders, 2026 — the top barrier to getting agents into production.

most
agent pilots never reach production

Quality and trust are the wall. "Usually right" isn't "dependable across a whole task."

Aug '26
the EU AI Act clock

High-risk systems need conformity evidence + automatic logging by 2 Aug 2026. Signed audit trails become a line item.

The fix the whole field converged on: a generator that proposes, a fresh checker that reviews, and a deterministic gate that only passes what's actually correct — logged to an audit trail. That's this.

Check an agent run — live, deterministic

A real agent trace (steps: input → tool calls → output) against a policy (deterministic rules). Edit either and re-check — the verdict is computed in your browser by the gated kernel, no model involved. Load the clean run or the rogue one:

The audit ledger — every run, hash-chained

Each checked run lands on a tamper-evident ledger. This is what an auditor (or the EU AI Act) wants: proof of what your agent did and whether it obeyed. Tamper with any entry and the chain breaks visibly — the page recomputes it.

The regression gate — a receipt that can say FAILS

Run your agent across a fixed scenario suite. It's certified only if every scenario obeys the policy — and one rogue makes the whole thing FAIL. That verdict is the artifact you ship to production or an auditor. A gate that can only say yes is theatre.

How it stays honest

  1. No LLM judge. Every check is a deterministic rule — allowed/forbidden/required tools, output shape, forbidden content, numeric invariants, a step ceiling. A model grading a model gives false confidence; the market's cheap default is exactly the thing you can't trust for the high-stakes call.
  2. It proves compliance, not "good". agent-proof proves your agent obeyed its rules and didn't regress — deterministically. It does not judge whether an answer is subtly wise. That boundary is the point: what it certifies, it certifies for real.
  3. Tamper-evident by construction. The audit ledger is a SHA-256 hash chain; the receipt is content-addressed. Change a verdict, a count, a violation — the hash breaks, and anyone can see it.
  4. The gate can refuse. Fewer than 3 scenarios won't certify (not evidence). One violation fails the suite. FAILS is a first-class, verifiable receipt.

Honest limits (v1)

  • Deterministic policy only — shape, tools, invariants, forbidden content, step ceiling. Fuzzy "is this answer good?" is deliberately out (that's the LLM-judge trap).
  • Tamper-evident, not yet signed/anchored — the ledger and receipt hash-verify; Ed25519 issuer signing + external anchoring is v2.
  • You bring the trace — agent-proof checks a run you capture; wiring it into your agent framework (LangChain, custom, etc.) is an adapter, not the kernel.
  • Not audited. Proof of construction, not a security certification.