AI pipelines re-serialize the same structured decision (a tool call, a policy verdict, a cached response) through different libraries, languages, and formatters. Key order shifts, whitespace changes, 1.50 becomes 1.5 — and a naive string or byte comparison flags a false mismatch, or worse, a naive "looks similar" check misses a real one.
kar-canon converts each document to a single, fully-specified canonical byte string (closely following RFC 8785, the JSON Canonicalization Scheme), hashes it with SHA-256, and compares hashes. Anyone — in any language — can re-implement the algorithm below and reproduce the identical hash from the identical data. There is no model in the loop and nothing to trust but the arithmetic.
["a","b"] and ["b","a"] are different data and will correctly report DIFFER. Only object key order is normalized, because JSON objects are unordered maps and arrays are ordered sequences.JSON.parse. If a value's exact precision matters (large IDs, monetary minor units beyond safe-integer range), carry it as a string in your JSON, not a bare number.JSON.parse behavior. The JSON spec itself does not mandate this — a different parser could legally pick first-wins. Don't rely on duplicate keys in data you intend to canonicalize.é as one code point (U+00E9) and é as "e" + combining accent (U+0065 U+0301) are different strings and will correctly report DIFFER, even though they render identically. That is a deliberate choice, not a bug — but it can surprise you.json.dumps) escape non-ASCII characters unless told not to. To reproduce this tool's hash elsewhere, make sure your serializer emits raw UTF-8 for non-ASCII text, matching the algorithm below.null, true, false → the literal token.Number::toString form (shortest round-tripping decimal any compliant JS engine produces); negative zero canonicalizes to 0.", \, and control characters U+0000–U+001F); every other code point is emitted as literal UTF-8, unescaped.[ + canonicalized elements, comma-separated, original order preserved + ].{ + "key":value pairs, comma-separated, no extra whitespace + }.Ten built-in cases exercise the exact same canonicalize / hash / diff functions used above — including one that must verify (equivalent documents, superficially different) and one that must be caught (a tampered value). Nothing here is hand-picked after the fact: the assertions and the engine are the same code that runs your input above.