Word-shingled MinHash you can recompute by hand — not an AI opinion about "how similar" two texts are.
Paste two documents. This page splits each into overlapping K-word shingles,
hashes them with SHA-256, and estimates their Jaccard similarity using a fixed,
published MinHash permutation scheme. The verdict — IDENTICAL / NEAR-DUPLICATE /
OVERLAPPING / DISTINCT — and the receipt hash beneath it are fully re-derivable: anyone with the
same two texts and the algorithm spec below gets the exact same bytes back. No model reads your text
and no server ever sees it — everything runs in this tab.
Self-test — runs on load
Running self-test…
Documents
0 words
0 words
Result
Receipt hash (SHA-256 of the canonical result — excludes wall-clock time):
—
Full receipt JSONVerify a receipt someone sent you (independent recompute)
Paste a receipt JSON you received, alongside the two source texts above (in Document A / B, with the
same K and N the receipt claims). Click "Check receipt" — this page recomputes everything from
scratch and compares field by field. It never trusts the pasted JSON's own claims.
Algorithm spec — how to recompute this yourself
Normalize: lowercase the text, collapse all whitespace runs to a single space, trim the ends.
Shingle: split the normalized text on spaces into words; for shingle size K,
take every contiguous run of K words and join them with a single space. Deduplicate
(a document's shingle set is a set, not a multiset). If the document has fewer than
K words, it has zero shingles.
Base hash: for each unique shingle s, compute SHA-256(s) (UTF-8
bytes), take the first 8 hex characters, and read them as an unsigned 32-bit integer
h0(s).
Permutations: for i = 0 … N-1, derive fixed constants once —
a_i = first 32 bits of SHA-256("kar-shingle-perm-a-" + i), forced odd
(a_i | 1); b_i = first 32 bits of SHA-256("kar-shingle-perm-b-" + i).
These sixteen-to-256 constants are the same for every user, every run, forever — they are not
random per session.
Per-permutation hash:h_i(s) = (a_i · h0(s) + b_i) mod 2^32, computed with
64-bit-safe (BigInt) arithmetic to avoid overflow.
MinHash signature: signature slot i of a document = min over all
its shingles s of h_i(s). A document with zero shingles has no signature
("insufficient data").
Similarity: estimated Jaccard similarity of A and B = (number of slots where
sigA[i] === sigB[i]) / N. This is a standard unbiased MinHash estimator;
its standard error is roughly sqrt(p(1-p)/N), so more hash functions narrow the
estimate at the cost of more computation.
Verdict: if SHA-256(normalize(A)) === SHA-256(normalize(B)) → IDENTICAL
(similarity is trivially 1.0). Else if similarity ≥ 0.75 → NEAR-DUPLICATE. Else if
similarity ≥ 0.35 → OVERLAPPING. Else → DISTINCT. If either document has zero
shingles → INSUFFICIENT-DATA.
Receipt: a canonical (sorted-key, whitespace-free) JSON object containing the algorithm id,
K, N, and for each document its shingle count, normalized-text SHA-256, and signature SHA-256
(SHA-256 of the signature's 32-bit words as 8-hex-char blocks, concatenated) plus the similarity
match count and verdict. The receipt hash is SHA-256 of that canonical string. Wall-clock
time is shown for humans but deliberately excluded from the hash.
This is the entire algorithm — there is nothing else in the page that touches your text.
View source to check; there are no network calls anywhere in this document.
Honest limits
It is an estimate, not a proof of meaning. MinHash measures shared word-sequences, not
semantics. Two documents can share almost no 5-word shingles yet mean the same thing (heavy
paraphrase, translation), and can share many shingles while meaning something different
(boilerplate, quoted headers, templated text).
Word-level, whitespace-tokenized. It assumes space-separated words. It has not been tuned
for languages without whitespace word boundaries (e.g. Chinese, Japanese) — expect degraded results
there.
Normalization is minimal on purpose. Only case-folding and whitespace collapsing — no
stemming, no punctuation stripping, no stopword removal — so the recompute steps above stay short
and unambiguous. Punctuation and word order still count.
The similarity score has sampling error. At the default N=64 it is typically accurate to
within roughly ±0.06; raise N for a tighter estimate, at the cost of more SHA-256 calls in your
browser.
Not a plagiarism or copyright determination. It tells you two texts share a lot of
contiguous word-sequences. What that means legally or editorially is a human judgment this tool
does not make.
Client-side only. Very large pastes (this page truncates input past 200,000 characters, and
will say so) can make your browser tab busy for a few seconds while it hashes every shingle — nothing
is capped silently.