kar-shingle v1 · offline · no LLM judge

Deterministic Near-Duplicate Verifier

Word-shingled MinHash you can recompute by hand — not an AI opinion about "how similar" two texts are.

Paste two documents. This page splits each into overlapping K-word shingles, hashes them with SHA-256, and estimates their Jaccard similarity using a fixed, published MinHash permutation scheme. The verdict — IDENTICAL / NEAR-DUPLICATE / OVERLAPPING / DISTINCT — and the receipt hash beneath it are fully re-derivable: anyone with the same two texts and the algorithm spec below gets the exact same bytes back. No model reads your text and no server ever sees it — everything runs in this tab.

Self-test — runs on load

Running self-test…

Documents

0 words
0 words

Algorithm spec — how to recompute this yourself

  1. Normalize: lowercase the text, collapse all whitespace runs to a single space, trim the ends.
  2. Shingle: split the normalized text on spaces into words; for shingle size K, take every contiguous run of K words and join them with a single space. Deduplicate (a document's shingle set is a set, not a multiset). If the document has fewer than K words, it has zero shingles.
  3. Base hash: for each unique shingle s, compute SHA-256(s) (UTF-8 bytes), take the first 8 hex characters, and read them as an unsigned 32-bit integer h0(s).
  4. Permutations: for i = 0 … N-1, derive fixed constants once — a_i = first 32 bits of SHA-256("kar-shingle-perm-a-" + i), forced odd (a_i | 1); b_i = first 32 bits of SHA-256("kar-shingle-perm-b-" + i). These sixteen-to-256 constants are the same for every user, every run, forever — they are not random per session.
  5. Per-permutation hash: h_i(s) = (a_i · h0(s) + b_i) mod 2^32, computed with 64-bit-safe (BigInt) arithmetic to avoid overflow.
  6. MinHash signature: signature slot i of a document = min over all its shingles s of h_i(s). A document with zero shingles has no signature ("insufficient data").
  7. Similarity: estimated Jaccard similarity of A and B = (number of slots where sigA[i] === sigB[i]) / N. This is a standard unbiased MinHash estimator; its standard error is roughly sqrt(p(1-p)/N), so more hash functions narrow the estimate at the cost of more computation.
  8. Verdict: if SHA-256(normalize(A)) === SHA-256(normalize(B)) → IDENTICAL (similarity is trivially 1.0). Else if similarity ≥ 0.75 → NEAR-DUPLICATE. Else if similarity ≥ 0.35 → OVERLAPPING. Else → DISTINCT. If either document has zero shingles → INSUFFICIENT-DATA.
  9. Receipt: a canonical (sorted-key, whitespace-free) JSON object containing the algorithm id, K, N, and for each document its shingle count, normalized-text SHA-256, and signature SHA-256 (SHA-256 of the signature's 32-bit words as 8-hex-char blocks, concatenated) plus the similarity match count and verdict. The receipt hash is SHA-256 of that canonical string. Wall-clock time is shown for humans but deliberately excluded from the hash.

This is the entire algorithm — there is nothing else in the page that touches your text. View source to check; there are no network calls anywhere in this document.

Honest limits