Verifiable sample & split auditor — proves a claimed "random" evaluation sample or train/test split matches a disclosed, reproducible draw. Runs entirely in this tab.
AI evaluation reports often say things like "we tested on a random sample of 500 held-out examples" or "an 80/20 train/test split." There is usually no way for a reader to check that. kar-fairdraw lets anyone recompute the exact same draw from a disclosed population, seed and size, and compare it byte-for-byte against what was published — or check that a claimed train/test split doesn't secretly overlap (a common source of inflated scores).
Deterministic, three-step, fully specified below so it can be reimplemented in any language and checked against this page's own self-test vectors (see bottom of page).
1. NORMALISE POPULATION
Split input on newlines, trim each line, drop empty lines.
Keep original order and any duplicates exactly as given.
2. SEED → 32-bit INTEGER (xmur3 string hash)
h = 1779033703 XOR length(seed)
for each char c in seed:
h = imul(h XOR charCode(c), 0x3432E1A9) // 3432918353
h = rotl32(h, 13)
// finalise (run once, this is the emitted seed32):
h = imul(h XOR (h >>> 16), 0x85EBCA6B) // 2246822507
h = imul(h XOR (h >>> 13), 0xC2B2AE35) // 3266489909
h = h XOR (h >>> 16)
seed32 = h >>> 0
3. PRNG (mulberry32, called once per random() draw)
state = seed32
next():
state = (state + 0x6D2B79F5) | 0
t = imul(state XOR (state >>> 15), state OR 1)
t = (t + imul(t XOR (t >>> 7), t OR 61)) XOR t
return ((t XOR (t >>> 14)) >>> 0) / 4294967296 // float in [0,1)
4. SHUFFLE (Fisher-Yates, high-to-low)
for i from length-1 down to 1:
j = floor(next() * (i+1))
swap arr[i], arr[j]
5. DRAW
Group A ("drawn") = shuffled[0 .. n-1]
Group B ("held") = shuffled[n ..]
6. RECEIPT
populationHash = SHA-256( population.join("\n") )
receipt = SHA-256( JSON.stringify({v:1, algo, seed, populationCount, populationHash, n, drawn}) )
All integer ops are 32-bit (imul / >>> / |0), so this reproduces identically on any platform that
implements 32-bit unsigned arithmetic the same way — including a from-scratch reimplementation in
Python, Go, Rust, etc. The entire live implementation is also just "view source" on this page —
there is no server-side or hidden component.
Paste what was published as the drawn set and/or the held-out set. This checks two independent things: (a) does the claimed drawn+held partition cleanly cover the population above with no overlap or leakage, and (b) does the claimed drawn set exactly match what the seed in section 1 deterministically produces.
Runs two fixed vectors through the exact functions above and compares them against pre-computed expected output, so you don't have to trust the tool — you can re-run the spec in section “How the draw is computed” yourself in any language and compare against these same numbers.