The changelog says why, not just what. When a criterion changes because it met a real codebase and lost, the entry names the failure that caused it.
Two criteria could not be satisfied by whole classes of repository — not for anything about their code, but for the forge they use and the language they are written in. Criteria text is unchanged, so the fingerprint is stable; results move, so the version is bumped.
⚑ Five declared detectors were unreachable, and one of them gated the badge. The file walk skipped
every dot-named entry except .github. The detectors, meanwhile, explicitly named .gitlab-ci.yml,
.travis.yml, .circleci/config.yml, .cursorrules and .aider.conf.yml — files that could never
appear in the list they were matched against. So a GitLab project running a full pipeline scored no
CI (VER-04), and had an empty CI text for VER-06, which is core — meaning it could not earn the
badge at any level of real verification. A Cursor or Aider user’s committed configuration was
invisible to BND-01. The walk now keeps an explicit DOT_KEEP set of the dot-named entries the
criteria ask about.
⚑ And the two CI checks disagreed with each other. VER-04 asked “is there CI?” by filename;
VER-06 asked “does CI run the tests?” by path. CircleCI’s file is .circleci/config.yml — a generic
name in a specific directory — so it answered no to the first and yes to the second. There is
now one isCIConfig predicate, used by both. Two lists that are meant to mean the same thing are one
bug waiting for the case that separates them.
⚑ Python’s assert was not counted as an assertion. The pattern required a ( or . after the
word, which idiomatic Python never has, while def test_ was counted as a test case. So a genuine
Python suite read as cases-with-no-assertions and failed VER-07, also core — for being written in
Python. The statement form is now matched, anchored to the start of a line so that import assert
from … and const assert = require(…) still do not count: importing an assertion library is not
making an assertion.
19 new tests covering what the assessor recognises — licence spellings, lockfiles per ecosystem,
agent configs per tool, spec-by-directory-or-name, ancestor directories in the path set, links inside
code fences, link decoration, per-language test and assertion counting, the root test.mjs
convention, object-form authors, sorted dependencies, and a malformed manifest not halting the run.
113 tests total.
Three fixes, all found by walking the assessor’s own lines with a mutation gate rather than by reading them. The criteria text is unchanged, so the fingerprint is stable — but results move, which is the same reason v0.7 was bumped.
⚑ Every duplication location the tool reported was wrong. EVO-01 and EVO-03 hashed a sliding
window over the filtered line array and then reported i + 1 as the line number. Since filtering
drops comment and short lines, the number reported was the position in the filtered array, not in the
file — off by however many comments sat above it, and further off the deeper into the file you went.
A user told “this block is repeated 5×, see file.mjs:151” opened line 151 and found unrelated code.
That is worse than no location: it looks actionable, and it cannot be argued with. Both scans now
carry the true line number.
⚑ And EVO-01 reported blocks that existed nowhere. Because filtering removed lines rather than breaking the window, two statements separated by a comment block were hashed as an adjacent pair. The tool reported a duplicated two-line block that appears in no file in the repository. A two-line window is now only a block when the two lines are genuinely adjacent. EVO-03 deliberately keeps closing gaps over blank and comment lines: at eight lines the signature is the code itself, and regenerated code commonly differs only in how it is spaced.
Import lines are excluded from EVO-01, for the same reason short lines and comment lines already were: every file that uses a module repeats them, they cannot be factored out of the files that need them, and counting them made any project with four test files fail the criterion for having four test files. This assessor was failing its own EVO-01 largely on its own import headers.
⚑ --json exited 0 on a FAIL. The README offers --json as the machine-readable path, which is
the one a CI would wire up; a pipeline built on it would have been green forever regardless of the
verdict. The exit code is now the verdict in every output mode. This is precisely the test-theatre
the tool exists to detect, shipped inside the tool.
Honest limitation, unchanged: the spec fingerprint hashes the criteria, not the evidence-gathering. A change like this one moves every verdict without tripping that guard. What now holds it is the test suite — 93 tests, including the boundary cases above — not discipline.
A VER-01 detection fix, found by turning the assessor on its own estate. VER-01 (“the repository
contains tests”) missed a suite named exactly test.mjs at the repo root — the konomi single-file
convention — because the old pattern required a separator before test (foo.test.mjs) or a test/
directory. A dozen real, test-bearing repos were being scored UNOPENED for tests they actually had,
skewing every verdict pessimistic. The criteria text is unchanged; only evidence-gathering was corrected,
so the fingerprint is stable — but results move, so the version is bumped and the fix is locked by a
regression fixture (test-corpus/root-test-demo/) with a repo-root test.mjs.
test.mjs, tests.js, spec.rb, etc., at
any depth including root. Precise — latest.mjs, contest.js, testimonials.mjs are not matched, and
a root test.mjs is classified as a test, never as source (so VER-05 does not double-count it).The git-history / provenance domain (24 → 27 criteria), behind a determinism harness: git is read
only for commits reachable from HEAD, and only counts, subjects, and author strings (for matching) are
read — never a commit hash, timestamp, or raw identity enters the verdict. Verified: this repo
assessed twice yields the same hash. The criteria are N/A when .git is absent or git is unavailable,
so the tool still runs without git.
wip/fix/update…).Your Name, root@localhost, …).Deferred honestly (each for a real reason, not omission):
Self-assessment stays PASS: 10/10 core, 15/15 non-core (100%). SPEC_VERSION → v0.6, re-locked. The
rubric is now 27 criteria across all six domains.
Third expansion batch (20 → 24 criteria), the no-git remainder — all deterministic against the walk.
main/bin at a file they never generated.REV-06 — SPEC-05 read prose and doc-examples as commands (found by self-assessment). Its first run
flagged English prose (“make the …”) and the plan doc’s own example commands (npm run build) as
undefined. Run commands are now extracted only from code regions of README-family docs — not prose,
not design docs that merely describe commands. The fourth self-found fix (REV-03/04/05/06).
Self-assessment stays PASS: 9/9 core, 13/13 non-core (100%). SPEC_VERSION → v0.5, re-locked.
Second expansion batch (17 → 20 criteria), still zero-git and deterministic against the static walk.
require
is excluded from the assertion set.)Self-assessment stays PASS: 8/8 core, 11/11 non-core (100%). SPEC_VERSION → v0.4, re-locked.
First batch of the rubric expansion (13 → 17 criteria), the zero-risk, no-git group. Each is deterministic against a static walk; the corpus and self-assessment stayed green throughout.
existsSync (a case-insensitive filesystem
would otherwise make the verdict OS-dependent).author/description, if present, are not known scaffold
placeholders. Placeholder lists are frozen and closed — an open list lets two assessors diverge.REV-05 — the link checker flagged its own example (found by self-assessment). SPEC-03’s first run
against this repo failed on a [text](target) written as an example inside a code span in the v0.3
plan doc. A link inside a code fence or `span` is an example, not a link; the checker now strips
code before extracting links (standard link-checker behaviour). Same pattern as REV-04 — a finding the
tool made against itself, fixed in the open.
Remaining v0.3 criteria (BND-03/04/05, EVO-02/03, SPEC-04/05, VER-06/07) and the git-history domain
(PRV-03/04/05, ACC-05) are designed in docs/RUBRIC-v0.3-plan.md and land in later batches — the git
ones behind their determinism rules.
The version that submits to its own rubric. v0.1 failed its own standard (1 of 3 core criteria);
this release fixes those failures in public before the assessor is pointed at anyone else.
Made it pass its own rubric (P0). Added the things whose absence it flags in others: a spec
(SPEC.md), tests (tests/), CI (.github/workflows/ci.yml), a README, a licence, and committed
agent instructions (CLAUDE.md). npm run self now returns PASS; CI fails if that ever regresses.
SPEC_VERSION is now enforced (REV-03). In v0.1, patching two criteria silently changed the
verdict hash while the version string stayed put — a verdict that can’t be reproduced against a
stated version is worthless. criteriaFingerprint() now hashes the criteria definitions,
spec-lock.json pins that fingerprint to the version, and scripts/check-spec-version.mjs (in CI)
fails on any silent drift. The fingerprint is also stamped into every verdict.
Carried forward from v0.1, found by running it against a repo built to fail:
REV-04 — the marker detector flagged its own definition (found by self-assessment). Running the
assessor on itself, SPEC-02 reported three abandoned-subgoal markers in assessor.mjs — but they were
the literal strings (TODO/FIXME/XXX/HACK) that define the detector, not real abandoned work. A
marker now counts only in tag context (TODO:, TODO(, TODO!, TODO before whitespace/end), not
when the words merely appear inside a /- or |-delimited list. The badge already passed (SPEC-02 is
non-core), but a false positive in the highest-noise detector is worth removing. The proper fix is
comment-aware parsing (a later revision); this is the cheap, correct tightening. This is the method: a
finding the tool made against itself, fixed in the open.
Portability fix. Paths are now normalised to forward slashes, so criteria that match on /
(test-file detection, CI detection) work on Windows as well as POSIX. A verdict must not depend on
the operating system it was produced on.
.assessorignore. A gitignore-style prefix list, read from the repository root, for fixtures the
assessor should not treat as product code. This repository uses it to exclude test-corpus/ — the
deliberately-broken demo repos it ships as its regression suite — from its own self-assessment.
Initial prototype. Deterministic single-pass assessment, binary + N/A + threshold scoring,
content-addressed verdict, thirteen criteria across six domains, seven behavioural tells. Shipped
while still embarrassing — on purpose. It failed its own rubric, and that failure is what v0.2
addresses first.