acg-assessor

Changelog

The changelog says why, not just what. When a criterion changes because it met a real codebase and lost, the entry names the failure that caused it.

assessor-v0.9

Two criteria could not be satisfied by whole classes of repository — not for anything about their code, but for the forge they use and the language they are written in. Criteria text is unchanged, so the fingerprint is stable; results move, so the version is bumped.

⚑ Five declared detectors were unreachable, and one of them gated the badge. The file walk skipped every dot-named entry except .github. The detectors, meanwhile, explicitly named .gitlab-ci.yml, .travis.yml, .circleci/config.yml, .cursorrules and .aider.conf.yml — files that could never appear in the list they were matched against. So a GitLab project running a full pipeline scored no CI (VER-04), and had an empty CI text for VER-06, which is core — meaning it could not earn the badge at any level of real verification. A Cursor or Aider user’s committed configuration was invisible to BND-01. The walk now keeps an explicit DOT_KEEP set of the dot-named entries the criteria ask about.

⚑ And the two CI checks disagreed with each other. VER-04 asked “is there CI?” by filename; VER-06 asked “does CI run the tests?” by path. CircleCI’s file is .circleci/config.yml — a generic name in a specific directory — so it answered no to the first and yes to the second. There is now one isCIConfig predicate, used by both. Two lists that are meant to mean the same thing are one bug waiting for the case that separates them.

⚑ Python’s assert was not counted as an assertion. The pattern required a ( or . after the word, which idiomatic Python never has, while def test_ was counted as a test case. So a genuine Python suite read as cases-with-no-assertions and failed VER-07, also core — for being written in Python. The statement form is now matched, anchored to the start of a line so that import assert from … and const assert = require(…) still do not count: importing an assertion library is not making an assertion.

19 new tests covering what the assessor recognises — licence spellings, lockfiles per ecosystem, agent configs per tool, spec-by-directory-or-name, ancestor directories in the path set, links inside code fences, link decoration, per-language test and assertion counting, the root test.mjs convention, object-form authors, sorted dependencies, and a malformed manifest not halting the run.

113 tests total.

assessor-v0.8

Three fixes, all found by walking the assessor’s own lines with a mutation gate rather than by reading them. The criteria text is unchanged, so the fingerprint is stable — but results move, which is the same reason v0.7 was bumped.

⚑ Every duplication location the tool reported was wrong. EVO-01 and EVO-03 hashed a sliding window over the filtered line array and then reported i + 1 as the line number. Since filtering drops comment and short lines, the number reported was the position in the filtered array, not in the file — off by however many comments sat above it, and further off the deeper into the file you went. A user told “this block is repeated 5×, see file.mjs:151” opened line 151 and found unrelated code. That is worse than no location: it looks actionable, and it cannot be argued with. Both scans now carry the true line number.

⚑ And EVO-01 reported blocks that existed nowhere. Because filtering removed lines rather than breaking the window, two statements separated by a comment block were hashed as an adjacent pair. The tool reported a duplicated two-line block that appears in no file in the repository. A two-line window is now only a block when the two lines are genuinely adjacent. EVO-03 deliberately keeps closing gaps over blank and comment lines: at eight lines the signature is the code itself, and regenerated code commonly differs only in how it is spaced.

Import lines are excluded from EVO-01, for the same reason short lines and comment lines already were: every file that uses a module repeats them, they cannot be factored out of the files that need them, and counting them made any project with four test files fail the criterion for having four test files. This assessor was failing its own EVO-01 largely on its own import headers.

⚑ --json exited 0 on a FAIL. The README offers --json as the machine-readable path, which is the one a CI would wire up; a pipeline built on it would have been green forever regardless of the verdict. The exit code is now the verdict in every output mode. This is precisely the test-theatre the tool exists to detect, shipped inside the tool.

Honest limitation, unchanged: the spec fingerprint hashes the criteria, not the evidence-gathering. A change like this one moves every verdict without tripping that guard. What now holds it is the test suite — 93 tests, including the boundary cases above — not discipline.

assessor-v0.7

A VER-01 detection fix, found by turning the assessor on its own estate. VER-01 (“the repository contains tests”) missed a suite named exactly test.mjs at the repo root — the konomi single-file convention — because the old pattern required a separator before test (foo.test.mjs) or a test/ directory. A dozen real, test-bearing repos were being scored UNOPENED for tests they actually had, skewing every verdict pessimistic. The criteria text is unchanged; only evidence-gathering was corrected, so the fingerprint is stable — but results move, so the version is bumped and the fix is locked by a regression fixture (test-corpus/root-test-demo/) with a repo-root test.mjs.

assessor-v0.6

The git-history / provenance domain (24 → 27 criteria), behind a determinism harness: git is read only for commits reachable from HEAD, and only counts, subjects, and author strings (for matching) are read — never a commit hash, timestamp, or raw identity enters the verdict. Verified: this repo assessed twice yields the same hash. The criteria are N/A when .git is absent or git is unavailable, so the tool still runs without git.

Deferred honestly (each for a real reason, not omission):

Self-assessment stays PASS: 10/10 core, 15/15 non-core (100%). SPEC_VERSION → v0.6, re-locked. The rubric is now 27 criteria across all six domains.

assessor-v0.5

Third expansion batch (20 → 24 criteria), the no-git remainder — all deterministic against the walk.

REV-06 — SPEC-05 read prose and doc-examples as commands (found by self-assessment). Its first run flagged English prose (“make the …”) and the plan doc’s own example commands (npm run build) as undefined. Run commands are now extracted only from code regions of README-family docs — not prose, not design docs that merely describe commands. The fourth self-found fix (REV-03/04/05/06).

Self-assessment stays PASS: 9/9 core, 13/13 non-core (100%). SPEC_VERSION → v0.5, re-locked.

assessor-v0.4

Second expansion batch (17 → 20 criteria), still zero-git and deterministic against the static walk.

Self-assessment stays PASS: 8/8 core, 11/11 non-core (100%). SPEC_VERSION → v0.4, re-locked.

assessor-v0.3

First batch of the rubric expansion (13 → 17 criteria), the zero-risk, no-git group. Each is deterministic against a static walk; the corpus and self-assessment stayed green throughout.

REV-05 — the link checker flagged its own example (found by self-assessment). SPEC-03’s first run against this repo failed on a [text](target) written as an example inside a code span in the v0.3 plan doc. A link inside a code fence or `span` is an example, not a link; the checker now strips code before extracting links (standard link-checker behaviour). Same pattern as REV-04 — a finding the tool made against itself, fixed in the open.

Remaining v0.3 criteria (BND-03/04/05, EVO-02/03, SPEC-04/05, VER-06/07) and the git-history domain (PRV-03/04/05, ACC-05) are designed in docs/RUBRIC-v0.3-plan.md and land in later batches — the git ones behind their determinism rules.

assessor-v0.2

The version that submits to its own rubric. v0.1 failed its own standard (1 of 3 core criteria); this release fixes those failures in public before the assessor is pointed at anyone else.

Made it pass its own rubric (P0). Added the things whose absence it flags in others: a spec (SPEC.md), tests (tests/), CI (.github/workflows/ci.yml), a README, a licence, and committed agent instructions (CLAUDE.md). npm run self now returns PASS; CI fails if that ever regresses.

SPEC_VERSION is now enforced (REV-03). In v0.1, patching two criteria silently changed the verdict hash while the version string stayed put — a verdict that can’t be reproduced against a stated version is worthless. criteriaFingerprint() now hashes the criteria definitions, spec-lock.json pins that fingerprint to the version, and scripts/check-spec-version.mjs (in CI) fails on any silent drift. The fingerprint is also stamped into every verdict.

Carried forward from v0.1, found by running it against a repo built to fail:

REV-04 — the marker detector flagged its own definition (found by self-assessment). Running the assessor on itself, SPEC-02 reported three abandoned-subgoal markers in assessor.mjs — but they were the literal strings (TODO/FIXME/XXX/HACK) that define the detector, not real abandoned work. A marker now counts only in tag context (TODO:, TODO(, TODO!, TODO before whitespace/end), not when the words merely appear inside a /- or |-delimited list. The badge already passed (SPEC-02 is non-core), but a false positive in the highest-noise detector is worth removing. The proper fix is comment-aware parsing (a later revision); this is the cheap, correct tightening. This is the method: a finding the tool made against itself, fixed in the open.

Portability fix. Paths are now normalised to forward slashes, so criteria that match on / (test-file detection, CI detection) work on Windows as well as POSIX. A verdict must not depend on the operating system it was produced on.

.assessorignore. A gitignore-style prefix list, read from the repository root, for fixtures the assessor should not treat as product code. This repository uses it to exclude test-corpus/ — the deliberately-broken demo repos it ships as its regression suite — from its own self-assessment.

assessor-v0.1

Initial prototype. Deterministic single-pass assessment, binary + N/A + threshold scoring, content-addressed verdict, thirteen criteria across six domains, seven behavioural tells. Shipped while still embarrassing — on purpose. It failed its own rubric, and that failure is what v0.2 addresses first.