acg-assessor

The Assessor — Rubric Specification

binary · deterministic · threshold-gated · the spec version is stamped on the criteria list below and on every verdict, so this document never carries a version of its own to go stale

This document is the rubric, and §6 is generated from assessor.mjs — the criteria listed there are the criteria the program applies, rendered, not a description of them maintained alongside.

⚑ It was not always. Until v0.8 this file described thirteen criteria while the program applied twenty-seven, so fourteen rules — five of them core — could fail a repository without appearing in the published rubric at all. Prose and code cannot be kept in step by intention; scripts/sync-spec.mjs --check runs in CI and fails the build if they part.

1 · Why this exists

Two independent assessors examining the same repository should reach substantially the same verdict. If two practitioners can look at the same codebase and disagree about whether it passes, we don’t have a rubric — we have opinions with a logo on them.

Two humans diverge. A deterministic program cannot. Same repository → same verdict → same hash. Inter-assessor agreement is 1.00 by construction, not by training. Everything below serves that.

2 · Scoring model

Binary, with an explicit N/A, and a published threshold. No grading, ever.

N/A criteria are excluded from both counts — they neither help nor harm the badge.

3 · Determinism guarantees

Violate any of these and the thesis dies:

Fixtures the assessor should not treat as product code (test corpora, vendored copies) are listed in .assessorignore, a gitignore-style prefix list read from the repository root.

4 · The six domains

domain prefix asks
specification integrity SPEC- is there a durable record of what the code was meant to do?
verification integrity VER- is the code actually exercised, and is that exercise honest?
agent boundaries BND- is the agent constrained, committed, and reviewable?
human accountability ACC- did a human own this?
evolvability EVO- can the next change be made safely?
provenance PRV- is this the artefact that was reviewed, and can you trace how it got here?

Criterion IDs are stable, permanent, and never reused. SPEC-03 retired is SPEC-03 retired forever; the next new one is the next unused number. Assessors and clients cite them for years.

5 · The seven tells

Behavioural signatures of agent-generated code. Each criterion is tagged with the tell it detects; the verdict reports the dominant tell — the failure mode that dominates the NOT_MET results.

tell signature
UNSPENT declared and never used (deps, imports, exports)
UNOPENED code paths with no test that exercises them
REPEAT near-identical blocks regenerated rather than factored
PASSED tests skipped, pending, or always-true
COLLAPSED abandoned subgoals left in place (TODO/FIXME/XXX)
ECHOED scaffold/boilerplate retained unmodified
INERT unreachable or no-op code

6 · The criteria

27 criteria · 11 core (marked ●) · spec assessor-v0.10.

This section is generated from the criteria the assessor actually applies. It is not a description of the program — it is the program’s own list, rendered. Each entry carries the tell it detects and the failure it was written from.

specification integrity

verification integrity

agent boundaries

human accountability

evolvability

provenance

Tells: UNSPENT · UNOPENED · REPEAT · PASSED · COLLAPSED · ECHOED · INERT.

N/A conditions are not listed here. They are decided per repository and printed with a written justification on every run, because whether a criterion applies is a fact about the repository in front of you, not about the rubric.

7 · Versioning and reproducibility

A verdict is only meaningful relative to the spec version it was issued under.

8 · Boundaries (liability, not preference)

9 · The regression corpus

test-corpus/clean-demo must PASS; test-corpus/slop-demo must FAIL (exit 1); this repository must PASS. Any criterion change is run against all three. If clean-demo ever fails, a criterion over-fires. If slop-demo ever passes, a criterion is toothless.

10 · Revision

Openly licensed and versioned in public. Every change is a PR with a rationale; the changelog says why. When a criterion changes because it met a real codebase and lost, publish the revision and the engagement that caused it (sanitised). A rubric that visibly changes on contact with reality is more credible than one that arrived complete.