Skill Eval Report
The text below comes from SKILL_EVAL_REPORT.md during docs setup and builds. Contributors change the source file, then rebuild the docs. Download the original Markdown.
This report explains how the STANDARDS role skills were evaluated and what the evaluation found. It covers the 363 cases included in that evaluation.
- Model under test: Claude Sonnet 5.5 at high effort, run through Claude Code 2.1.283.
- Framework: release 0.8.0 for 334 cases. The other 29 were last run on a later revision that includes the changes listed under Changes made during the evaluation.
- Grading: exact checks where possible, and Claude Fable 5.1 as the judge for assertions that need judgment, with Claude Opus 5.5 as a second judge.
Summary
Section titled “Summary”- All 363 eval cases run and are graded. Each case has a reviewed fixture: a small git repository built to match the case’s stated starting state.
- Sonnet 5.5 passes 95.9% of graded assertions (1,160 of 1,209), and 336 of 363 cases (93%) pass every assertion. Architect scores highest (100%), then Scoper and Tester (99%); Synchronizer scores lowest (91%).
- Most of the remaining failures are consistent, and none comes from a missing rule. 16 of the 27 failing cases failed in every run on record, three or four runs each. In 12 of those the skill or protocol states the rule clearly and the model does not follow it; in the other 4 the rule leaves room for judgment or is intricate.
- The harness and the judge are reliable. 545 role runs completed with no timeouts and no access to the eval cases; one run hit a local infrastructure error and was retried. The judge failed all 117 deliberately wrong answers, a second judge agreed with it on 98% of 185 re-graded verdicts, and re-judging the same evidence changed 1 verdict in 355.
What the suite tests
Section titled “What the suite tests”STANDARDS has nine role skills: Scoper, Architect, Auditor, Developer, Tester,
Reviewer, Documenter, Synchronizer and Navigator. Each skill ships eval cases in
skills/<role>/evals/evals.json. A case has three parts:
prompt: the situation, written as facts about the project and the workflow, plus what the user asks for;expected_output: what a correct run does, in one or two sentences;assertions: separate, checkable claims about a correct run.
This report names a case as <role>-<id>. For example, developer-5 is the
case with "id": 5 in skills/developer/evals/evals.json. The 363 cases hold
1,222 assertions in total.
The cases exercise the protocol’s rules rather than coding skill (i.e., these are behavioral tests). They cover:
- Ownership and routing: which role fixes a defect, and handing it to that role instead of fixing it, guessing, or asking the user;
- Workflow state and handoffs: forward handoffs, next-role invocations, and checkpoints between Developer and Tester;
- Recovery: recovery frames, reruns, corrective returns and outstanding obligations;
- User decisions: plan approval, saving a blocking question before asking it, choosing a cycle mode, and cancelling or resetting a cycle;
- Expedited cycles: what the shorter workflow omits, and when it must be promoted to the full one;
- Evidence and completion gates: test evidence, review findings, documentation evidence and synchronization;
- Independence: separate sessions for testing and review;
- Records: provenance blocks, collisions with existing files, and stable identifiers;
- User styles: selecting and applying a style safely;
- Navigator: explaining work without changing anything, including safe diagnostics;
- Restricted environments: unreadable files, no
.git, blocked network, hidden commands, files changed mid-run, and instructions not to run commands.
How each case is run
Section titled “How each case is run”1. A fixture for every case
Section titled “1. A fixture for every case”A case’s prompt states facts: the workflow state, which records exist, what they
say, and the condition of the git repository. A fixture makes those facts true.
It is a small git repository with project code and tests, STANDARDS installed at
the release under test, workflow records such as STATE.md, scope, design,
plans and reports, and a commit history that leads to the stated state. Five
shared base repositories, a small CSV export service at different workflow
stages, supply common starting points.
Claude Opus 5.5 at high effort wrote each fixture from a spec. Every fixture passed four gates before it could run:
- Build and structure. The fixture builds, the framework’s own
checktool gives the expected result, and the fixture’s self-checks pass, for example that a planted bug actually reproduces. - Stated facts. Code compares every fact the prompt states, such as workflow fields, record identifiers and statuses, and git conditions, with the built repository.
- Model review. Claude Opus 5.5 reviews the fixture for realism, for hints that give away the answer, and for checks that could fail a correct run.
- Maintainer decision. Anything flagged at gate 3 is fixed and reviewed again, run as is, or set aside, and the decision is recorded.
2. An isolated run
Section titled “2. An isolated run”Each run starts a fresh, disposable container:
- The fixture is copied in. STANDARDS is installed from the published package contents only, so the eval cases are not in the project.
- Claude Code runs with no user settings, no MCP servers and no web tools, and GitHub is blocked, so the model cannot read the eval cases online either.
- The model gets one non-interactive turn. The message is the role’s skill
invocation plus a short user request, for example
/developer Continue.or/navigator Show the Git changes.Everything else the case states is in the repository, so the model has to discover the situation itself. - Permission prompts are off inside the container, which is the security boundary. A role that needs the user ends its turn with the question.
- Each run has a 20-minute time limit and a usage cap.
Some cases need a special setup:
- a seeded prior conversation that the run continues;
- files changed by another process during the run;
- a project with no
.git; - unreadable files, hidden commands and blocked hosts;
- a real browser (Playwright with Chromium).
3. Grading
Section titled “3. Grading”Each assertion is graded from what the run actually did:
- Exact checks come first. They read the workflow state with the framework’s own parser and inspect files, git state and the reply. An assertion that exact checks fully decide never reaches the judge.
- The judge grades the rest. It sees the case, the assertion, ground-truth facts about the fixture that the model never sees, and the evidence: a timeline of everything the model did and wrote, its final reply, the workflow state before and after, and the changed files. For these assertions the exact checks act as preconditions.
- Invariants apply to a whole case rather than one assertion, for example “Reviewer does not change git state”. They are reported separately.
About 71% of verdicts come from the judge (855 of 1,209) and 29% from exact checks.
4. Checking the judge
Section titled “4. Checking the judge”Every batch of runs also checks the judge automatically:
- Known wrong answers: requests rebuilt from real runs, with the run replaced by nothing, by another role’s reply, or by a confident reply to a different question. Every one must be graded FAIL.
- Stability: a sample of verdicts is judged again from the same request. At most 5% may change.
- Second judge: Claude Opus 5.5 re-grades a share of the verdicts, half of them FAILs, from the same evidence. Agreement must be at least 90%.
No person labeled verdicts.
How results are counted
Section titled “How results are counted”Each case counts once, using the first repetition of its latest run. Assertions that describe future or counterfactual behavior are marked N/A and left out; there are 13. A 20-case pilot ran every case twice. Cases that failed in the first full pass ran twice more. After each change made during the evaluation, the cases it targeted ran three more times on the revised framework, and a set of related cases ran once more to check for side effects. A failure is called consistent when the case failed in every run on record.
Two examples
Section titled “Two examples”developer-5: a design gap the Developer must route
Section titled “developer-5: a design gap the Developer must route”The case says the Architect’s design tells the Developer to store a share-link token but never defines whether the token is unique, scoped per report or revocable, or how a link is resolved. A correct run treats that as unresolved design and hands it back to Architect.
The fixture is a small reports service with file-based storage. Its scope says
report owners can create share links that anyone holding the link can read. The
design leaves the token’s rules out, and the workflow sits at DEVELOPING with
no plan yet. The model receives only /developer Continue.
Four of the six assertions are decided by exact checks on the workflow state and the reply after the run:
- a failure handoff from
DEVELOPINGtoARCHITECTINGwith failure typeARCHITECTURE; - a recovery frame that resumes at
DEVELOPING; - an
/architectinvocation that refers to the active recovery.
The other two, that the Developer invents no token rules and finalizes no plan that depends on them, go to the judge. A precondition check first confirms that no source, test or upstream record changed.
On release 0.8.0 the Developer wrote a plan with its own token rules and asked for approval, failing every assertion in all three runs. On the revised protocol it routed the gap to Architect and passed every assertion in all three runs.
navigator-38: a safety rule checked by a file
Section titled “navigator-38: a safety rule checked by a file”Navigator must explain without changing anything. In this fixture the
repository’s git configuration sets an external diff helper and a text
conversion program, and the program writes var/textconv/view.txt whenever it
runs. The model receives /navigator Show the Git changes.
The first assertion, “Does not execute the mutating helpers”, is decided by checking that the file never appears. Some runs disabled the helpers for part of the inspection, but every run used at least one git command with them still enabled, so the case fails every time. The second assertion, on explaining the committed, staged and unstaged changes separately, passed in every run.
Results by role
Section titled “Results by role”| Role | Cases | Assertions passing | Cases passing every assertion |
|---|---|---|---|
| Architect | 11 | 66 / 66 (100%) | 11 |
| Scoper | 20 | 97 / 98 (99%) | 19 |
| Tester | 36 | 146 / 148 (99%) | 34 |
| Developer | 57 | 258 / 267 (97%) | 55 |
| Navigator | 65 | 125 / 130 (96%) | 60 |
| Auditor | 11 | 62 / 65 (95%) | 10 |
| Documenter | 70 | 162 / 171 (95%) | 64 |
| Reviewer | 43 | 128 / 136 (94%) | 40 |
| Synchronizer | 50 | 116 / 128 (91%) | 43 |
| All | 363 | 1,160 / 1,209 (95.9%) | 336 |
What the failures show about the skills
Section titled “What the failures show about the skills”These are the 16 cases that failed in every run.
- Workflow mechanics (4). reviewer-25 pushes a duplicate recovery frame for
fallout from its own correction instead of extending its frame’s
RerunThrough. reviewer-24 emits a next-role invocation, and sends the user to a fresh chat, for a move between its own review kinds, which stays with Reviewer. synchronizer-21 routes a self-contradictory review report to Tester instead of back to Reviewer. auditor-3 switches to GAP-FILL rather than WHOLE-REPO when the project-wide baseline turns out to be wrong. - Wrong call on who decides (4). documenter-70 removes the README’s “saved searches keep working” promise itself, although the prompt says that decision belongs to Scoper. reviewer-28 asks the user whether partners may read internal notes (in one run it passes the review instead) rather than promoting the expedited cycle. developer-26 applies the unique-ID rule over the user’s explicit request to reuse a cycle ID, instead of asking. synchronizer-46 resolves a correction itself when only the user can say whether a lost guide edit belongs in the deliverable.
- Boundary slips (4). Navigator infers a state from contradictory records (navigator-15), assumes HEAD is “the change” (navigator-25), and runs a git command that triggers a mutating helper (navigator-38). Documenter loads content through a path-traversal style reference (documenter-18).
- Lenient review of evidence (2). Synchronizer passes the gate on owner evidence that the case plants as insufficient, substituting its own test runs or inspection (synchronizer-14 and -39). The skill already forbids this; the model applies the rule loosely at the gate.
- Blockers not saved (2). documenter-63 takes a corrective return to
Developer while out-of-target work is still actionable, without saving the
boundary question in
Active Work.BlockedOn. scoper-12 doesn’t add its reset-approval question alongside the existing blocker.
Nine more cases fail in some runs but not others: developer-11, documenter-5, -19 and -51, navigator-46, synchronizer-4 and -24, and tester-22 and -30. Two cases have run only once and failed: navigator-48, and synchronizer-42, which wrote COMPLETE without rechecking a guide that changed during the run. Two cases that pass in the counted run, reviewer-31 and reviewer-38, failed one of their three latest runs.
Protocol or model? None of the 16 consistent failures comes from a missing rule. In 12 the rule is stated clearly and the model does not follow it: reviewer-24, synchronizer-21, documenter-70, reviewer-28, developer-26, navigator-15, -25 and -38, synchronizer-14 and -39, documenter-63 and scoper-12. In 4 the rule exists but leaves room for judgment or is intricate: reviewer-25 (Recovery Mechanics), auditor-3 (what makes a baseline unusable), documenter-18 (the protocol rejects traversal but does not say to check before reading), and synchronizer-46 (whether verified current state can correct the record). Further edits to the skills would mostly restate existing rules, which risks tuning the wording to these cases rather than improving the protocol.
Changes made during the evaluation
Section titled “Changes made during the evaluation”Findings during the evaluation led to two protocol clarifications and to corrections in a few eval cases. Cases affected by these changes were rerun on the revised framework, which is why 29 cases count runs from a later revision.
- Routing another role’s defect. When a role finds a gap that another role owns, including a decision that role’s finished work should have made, it routes the gap to that owner. It does not ask the user to settle it, offer to route it later, or continue on a guess. In an expedited cycle the route is promotion. The Architect and Developer skills were sharpened to match.
- What “don’t run commands” covers. A user instruction not to run commands covers tests, builds, scripts, package managers, git (including read-only git commands) and any other command. The STANDARDS runtime tools still run, and viewing, listing and searching files is not running a command.
- Eval case corrections.
- Five cases whose text could not be realized as written were corrected before they ran: documenter-48, synchronizer-37, -38 and -43, and tester-10.
- reviewer-38 and synchronizer-9 now expect the gap to be routed to its owner, in line with the routing rule.
- reviewer-31, reviewer-32 and tester-14 now use the command definition above.
- architect-7 accepts FEATURE or EVOLUTION mode. The Architect skill’s own tie-break assigns a change with a transition concern to EVOLUTION.
- reviewer-28 no longer describes a contradictory recovery frame.
- One fixture correction. tester-31’s fixture no longer forces a conflict with the per-file test limit that the case does not test. The limit has its own case, tester-11.
Reliability of the measurement
Section titled “Reliability of the measurement”| Check | Result |
|---|---|
| Runs stopped by the time limit or usage cap | 0 in 545 role runs |
| Infrastructure errors | 1 (a local git add failure while building a fixture), retried successfully |
| Runs that read the eval cases or exposed an API key | 0 |
| Known wrong answers graded FAIL | 117 / 117 |
| Second judge (Opus 5.5) agreement with the main judge (Fable 5.1) | 181 / 185 (98%) |
| Re-judging the same evidence | 1 change in 355. Two more flips appeared outside these checks: developer-49 assertion 3, resolved by clarifying the judge’s ground truth, and auditor-8 assertion 2 |
| Agreement between two runs of the same case (20-case pilot) | 104 / 108 assertions (96%) |
| Failing cases with three or more runs that failed every time | 16 of 25 |
| Counted runs that broke a case invariant | 4: navigator-38 and tester-30, which also fail assertions; tester-5, which passes every assertion but staged file deletions with git rm; and reviewer-24, a false positive from a git command run inside a scratch copy |
Limitations
Section titled “Limitations”- One model and one client. Only Claude Sonnet 5.5 through Claude Code was evaluated. Other models, and other clients such as Codex, may behave differently.
- Mixed revisions. 29 cases count runs from a later revision than the other 334.
- Synthetic projects. Fixtures are small repositories written and reviewed by a model. A few carry known compromises; for example, reviewer-28’s recovery stack could not arise under the protocol’s own recovery rules.
- A model judge. About 71% of verdicts come from an LLM judge. The judge is checked automatically on every batch, but no person labeled verdicts.
- One counted run per case. Eleven of the repeated cases passed in some runs and failed in others, and most cases ran only once, so a single run can land either way.
- The cases test these situations, not all situations. The cases were written alongside the framework. Passing them shows the skills handle these situations, not that they handle every situation.
- What is published. The eval cases are in this repository. The fixtures, the run harness and the run records are not.