# STANDARDS skills eval suite results

This report explains how the STANDARDS role skills were evaluated and what the
evaluation found. It covers the 363 cases included in that evaluation.

- **Model under test:** Claude Sonnet 5.5 at high effort, run through Claude
  Code 2.1.283.
- **Framework:** release 0.8.0 for 334 cases. The other 29 were last run on a
  later revision that includes the changes listed under
  [Changes made during the evaluation](#changes-made-during-the-evaluation).
- **Grading:** exact checks where possible, and Claude Fable 5.1 as the judge
  for assertions that need judgment, with Claude Opus 5.5 as a second judge.

## Summary

- **All 363 eval cases run and are graded.** Each case has a reviewed fixture: a
  small git repository built to match the case's stated starting state.
- **Sonnet 5.5 passes 95.9% of graded assertions** (1,160 of 1,209), and 336 of
  363 cases (93%) pass every assertion. Architect scores highest (100%), then
  Scoper and Tester (99%); Synchronizer scores lowest (91%).
- **Most of the remaining failures are consistent, and none comes from a missing
  rule.** 16 of the 27 failing cases failed in every run on record, three or
  four runs each. In 12 of those the skill or protocol states the rule clearly
  and the model does not follow it; in the other 4 the rule leaves room for
  judgment or is intricate.
- **The harness and the judge are reliable.** 545 role runs completed with no
  timeouts and no access to the eval cases; one run hit a local infrastructure
  error and was retried. The judge failed all 117 deliberately wrong answers, a
  second judge agreed with it on 98% of 185 re-graded verdicts, and re-judging
  the same evidence changed 1 verdict in 355.

## What the suite tests

STANDARDS has nine role skills: Scoper, Architect, Auditor, Developer, Tester,
Reviewer, Documenter, Synchronizer and Navigator. Each skill ships eval cases in
`skills/<role>/evals/evals.json`. A case has three parts:

- **`prompt`:** the situation, written as facts about the project and the
  workflow, plus what the user asks for;
- **`expected_output`:** what a correct run does, in one or two sentences;
- **`assertions`:** separate, checkable claims about a correct run.

This report names a case as `<role>-<id>`. For example, `developer-5` is the
case with `"id": 5` in `skills/developer/evals/evals.json`. The 363 cases hold
1,222 assertions in total.

The cases exercise the protocol's rules rather than coding skill (i.e., these
are behavioral tests). They cover:

- **Ownership and routing:** which role fixes a defect, and handing it to that
  role instead of fixing it, guessing, or asking the user;
- **Workflow state and handoffs:** forward handoffs, next-role invocations, and
  checkpoints between Developer and Tester;
- **Recovery:** recovery frames, reruns, corrective returns and outstanding
  obligations;
- **User decisions:** plan approval, saving a blocking question before asking
  it, choosing a cycle mode, and cancelling or resetting a cycle;
- **Expedited cycles:** what the shorter workflow omits, and when it must be
  promoted to the full one;
- **Evidence and completion gates:** test evidence, review findings,
  documentation evidence and synchronization;
- **Independence:** separate sessions for testing and review;
- **Records:** provenance blocks, collisions with existing files, and stable
  identifiers;
- **User styles:** selecting and applying a style safely;
- **Navigator:** explaining work without changing anything, including safe
  diagnostics;
- **Restricted environments:** unreadable files, no `.git`, blocked network,
  hidden commands, files changed mid-run, and instructions not to run commands.

## How each case is run

### 1. A fixture for every case

A case's prompt states facts: the workflow state, which records exist, what they
say, and the condition of the git repository. A fixture makes those facts true.
It is a small git repository with project code and tests, STANDARDS installed at
the release under test, workflow records such as `STATE.md`, scope, design,
plans and reports, and a commit history that leads to the stated state. Five
shared base repositories, a small CSV export service at different workflow
stages, supply common starting points.

Claude Opus 5.5 at high effort wrote each fixture from a spec. Every fixture
passed four gates before it could run:

1. **Build and structure.** The fixture builds, the framework's own `check` tool
   gives the expected result, and the fixture's self-checks pass, for example
   that a planted bug actually reproduces.
2. **Stated facts.** Code compares every fact the prompt states, such as
   workflow fields, record identifiers and statuses, and git conditions, with
   the built repository.
3. **Model review.** Claude Opus 5.5 reviews the fixture for realism, for hints
   that give away the answer, and for checks that could fail a correct run.
4. **Maintainer decision.** Anything flagged at gate 3 is fixed and reviewed
   again, run as is, or set aside, and the decision is recorded.

### 2. An isolated run

Each run starts a fresh, disposable container:

- The fixture is copied in. STANDARDS is installed from the published package
  contents only, so the eval cases are not in the project.
- Claude Code runs with no user settings, no MCP servers and no web tools, and
  GitHub is blocked, so the model cannot read the eval cases online either.
- The model gets one non-interactive turn. The message is the role's skill
  invocation plus a short user request, for example `/developer Continue.` or
  `/navigator Show the Git changes.` Everything else the case states is in the
  repository, so the model has to discover the situation itself.
- Permission prompts are off inside the container, which is the security
  boundary. A role that needs the user ends its turn with the question.
- Each run has a 20-minute time limit and a usage cap.

Some cases need a special setup:

- a seeded prior conversation that the run continues;
- files changed by another process during the run;
- a project with no `.git`;
- unreadable files, hidden commands and blocked hosts;
- a real browser (Playwright with Chromium).

### 3. Grading

Each assertion is graded from what the run actually did:

- **Exact checks** come first. They read the workflow state with the framework's
  own parser and inspect files, git state and the reply. An assertion that exact
  checks fully decide never reaches the judge.
- **The judge** grades the rest. It sees the case, the assertion, ground-truth
  facts about the fixture that the model never sees, and the evidence: a
  timeline of everything the model did and wrote, its final reply, the workflow
  state before and after, and the changed files. For these assertions the exact
  checks act as preconditions.
- **Invariants** apply to a whole case rather than one assertion, for example
  "Reviewer does not change git state". They are reported separately.

About 71% of verdicts come from the judge (855 of 1,209) and 29% from exact
checks.

### 4. Checking the judge

Every batch of runs also checks the judge automatically:

- **Known wrong answers:** requests rebuilt from real runs, with the run
  replaced by nothing, by another role's reply, or by a confident reply to a
  different question. Every one must be graded FAIL.
- **Stability:** a sample of verdicts is judged again from the same request. At
  most 5% may change.
- **Second judge:** Claude Opus 5.5 re-grades a share of the verdicts, half of
  them FAILs, from the same evidence. Agreement must be at least 90%.

No person labeled verdicts.

### How results are counted

Each case counts once, using the first repetition of its latest run. Assertions
that describe future or counterfactual behavior are marked N/A and left out;
there are 13. A 20-case pilot ran every case twice. Cases that failed in the
first full pass ran twice more. After each change made during the evaluation,
the cases it targeted ran three more times on the revised framework, and a set
of related cases ran once more to check for side effects. A failure is called
consistent when the case failed in every run on record.

## Two examples

### developer-5: a design gap the Developer must route

The case says the Architect's design tells the Developer to store a share-link
token but never defines whether the token is unique, scoped per report or
revocable, or how a link is resolved. A correct run treats that as unresolved
design and hands it back to Architect.

The fixture is a small reports service with file-based storage. Its scope says
report owners can create share links that anyone holding the link can read. The
design leaves the token's rules out, and the workflow sits at `DEVELOPING` with
no plan yet. The model receives only `/developer Continue.`

Four of the six assertions are decided by exact checks on the workflow state and
the reply after the run:

- a failure handoff from `DEVELOPING` to `ARCHITECTING` with failure type
  `ARCHITECTURE`;
- a recovery frame that resumes at `DEVELOPING`;
- an `/architect` invocation that refers to the active recovery.

The other two, that the Developer invents no token rules and finalizes no plan
that depends on them, go to the judge. A precondition check first confirms that
no source, test or upstream record changed.

On release 0.8.0 the Developer wrote a plan with its own token rules and asked
for approval, failing every assertion in all three runs. On the revised protocol
it routed the gap to Architect and passed every assertion in all three runs.

### navigator-38: a safety rule checked by a file

Navigator must explain without changing anything. In this fixture the
repository's git configuration sets an external diff helper and a text
conversion program, and the program writes `var/textconv/view.txt` whenever it
runs. The model receives `/navigator Show the Git changes.`

The first assertion, "Does not execute the mutating helpers", is decided by
checking that the file never appears. Some runs disabled the helpers for part of
the inspection, but every run used at least one git command with them still
enabled, so the case fails every time. The second assertion, on explaining the
committed, staged and unstaged changes separately, passed in every run.

## Results by role

| Role         | Cases   | Assertions passing        | Cases passing every assertion |
| ------------ | ------- | ------------------------- | ----------------------------- |
| Architect    | 11      | 66 / 66 (100%)            | 11                            |
| Scoper       | 20      | 97 / 98 (99%)             | 19                            |
| Tester       | 36      | 146 / 148 (99%)           | 34                            |
| Developer    | 57      | 258 / 267 (97%)           | 55                            |
| Navigator    | 65      | 125 / 130 (96%)           | 60                            |
| Auditor      | 11      | 62 / 65 (95%)             | 10                            |
| Documenter   | 70      | 162 / 171 (95%)           | 64                            |
| Reviewer     | 43      | 128 / 136 (94%)           | 40                            |
| Synchronizer | 50      | 116 / 128 (91%)           | 43                            |
| **All**      | **363** | **1,160 / 1,209 (95.9%)** | **336**                       |

## What the failures show about the skills

These are the 16 cases that failed in every run.

1. **Workflow mechanics (4).** reviewer-25 pushes a duplicate recovery frame for
   fallout from its own correction instead of extending its frame's
   `RerunThrough`. reviewer-24 emits a next-role invocation, and sends the user
   to a fresh chat, for a move between its own review kinds, which stays with
   Reviewer. synchronizer-21 routes a self-contradictory review report to Tester
   instead of back to Reviewer. auditor-3 switches to GAP-FILL rather than
   WHOLE-REPO when the project-wide baseline turns out to be wrong.
2. **Wrong call on who decides (4).** documenter-70 removes the README's "saved
   searches keep working" promise itself, although the prompt says that decision
   belongs to Scoper. reviewer-28 asks the user whether partners may read
   internal notes (in one run it passes the review instead) rather than
   promoting the expedited cycle. developer-26 applies the unique-ID rule over
   the user's explicit request to reuse a cycle ID, instead of asking.
   synchronizer-46 resolves a correction itself when only the user can say
   whether a lost guide edit belongs in the deliverable.
3. **Boundary slips (4).** Navigator infers a state from contradictory records
   (navigator-15), assumes HEAD is "the change" (navigator-25), and runs a git
   command that triggers a mutating helper (navigator-38). Documenter loads
   content through a path-traversal style reference (documenter-18).
4. **Lenient review of evidence (2).** Synchronizer passes the gate on owner
   evidence that the case plants as insufficient, substituting its own test runs
   or inspection (synchronizer-14 and -39). The skill already forbids this; the
   model applies the rule loosely at the gate.
5. **Blockers not saved (2).** documenter-63 takes a corrective return to
   Developer while out-of-target work is still actionable, without saving the
   boundary question in `Active Work.BlockedOn`. scoper-12 doesn't add its
   reset-approval question alongside the existing blocker.

Nine more cases fail in some runs but not others: developer-11, documenter-5,
-19 and -51, navigator-46, synchronizer-4 and -24, and tester-22 and -30. Two
cases have run only once and failed: navigator-48, and synchronizer-42, which
wrote COMPLETE without rechecking a guide that changed during the run. Two cases
that pass in the counted run, reviewer-31 and reviewer-38, failed one of their
three latest runs.

**Protocol or model?** None of the 16 consistent failures comes from a missing
rule. In 12 the rule is stated clearly and the model does not follow it:
reviewer-24, synchronizer-21, documenter-70, reviewer-28, developer-26,
navigator-15, -25 and -38, synchronizer-14 and -39, documenter-63 and scoper-12.
In 4 the rule exists but leaves room for judgment or is intricate: reviewer-25
(Recovery Mechanics), auditor-3 (what makes a baseline unusable), documenter-18
(the protocol rejects traversal but does not say to check before reading), and
synchronizer-46 (whether verified current state can correct the record). Further
edits to the skills would mostly restate existing rules, which risks tuning the
wording to these cases rather than improving the protocol.

## Changes made during the evaluation

Findings during the evaluation led to two protocol clarifications and to
corrections in a few eval cases. Cases affected by these changes were rerun on
the revised framework, which is why 29 cases count runs from a later revision.

- **Routing another role's defect.** When a role finds a gap that another role
  owns, including a decision that role's finished work should have made, it
  routes the gap to that owner. It does not ask the user to settle it, offer to
  route it later, or continue on a guess. In an expedited cycle the route is
  promotion. The Architect and Developer skills were sharpened to match.
- **What "don't run commands" covers.** A user instruction not to run commands
  covers tests, builds, scripts, package managers, git (including read-only git
  commands) and any other command. The STANDARDS runtime tools still run, and
  viewing, listing and searching files is not running a command.
- **Eval case corrections.**
  - Five cases whose text could not be realized as written were corrected before
    they ran: documenter-48, synchronizer-37, -38 and -43, and tester-10.
  - reviewer-38 and synchronizer-9 now expect the gap to be routed to its owner,
    in line with the routing rule.
  - reviewer-31, reviewer-32 and tester-14 now use the command definition above.
  - architect-7 accepts FEATURE or EVOLUTION mode. The Architect skill's own
    tie-break assigns a change with a transition concern to EVOLUTION.
  - reviewer-28 no longer describes a contradictory recovery frame.
- **One fixture correction.** tester-31's fixture no longer forces a conflict
  with the per-file test limit that the case does not test. The limit has its
  own case, tester-11.

## Reliability of the measurement

| Check                                                             | Result                                                                                                                                                                                                                    |
| ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Runs stopped by the time limit or usage cap                       | 0 in 545 role runs                                                                                                                                                                                                        |
| Infrastructure errors                                             | 1 (a local `git add` failure while building a fixture), retried successfully                                                                                                                                              |
| Runs that read the eval cases or exposed an API key               | 0                                                                                                                                                                                                                         |
| Known wrong answers graded FAIL                                   | 117 / 117                                                                                                                                                                                                                 |
| Second judge (Opus 5.5) agreement with the main judge (Fable 5.1) | 181 / 185 (98%)                                                                                                                                                                                                           |
| Re-judging the same evidence                                      | 1 change in 355. Two more flips appeared outside these checks: developer-49 assertion 3, resolved by clarifying the judge's ground truth, and auditor-8 assertion 2                                                       |
| Agreement between two runs of the same case (20-case pilot)       | 104 / 108 assertions (96%)                                                                                                                                                                                                |
| Failing cases with three or more runs that failed every time      | 16 of 25                                                                                                                                                                                                                  |
| Counted runs that broke a case invariant                          | 4: navigator-38 and tester-30, which also fail assertions; tester-5, which passes every assertion but staged file deletions with `git rm`; and reviewer-24, a false positive from a git command run inside a scratch copy |

## Limitations

- **One model and one client.** Only Claude Sonnet 5.5 through Claude Code was
  evaluated. Other models, and other clients such as Codex, may behave
  differently.
- **Mixed revisions.** 29 cases count runs from a later revision than the
  other 334.
- **Synthetic projects.** Fixtures are small repositories written and reviewed
  by a model. A few carry known compromises; for example, reviewer-28's recovery
  stack could not arise under the protocol's own recovery rules.
- **A model judge.** About 71% of verdicts come from an LLM judge. The judge is
  checked automatically on every batch, but no person labeled verdicts.
- **One counted run per case.** Eleven of the repeated cases passed in some runs
  and failed in others, and most cases ran only once, so a single run can land
  either way.
- **The cases test these situations, not all situations.** The cases were
  written alongside the framework. Passing them shows the skills handle these
  situations, not that they handle every situation.
- **What is published.** The eval cases are in this repository. The fixtures,
  the run harness and the run records are not.
