Files
magnus919_agent-skills/neckbeard/references/evaluation.md
T
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> a315b77bd8 feat(flatten): rewrite moved-file paths for flat layout
Re-point relative references inside the six non-tailscale/workflow-architect
bundle dirs now that they live one level shallower at the repo root:
- bundle-root SKILL.md/manifest.yaml/README.md: ../../<target> -> ../<target>,
  cross-bundle ../../bundles/<x> -> ../<x>
- forward-deployed-engineering references/: ../../../<target> -> ../../<target>
- product-lifecycle references cross-bundle ../../bundles/neckbeard -> ../../neckbeard
- product-lifecycle references/discovery-brief.md prose headings drop bundles/ prefix
- neckbeard eval shell commands bundles/neckbeard/eval -> neckbeard/eval
- production-excellence AGENTS.md depth note updated
- manifest header comments point at ../schemas/bundle-manifest-v1.schema.json
- regenerate docs/lifecycle-capability-matrix.{md,json}

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-14 15:59:56 -04:00

152 lines
6.3 KiB
Markdown

# Evaluation Methodology
The evaluation harness lives in [../eval/](../eval/). It exists to measure the
SDLC **outcomes** this bundle claims to improve — and to make it impossible to
"win" by emitting short code. This file is the method; `eval/` is the tooling.
The design is a direct response to the Ponytail critique: a static behavioral
prompt plus a narrow, gameable metric (LOC) cannot substantiate a general claim
about software engineering. So the metric here is never LOC or brevity.
## What is measured (outcome rubric)
Score each run on these dimensions. LOC may appear only as **diagnostic
metadata**, never as a success proxy.
| Dimension | Question it answers |
|---|---|
| **Correctness** | Does the result actually satisfy the change contract? |
| **Regression safety** | Did it avoid breaking existing behavior/tests? |
| **Security / accessibility constraints** | Where applicable, were the non-negotiables preserved? |
| **Test adequacy** | Are the checks sufficient for the declared boundary? |
| **Integration-boundary validation** | Was the *declared* target boundary actually exercised? |
| **Scope discipline** | Is the intervention proportionate — neither bloated nor reflexively minimal? |
| **Maintainability** | Can a human read, review, and extend it? |
| **Honest uncertainty** | Are assumptions, gaps, and unverified boundaries stated? |
| **Time / cost** | Only if measured; reported, never used alone to claim a win. |
Full rubric with scoring anchors: [../eval/rubric.md](../eval/rubric.md).
**Trajectory scoring.** These dimensions apply to trajectory runs exactly as
they apply to single-task runs. A trajectory is scored on its *outcomes*
correctness, regression safety, security/accessibility, test adequacy,
integration-boundary validation, scope discipline, maintainability, and honest
uncertainty — not on phase count, response length, or response shape. A
trajectory that reaches the correct terminal state with proper gate evidence
and skip transparency scores well regardless of how many phases it enumerated.
Shape-neutrality (see [../eval/baseline-protocol.md](../eval/baseline-protocol.md))
extends to trajectory comparisons: do not reward a run for visiting more phases
or penalize it for a different recording format.
## Task fixtures
Representative, repository-backed tasks across these classes:
- bug diagnosis
- feature change
- refactor
- specification ambiguity
- regression prevention, including test-hardening guards for already-correct
production behavior
- review finding
- release verification
- **"no change needed"** cases
Plus **adversarial / counterfactual** cases where the correct answer is a
*larger* change, a new dependency, a non-code process change, or no code change
at all. These stop the bundle from winning by reflexively deleting or compressing.
Each fixture carries its repository context and harness constraints. Schema:
[../eval/task-schema.md](../eval/task-schema.md). Fixtures: [../eval/fixtures/](../eval/fixtures/).
## Trajectory fixtures
In addition to single-task fixtures, the suite includes **trajectory
fixtures** — multi-phase change-request journeys that record the full sequence
of journey phases, gates, routing decisions, and terminal state. Where a
single-task fixture has one `class` and `expected_boundary`, a
trajectory fixture describes an end-to-end delivery (e.g. all nine phases to a
released state, a lightweight path ending closed with recorded skips, or a
lightweight test-hardening path with clean-baseline and targeted-mutant
verification).
Trajectory fixtures are validated by the same runner using `kind: trajectory`
as the discriminator. Their schema is documented in
[../eval/task-schema.md](../eval/task-schema.md#trajectory-fixtures).
## Running the integrated suite
The runner ([../eval/run_eval.py](../eval/run_eval.py)) validates both
single-task and trajectory fixtures in one pass:
```sh
.venv/bin/python3 neckbeard/eval/run_eval.py \
--suite neckbeard/eval/fixtures --validate-only
```
Expected output (current fixture set):
```
OK: 14 fixture(s) valid.
11 single-task fixture(s) valid.
3 trajectory fixture(s) valid.
by class: {'adversarial': 2, 'bug-diagnosis': 1, 'feature-change': 1, 'no-change-needed': 1, 'refactor': 1, 'regression-prevention': 2, 'release-verification': 1, 'review-finding': 1, 'spec-ambiguity': 1}
by visibility: {'public': 11}
adversarial: 2
trajectory paths: {'full': 1, 'lightweight': 2}
```
To scaffold a scoring report for manual evaluation:
```sh
.venv/bin/python3 neckbeard/eval/run_eval.py \
--suite neckbeard/eval/fixtures \
--report neckbeard/eval/out/report.md
```
## Holdout discipline
Keep a task set **separate** from author iteration. Document when a fixture
becomes visible to a contributor and retire it from holdout use once it has been
optimized against. Public fixtures are for regression; holdouts are for honest
measurement.
## Fair baselines
Compare against a **context-equivalent** agent/harness. Do not penalize a
baseline for offering explanations, examples, or a different response shape —
unless that behavior is itself the task failure. The baseline must see the same
repository context and constraints.
## Multi-run, multi-model reporting
Report, for every result:
- model and model version (where available)
- harness / system prompt
- tools available
- fixture revision
- randomization settings
- run count and variance / confidence intervals
Never collapse one favorable point estimate into a universal claim.
## Reproducible artifacts
Retain: prompts, fixtures, scoring rubric, commands, raw anonymized outputs
(when licensing permits), and the aggregation script. Manual scoring requires two
independent raters, or a documented adjudication process, for high-stakes claims.
## Regression gate
A change to the bundle cannot claim improvement without running the public suite
and reporting holdout results through the maintainers' controlled workflow.
## Claims policy
Scope every performance claim to the evaluated **models, harnesses, fixture
revision (git SHA), repositories, task classes, and run dates**. No result
generalizes beyond the specific harness version, model, fixture set, and
revision that produced it. Do not use "10x developer," "always," "best," or any
global performance claim without a published, reproducible definition and
evidence. Report template: [../templates/eval-report.md](../templates/eval-report.md).