Files
magnus919_agent-skills/neckbeard/references/evaluation.md
T
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> a315b77bd8 feat(flatten): rewrite moved-file paths for flat layout
Re-point relative references inside the six non-tailscale/workflow-architect
bundle dirs now that they live one level shallower at the repo root:
- bundle-root SKILL.md/manifest.yaml/README.md: ../../<target> -> ../<target>,
  cross-bundle ../../bundles/<x> -> ../<x>
- forward-deployed-engineering references/: ../../../<target> -> ../../<target>
- product-lifecycle references cross-bundle ../../bundles/neckbeard -> ../../neckbeard
- product-lifecycle references/discovery-brief.md prose headings drop bundles/ prefix
- neckbeard eval shell commands bundles/neckbeard/eval -> neckbeard/eval
- production-excellence AGENTS.md depth note updated
- manifest header comments point at ../schemas/bundle-manifest-v1.schema.json
- regenerate docs/lifecycle-capability-matrix.{md,json}

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-14 15:59:56 -04:00

6.3 KiB

Evaluation Methodology

The evaluation harness lives in ../eval/. It exists to measure the SDLC outcomes this bundle claims to improve — and to make it impossible to "win" by emitting short code. This file is the method; eval/ is the tooling.

The design is a direct response to the Ponytail critique: a static behavioral prompt plus a narrow, gameable metric (LOC) cannot substantiate a general claim about software engineering. So the metric here is never LOC or brevity.

What is measured (outcome rubric)

Score each run on these dimensions. LOC may appear only as diagnostic metadata, never as a success proxy.

Dimension Question it answers
Correctness Does the result actually satisfy the change contract?
Regression safety Did it avoid breaking existing behavior/tests?
Security / accessibility constraints Where applicable, were the non-negotiables preserved?
Test adequacy Are the checks sufficient for the declared boundary?
Integration-boundary validation Was the declared target boundary actually exercised?
Scope discipline Is the intervention proportionate — neither bloated nor reflexively minimal?
Maintainability Can a human read, review, and extend it?
Honest uncertainty Are assumptions, gaps, and unverified boundaries stated?
Time / cost Only if measured; reported, never used alone to claim a win.

Full rubric with scoring anchors: ../eval/rubric.md.

Trajectory scoring. These dimensions apply to trajectory runs exactly as they apply to single-task runs. A trajectory is scored on its outcomes — correctness, regression safety, security/accessibility, test adequacy, integration-boundary validation, scope discipline, maintainability, and honest uncertainty — not on phase count, response length, or response shape. A trajectory that reaches the correct terminal state with proper gate evidence and skip transparency scores well regardless of how many phases it enumerated. Shape-neutrality (see ../eval/baseline-protocol.md) extends to trajectory comparisons: do not reward a run for visiting more phases or penalize it for a different recording format.

Task fixtures

Representative, repository-backed tasks across these classes:

  • bug diagnosis
  • feature change
  • refactor
  • specification ambiguity
  • regression prevention, including test-hardening guards for already-correct production behavior
  • review finding
  • release verification
  • "no change needed" cases

Plus adversarial / counterfactual cases where the correct answer is a larger change, a new dependency, a non-code process change, or no code change at all. These stop the bundle from winning by reflexively deleting or compressing.

Each fixture carries its repository context and harness constraints. Schema: ../eval/task-schema.md. Fixtures: ../eval/fixtures/.

Trajectory fixtures

In addition to single-task fixtures, the suite includes trajectory fixtures — multi-phase change-request journeys that record the full sequence of journey phases, gates, routing decisions, and terminal state. Where a single-task fixture has one class and expected_boundary, a trajectory fixture describes an end-to-end delivery (e.g. all nine phases to a released state, a lightweight path ending closed with recorded skips, or a lightweight test-hardening path with clean-baseline and targeted-mutant verification).

Trajectory fixtures are validated by the same runner using kind: trajectory as the discriminator. Their schema is documented in ../eval/task-schema.md.

Running the integrated suite

The runner (../eval/run_eval.py) validates both single-task and trajectory fixtures in one pass:

.venv/bin/python3 neckbeard/eval/run_eval.py \
  --suite neckbeard/eval/fixtures --validate-only

Expected output (current fixture set):

OK: 14 fixture(s) valid.
  11 single-task fixture(s) valid.
  3 trajectory fixture(s) valid.
  by class: {'adversarial': 2, 'bug-diagnosis': 1, 'feature-change': 1, 'no-change-needed': 1, 'refactor': 1, 'regression-prevention': 2, 'release-verification': 1, 'review-finding': 1, 'spec-ambiguity': 1}
  by visibility: {'public': 11}
  adversarial: 2
  trajectory paths: {'full': 1, 'lightweight': 2}

To scaffold a scoring report for manual evaluation:

.venv/bin/python3 neckbeard/eval/run_eval.py \
  --suite neckbeard/eval/fixtures \
  --report neckbeard/eval/out/report.md

Holdout discipline

Keep a task set separate from author iteration. Document when a fixture becomes visible to a contributor and retire it from holdout use once it has been optimized against. Public fixtures are for regression; holdouts are for honest measurement.

Fair baselines

Compare against a context-equivalent agent/harness. Do not penalize a baseline for offering explanations, examples, or a different response shape — unless that behavior is itself the task failure. The baseline must see the same repository context and constraints.

Multi-run, multi-model reporting

Report, for every result:

  • model and model version (where available)
  • harness / system prompt
  • tools available
  • fixture revision
  • randomization settings
  • run count and variance / confidence intervals

Never collapse one favorable point estimate into a universal claim.

Reproducible artifacts

Retain: prompts, fixtures, scoring rubric, commands, raw anonymized outputs (when licensing permits), and the aggregation script. Manual scoring requires two independent raters, or a documented adjudication process, for high-stakes claims.

Regression gate

A change to the bundle cannot claim improvement without running the public suite and reporting holdout results through the maintainers' controlled workflow.

Claims policy

Scope every performance claim to the evaluated models, harnesses, fixture revision (git SHA), repositories, task classes, and run dates. No result generalizes beyond the specific harness version, model, fixture set, and revision that produced it. Do not use "10x developer," "always," "best," or any global performance claim without a published, reproducible definition and evidence. Report template: ../templates/eval-report.md.