Files
magnus919_agent-skills/neckbeard/templates/eval-report.md
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> c0c7690724 feat(flatten): move bundle dirs to repo root
Move the 8 directories under bundles/ to the repo root via git mv and
remove the now-empty bundles/ directory. Replace the "bundles" entry in
pyproject.toml [tool.deptry] extend_exclude with the 8 moved dir names so
the moved trees stay excluded from Python dependency analysis.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-14 15:26:27 -04:00

2.1 KiB

Evaluation Report

Every performance claim must be scoped to the evaluated models, harnesses, repositories, task classes, and dates below. No "10x," "always," or "best" without a published, reproducible definition and evidence. LOC is diagnostic metadata only — never a success proxy.

Run identity

Field Value
Bundle revision
Fixture revision
Date(s)
Rater(s)

Models and harnesses compared

Arm Model + version Harness / system prompt Tools available Randomization Run count
neckbeard
baseline (context-equivalent)

Task classes exercised

  • Public:
  • Holdout:

Outcome scores

Dimension neckbeard (mean ± spread) baseline (mean ± spread)
Correctness
Regression safety
Security / accessibility constraints
Test adequacy
Integration-boundary validation
Scope discipline
Maintainability
Honest uncertainty
Time / cost (if measured)

Diagnostic metadata (not a success proxy)

Metric neckbeard baseline
LOC (diagnostic only)

Adversarial / counterfactual behavior

Scoped claim

Artifacts retained