# Evaluation Report > Every performance claim must be scoped to the evaluated models, harnesses, > repositories, task classes, and dates below. No "10x," "always," or "best" > without a published, reproducible definition and evidence. LOC is diagnostic > metadata only — never a success proxy. ## Run identity | Field | Value | |---|---| | Bundle revision | | | Fixture revision | | | Date(s) | | | Rater(s) | | ## Models and harnesses compared | Arm | Model + version | Harness / system prompt | Tools available | Randomization | Run count | |---|---|---|---|---|---| | neckbeard | | | | | | | baseline (context-equivalent) | | | | | | ## Task classes exercised - Public: - Holdout: ## Outcome scores | Dimension | neckbeard (mean ± spread) | baseline (mean ± spread) | |---|---|---| | Correctness | | | | Regression safety | | | | Security / accessibility constraints | | | | Test adequacy | | | | Integration-boundary validation | | | | Scope discipline | | | | Maintainability | | | | Honest uncertainty | | | | Time / cost (if measured) | | | ## Diagnostic metadata (not a success proxy) | Metric | neckbeard | baseline | |---|---|---| | LOC (diagnostic only) | | | ## Adversarial / counterfactual behavior ## Scoped claim ## Artifacts retained