Files
magnus919_agent-skills/lifecycle-evals
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> b7aa174b84 fix(evals): re-point lifecycle corpus tooling at flat layout
Update the three hardcoded bundle manifest paths in run-corpus.sh and
validate-corpus-coverage.py from bundles/<name>/evals/evals.json to
<name>/evals/evals.json, refresh the coverage-index.json via --write-index,
and update the corpus prose (README, coverage-matrix, sources,
discovery-brief) to drop the bundles/ prefix.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-14 16:54:34 -04:00
..

Lifecycle Evaluation Corpus

A focused output-quality evaluation corpus for the milestone-4 product-to-production skills: the 14 new top-level skills (implementation-planning, product-analytics-and-measurement, product-roadmapping-and-portfolio, product-experimentation, product-adoption, conditional-customer-success, product-operations-and-governance, product-lifecycle-learning, production-readiness, migration-engineering, resilience-and-recovery, capacity-and-cost-engineering, incident-learning, privacy-engineering) and the 3 bundle umbrellas (product-lifecycle, production-excellence, agent-production-operations).

This directory is not a canonical skill: it deliberately has no SKILL.md, so the repository validators (which discover skills by SKILL.md) do not treat it as one. It is a corpus layer over the 17 per-skill eval manifests, plus run tooling, documentation, and committed reproducible run artifacts.

Why this exists

The 17 milestone-4 skills each ship evals/evals.json with at least five output-quality cases. This corpus layers the cross-skill requirements of issue #204 on top of those manifests:

  • every corpus case is an output-quality case with observable assertions in the canonical assertions field (never trigger-only, never the expectations alias);
  • the corpus as a whole covers five behavioral categories — ambiguity, conflicting evidence, unsafe authority, failure, and justified stop/retire — each with at least one case;
  • the three bundle manifests carry six integrated trajectory scenarios — product launch, failed experiment, migration with reconciliation failure, blocked production-readiness review, agent tool failure, and privacy-boundary escalation — each exercising real handoffs (evidence-ledger entries, production-evidence-packet fields, routing records, escalation records, trace-to-eval feedback records), not isolated wrapper output;
  • results are reproducible with the fake adapter, scoped, and honestly reported.

The machine-checkable map of which case covers which category/scenario is references/coverage-index.json, validated by scripts/validate-corpus-coverage.py; the human-readable version is references/coverage-matrix.md.

How to run the corpus

Prerequisite: the repository .venv (requirements-dev.txt installed). No credentials, API keys, or network access are required — the corpus runs with the fake adapter only.

Run all 17 manifests end-to-end:

bash lifecycle-evals/scripts/run-corpus.sh

This loops every manifest through eval_runner --adapter fake, writes per-trial manifests under ${CORPUS_OUT_DIR:-/tmp/lifecycle-evals-runs}/<skill>/manifests/, and exits 0 only when every trial completed with zero failures.

Run a single manifest:

.venv/bin/python -m eval_runner <skill>/evals/evals.json --adapter fake --output-dir /tmp/eval-smoke-<skill>

Validate the coverage index and category/scenario coverage:

.venv/bin/python lifecycle-evals/scripts/validate-corpus-coverage.py

After changing case content or tags, refresh the committed index with --write-index (see references/regression-detection.md).

Artifact layout

Path What it is
README.md This file: run instructions, scoping, status semantics, claims policy
references/coverage-matrix.md Human-readable matrix: every corpus case ID tagged with behavioral categories and integrated scenarios
references/coverage-index.json Machine-readable coverage index (generated by scripts/validate-corpus-coverage.py)
references/regression-detection.md Concrete regression-comparison procedure, case-ID stability rule, ratchet command, interpretation guidance
references/sources.md Fixture/source notes and provenance (corpus cases are self-contained; inputs inline in prompts)
references/discovery-brief.md Bounded discovery brief: surveyed surfaces, ownership boundaries, decisions
scripts/run-corpus.sh Fake-adapter loop over all 17 manifests, aggregate exit 0
scripts/validate-corpus-coverage.py Programmatic coverage validator (categories, scenarios, ID resolution, index currency)
run-artifacts/manifests/ One committed snapshot of fake-adapter run output (per-trial manifests), refreshed at merge time

The 17 eval manifests themselves live in their owning skills: <skill>/evals/evals.json for the 14 top-level skills and <bundle>/evals/evals.json for the 3 bundle umbrellas.

Per-case status semantics

Each eval_runner trial produces a per-trial manifest (see run-artifacts/manifests/*.json) whose status field is one of:

Status Meaning
completed The trial executed to completion under the adapter and was serialized. With the fake adapter this means the pipeline ran the case end-to-end with no runner error.
error / timeout / stopped The trial did not complete normally (execution error, timeout, or early stop) and counts as a failure for the run.

The runner reports done: N trial(s), 0 failure(s) and exits 0 when every trial is completed. A 0-failure run proves pipeline reproducibility and case executability: every case loads, runs through the harness, and serializes a scoped result. It is not a measure of model capability (see the claims policy below). The per-trial manifests also record prompt_hash and fixture_hashes so the exact case content under test is pinned.

Task-class scoping statement

All results produced by this corpus are scoped to:

  • Harness: the fake adapter v0.1.0 via eval_runner (.venv/bin/python -m eval_runner).
  • Model: none / unspecified for fake runs (model.provider and model.model_id are recorded per trial as unspecified unless --model is passed).
  • Task class: output-quality + integrated-trajectory evaluation of the milestone-4 product-to-production skills (the 14 product/production skills and 3 bundle umbrellas), covering the five behavioral categories and six integrated scenarios listed above.
  • Date: the started_at / finished_at timestamps recorded per trial.

No result is presented without this scope; per-trial manifests carry it mechanically.

Trigger-only prohibition

Trigger-only checks — "does the skill load when I mention X", "is frontmatter valid", "did the trigger match" — are not accepted as substitutes for output-quality evaluation anywhere in this corpus. Every case has at least one assertion verifiable from the produced output, artifact, decision outcome, or record content (a rejection, an escalation, a stop/retire decision, a routed artifact, a recorded evidence entry). Cases whose correct behavior is a negative outcome (reject / escalate / stop / retire / pause / block / decline) assert exactly that negative outcome.

Claims policy (non-claim statement)

This corpus is a small, fixed output-quality corpus (14 skills × ≥5 cases + 3 bundle umbrellas of integrated cases). Fake-adapter runs prove pipeline reproducibility and case executability only. They produce no pass-rate, accuracy, or capability claims about any model — no "10x", no "best", no universal performance claims. Any future real-adapter run (a model-backed harness) must be separately scoped, labeled (adapter + model + model version + date + task class), and reported under its own claims policy; results from such runs must never be conflated with the fake-adapter corpus results committed here.

Reference files