Update the three hardcoded bundle manifest paths in run-corpus.sh and validate-corpus-coverage.py from bundles/<name>/evals/evals.json to <name>/evals/evals.json, refresh the coverage-index.json via --write-index, and update the corpus prose (README, coverage-matrix, sources, discovery-brief) to drop the bundles/ prefix. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Lifecycle Evaluation Corpus
A focused output-quality evaluation corpus for the milestone-4 product-to-production skills: the 14 new top-level skills (implementation-planning, product-analytics-and-measurement, product-roadmapping-and-portfolio, product-experimentation, product-adoption, conditional-customer-success, product-operations-and-governance, product-lifecycle-learning, production-readiness, migration-engineering, resilience-and-recovery, capacity-and-cost-engineering, incident-learning, privacy-engineering) and the 3 bundle umbrellas (product-lifecycle, production-excellence, agent-production-operations).
This directory is not a canonical skill: it deliberately has no SKILL.md, so the
repository validators (which discover skills by SKILL.md) do not treat it as one. It is a
corpus layer over the 17 per-skill eval manifests, plus run tooling, documentation, and
committed reproducible run artifacts.
Why this exists
The 17 milestone-4 skills each ship evals/evals.json with at least five output-quality
cases. This corpus layers the cross-skill requirements of issue #204 on top of those
manifests:
- every corpus case is an output-quality case with observable assertions in the canonical
assertionsfield (never trigger-only, never theexpectationsalias); - the corpus as a whole covers five behavioral categories — ambiguity, conflicting evidence, unsafe authority, failure, and justified stop/retire — each with at least one case;
- the three bundle manifests carry six integrated trajectory scenarios — product launch, failed experiment, migration with reconciliation failure, blocked production-readiness review, agent tool failure, and privacy-boundary escalation — each exercising real handoffs (evidence-ledger entries, production-evidence-packet fields, routing records, escalation records, trace-to-eval feedback records), not isolated wrapper output;
- results are reproducible with the fake adapter, scoped, and honestly reported.
The machine-checkable map of which case covers which category/scenario is
references/coverage-index.json, validated by
scripts/validate-corpus-coverage.py; the
human-readable version is references/coverage-matrix.md.
How to run the corpus
Prerequisite: the repository .venv (requirements-dev.txt installed). No credentials, API
keys, or network access are required — the corpus runs with the fake adapter only.
Run all 17 manifests end-to-end:
bash lifecycle-evals/scripts/run-corpus.sh
This loops every manifest through eval_runner --adapter fake, writes per-trial manifests
under ${CORPUS_OUT_DIR:-/tmp/lifecycle-evals-runs}/<skill>/manifests/, and exits 0 only
when every trial completed with zero failures.
Run a single manifest:
.venv/bin/python -m eval_runner <skill>/evals/evals.json --adapter fake --output-dir /tmp/eval-smoke-<skill>
Validate the coverage index and category/scenario coverage:
.venv/bin/python lifecycle-evals/scripts/validate-corpus-coverage.py
After changing case content or tags, refresh the committed index with
--write-index (see references/regression-detection.md).
Artifact layout
| Path | What it is |
|---|---|
README.md |
This file: run instructions, scoping, status semantics, claims policy |
references/coverage-matrix.md |
Human-readable matrix: every corpus case ID tagged with behavioral categories and integrated scenarios |
references/coverage-index.json |
Machine-readable coverage index (generated by scripts/validate-corpus-coverage.py) |
references/regression-detection.md |
Concrete regression-comparison procedure, case-ID stability rule, ratchet command, interpretation guidance |
references/sources.md |
Fixture/source notes and provenance (corpus cases are self-contained; inputs inline in prompts) |
references/discovery-brief.md |
Bounded discovery brief: surveyed surfaces, ownership boundaries, decisions |
scripts/run-corpus.sh |
Fake-adapter loop over all 17 manifests, aggregate exit 0 |
scripts/validate-corpus-coverage.py |
Programmatic coverage validator (categories, scenarios, ID resolution, index currency) |
run-artifacts/manifests/ |
One committed snapshot of fake-adapter run output (per-trial manifests), refreshed at merge time |
The 17 eval manifests themselves live in their owning skills:
<skill>/evals/evals.json for the 14 top-level skills and <bundle>/evals/evals.json
for the 3 bundle umbrellas.
Per-case status semantics
Each eval_runner trial produces a per-trial manifest (see
run-artifacts/manifests/*.json) whose status field is one of:
| Status | Meaning |
|---|---|
completed |
The trial executed to completion under the adapter and was serialized. With the fake adapter this means the pipeline ran the case end-to-end with no runner error. |
error / timeout / stopped |
The trial did not complete normally (execution error, timeout, or early stop) and counts as a failure for the run. |
The runner reports done: N trial(s), 0 failure(s) and exits 0 when every trial is
completed. A 0-failure run proves pipeline reproducibility and case executability:
every case loads, runs through the harness, and serializes a scoped result. It is not a
measure of model capability (see the claims policy below). The per-trial manifests also
record prompt_hash and fixture_hashes so the exact case content under test is pinned.
Task-class scoping statement
All results produced by this corpus are scoped to:
- Harness: the
fakeadapter v0.1.0 viaeval_runner(.venv/bin/python -m eval_runner). - Model: none /
unspecifiedfor fake runs (model.providerandmodel.model_idare recorded per trial asunspecifiedunless--modelis passed). - Task class: output-quality + integrated-trajectory evaluation of the milestone-4 product-to-production skills (the 14 product/production skills and 3 bundle umbrellas), covering the five behavioral categories and six integrated scenarios listed above.
- Date: the
started_at/finished_attimestamps recorded per trial.
No result is presented without this scope; per-trial manifests carry it mechanically.
Trigger-only prohibition
Trigger-only checks — "does the skill load when I mention X", "is frontmatter valid", "did the trigger match" — are not accepted as substitutes for output-quality evaluation anywhere in this corpus. Every case has at least one assertion verifiable from the produced output, artifact, decision outcome, or record content (a rejection, an escalation, a stop/retire decision, a routed artifact, a recorded evidence entry). Cases whose correct behavior is a negative outcome (reject / escalate / stop / retire / pause / block / decline) assert exactly that negative outcome.
Claims policy (non-claim statement)
This corpus is a small, fixed output-quality corpus (14 skills × ≥5 cases + 3 bundle umbrellas of integrated cases). Fake-adapter runs prove pipeline reproducibility and case executability only. They produce no pass-rate, accuracy, or capability claims about any model — no "10x", no "best", no universal performance claims. Any future real-adapter run (a model-backed harness) must be separately scoped, labeled (adapter + model + model version + date + task class), and reported under its own claims policy; results from such runs must never be conflated with the fake-adapter corpus results committed here.
Reference files
references/coverage-matrix.md— which case covers which category/scenario.references/regression-detection.md— how to detect and interpret regressions across revisions.references/sources.md— fixture/source notes and provenance.references/discovery-brief.md— the bounded discovery brief for this corpus.