Files
magnus919_agent-skills/lifecycle-evals/scripts/run-corpus.sh
T
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
9d6bddad61 test: add lifecycle evaluation corpus for new product and production skills (#232)
* test(evals): scope claims to harness model fixtures and revision

Append the neckbeard claims-scoping sentence to one representative
expected_output per per-skill manifest so every corpus member states
VAL-EVL-032 scope (harness, model, fixtures, revision under test).

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(product-lifecycle): upgrade integrated launch trajectory

Add an explicit launch-decision assertion to the new-product lifecycle
case so the integrated product-launch scenario terminates in a launch
decision recorded as a lifecycle evidence-ledger entry (VAL-CRP-010),
and scope its expected_output claims per VAL-EVL-032.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(production-excellence): add integrated migration reconciliation failure case

Add integrated-migration-reconciliation-failure: the production-excellence
gate model returns No-go on a reconciliation mismatch, records the failure
evidence, produces a rollback/roll-forward decision with an accountable
owner, and does not proceed to launch (VAL-CRP-012).

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(agent-production-operations): add privacy boundary escalation case

Add integrated-privacy-boundary-escalation (VAL-CRP-015): the runtime
control plan halts a cross-boundary EU PII trace export before any data
processing, names the privacy boundary, and escalates to jurisdiction-
specific legal review and a human operator. Also add a tool-authority-
health handoff assertion to the read-only contract case (VAL-CRP-016).

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(lifecycle-evals): add lifecycle evaluation corpus

Add the #204 corpus home: run tooling (run-corpus.sh, fake adapter only),
programmatic coverage validator (validate-corpus-coverage.py), machine-
readable coverage index + human-readable coverage matrix, regression-
detection and fixture/source notes, the bounded discovery brief, and a
one-snapshot committed set of fake-adapter per-trial run artifacts with
harness/model/date scoping fields.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 20:13:36 -04:00

70 lines
2.7 KiB
Bash

#!/usr/bin/env bash
# Run the lifecycle evaluation corpus end-to-end with the fake adapter.
#
# Loops every corpus manifest (14 per-skill + 3 bundle umbrellas = 17) through
# eval_runner with --adapter fake and aggregates the exit status. A trial that
# reports a failure (non-"completed" status) or a runner error counts as a
# corpus failure. The script exits 0 only when every trial in every manifest
# completed with zero failures.
#
# No API keys, model credentials, or network access are required: the fake
# adapter is fully deterministic (see eval_runner/fake_adapter.py).
#
# Usage (from the repository root):
# bash lifecycle-evals/scripts/run-corpus.sh
#
# Output artifacts are written under ${CORPUS_OUT_DIR} (default: /tmp/lifecycle-evals-runs),
# one subdirectory per skill, with per-trial manifests under
# <out>/<skill>/manifests/*.manifest.json. The committed one-snapshot artifact
# copy lives in lifecycle-evals/run-artifacts/manifests/ and is refreshed only
# at merge time (VAL-CRP-021 churn policy: do NOT gate CI on artifact freshness).
set -euo pipefail
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
PYTHON="$ROOT/.venv/bin/python"
OUT_DIR="${CORPUS_OUT_DIR:-/tmp/lifecycle-evals-runs}"
if [[ ! -x "$PYTHON" ]]; then
echo "error: $PYTHON not found (is the repository .venv set up?)" >&2
exit 2
fi
MANIFESTS=(
"implementation-planning/evals/evals.json"
"product-analytics-and-measurement/evals/evals.json"
"product-roadmapping-and-portfolio/evals/evals.json"
"product-experimentation/evals/evals.json"
"product-adoption/evals/evals.json"
"conditional-customer-success/evals/evals.json"
"product-operations-and-governance/evals/evals.json"
"product-lifecycle-learning/evals/evals.json"
"production-readiness/evals/evals.json"
"migration-engineering/evals/evals.json"
"resilience-and-recovery/evals/evals.json"
"capacity-and-cost-engineering/evals/evals.json"
"incident-learning/evals/evals.json"
"privacy-engineering/evals/evals.json"
"bundles/product-lifecycle/evals/evals.json"
"bundles/production-excellence/evals/evals.json"
"bundles/agent-production-operations/evals/evals.json"
)
total_failures=0
for manifest in "${MANIFESTS[@]}"; do
skill_dir="$(dirname "$(dirname "$manifest")")"
echo "==> $manifest"
if ! "$PYTHON" -m eval_runner "$ROOT/$manifest" --adapter fake \
--output-dir "$OUT_DIR/$skill_dir"; then
echo "FAIL: $manifest" >&2
total_failures=$((total_failures + 1))
fi
done
if [[ "$total_failures" -gt 0 ]]; then
echo "corpus run FAILED: $total_failures manifest(s) reported failures" >&2
exit 1
fi
echo "corpus run OK: ${#MANIFESTS[@]} manifests, 0 failures (fake adapter, no credentials)"
exit 0