* feat(bundles): add bundle manifest schema, manifests, and validation (#203) Introduce a machine-readable composition contract for canonical bundles: purpose, audience, stages, included skills, prerequisites, outputs, handoffs, conflicts, and eval suite (schemas/bundle-manifest-v1.schema.json, following the evals-v1 versioned-schema convention). Ship the bounded design note (docs/bundle-manifest-design.md), a schema-conformant example, canonical manifests for the three new milestone bundles, and a stdlib-only validator (scripts/validate-bundles.rb) that rejects incomplete, contradictory, and undeclared-overlapping manifests while keeping bundles an optional layer. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> * feat(bundles): add lifecycle capability matrix generator and validators (#203) Add scripts/gen-lifecycle-matrix.rb, which deterministically produces the human-readable docs/lifecycle-capability-matrix.md (one row per canonical bundle) and the machine-readable docs/lifecycle-capability-matrix.json (with per-cell source provenance) reusing the gen-*.rb conventions. Add scripts/validate-lifecycle-matrix.rb to check bundle coverage, cell traceability, artifact currency, and catalog-exactness of nested bundle helpers. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> * test(bundles): add bundle manifest validation tests (#203) Add scripts/test-validate-bundles.rb covering schema conformance of the committed example, valid-manifest and declared-conflict positives, per-field incomplete-manifest rejections, contradictory-manifest rejections (missing skill, undeclared handoff artifact, non-catalog conflict), undeclared-overlap rejection naming both manifests, and matrix generator/validator completeness and drift detection. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> * ci(bundles): wire bundle manifest validation into the gate (#203) Add validate-bundles.rb, test-validate-bundles.rb, the lifecycle matrix generator check, and the matrix validator to .github/workflows/validate.yml alongside the existing validator steps. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> --------- Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
11 KiB
name, description, license, compatibility, metadata
| name | description | license | compatibility | metadata | ||
|---|---|---|---|---|---|---|
| production-excellence | Route cross-domain production evidence (readiness, migration, recovery, capacity/cost, incident-learning) into a launch or operational decision — go, no-go, defer, exception, or escalation — with an accountable owner and a post-launch learning path. Compose production specialists without copying their runbooks. Do not use for incident command, release-pipeline mechanics, platform architecture, threat modeling, data-pipeline design, or any task owned end-to-end by a single specialist skill; do not use as a generic checklist detached from service ownership, risk, evidence, and verification. | MIT | Platform-agnostic methodology. No runtime dependencies, API keys, or external services required. |
|
Production Excellence
A thin composition bundle that assembles cross-domain production evidence into a defensible launch or operational decision. It owns the acceptance and handoff layer — the gate model that reads evidence from specialist skills and produces go / no-go / defer / exception / escalation outcomes with accountable owners. It does not own any specialist's runbook.
When to load this
Load when:
- A service or change is approaching a launch decision and evidence from multiple production domains must be assembled.
- You need a structured gate model (go/no-go/defer/exception/escalation) with conditions, evidence, and accountable owners.
- Cross-domain evidence (readiness, migration, recovery, capacity/cost, incident history) must be combined into one operational handoff record.
- A launch or change needs a post-launch learning path routed to incident-learning and product-lifecycle-learning.
- You are coordinating a production change across SRE, release, platform, security, data, and QA specialists and need a single acceptance contract.
When not to use
Do not load this bundle for:
- Incident command or SLO operations — those are owned by site-reliability-engineering.
- Release-pipeline mechanics, versioning, or deployment strategies — those are owned by release-engineering.
- Platform architecture or internal-developer-platform design — those are owned by platform-engineering.
- Threat modeling, security review procedure, or vulnerability assessment — those are owned by secure-software-engineering.
- Data-pipeline design, ETL, or storage architecture — those are owned by data-engineering.
- Test-strategy design, regression-suite management, or test-automation framework design — those are owned by qa-methodology.
- A generic checklist detached from service ownership, risk, evidence, and verification — every gate in this bundle requires a named service owner, assessed risk, verified evidence, and a declaration of the verification boundary. A bare checklist is never a valid outcome.
This bundle composes specialists. It never replaces them and never re-derives their methods. If the task is wholly within one specialist's domain, load that specialist directly.
Readiness routing table
The bundle routes each production concern to the specialist that owns it. The bundle itself owns only the acceptance and handoff layer — the cross-domain assembly and the gate decision.
Primary production-domain routes
| Domain | Specialist skill | What the specialist owns | What the bundle adds |
|---|---|---|---|
| Production readiness | production-readiness | Risk-scaled evidence packet (11 categories), go/no-go/defer/exception launch decisions, accountable owners | Cross-domain assembly with migration, recovery, capacity/cost, and incident evidence; gate integration |
| Migration | migration-engineering | Expand/contract, compatibility windows, dual-running, backfills, reconciliation, cutover, recovery paths | Migration evidence as input to the gate model; handoff of migration verification to the operational record |
| Resilience and recovery | resilience-and-recovery | Failure modes, degradation choices, RTO/RPO, restore testing, DR, game days, failover, data integrity | Recovery evidence as a gate condition; exercise results feed the handoff record |
| Capacity and cost | capacity-and-cost-engineering | Demand/capacity/scaling/utilization models, unit-cost connection to SLO decisions, cost-constrained scenarios | Capacity/cost evidence as a gate condition; SLO/cost tradeoff decisions feed the gate model |
| Incident learning | incident-learning | Observed facts, causal hypotheses, contributing conditions, follow-up work mapping, verified closure | Pre-existing incident evidence as a gate condition; post-launch incidents routed back to incident-learning |
Supporting specialist routes
| Domain | Specialist skill | When routed |
|---|---|---|
| Reliability / SLOs | site-reliability-engineering | SLO/error-budget status required for gate entry; incident response for post-launch issues |
| Release mechanics | release-engineering | Release plan, rollout/rollback strategy required for gate entry |
| Platform | platform-engineering | Service-catalog entry, paved-road status for new services |
| Security | secure-software-engineering | Security review evidence for trust-boundary changes |
| Data | data-engineering | Data-quality and pipeline evidence for data-path changes |
| QA | qa-methodology | Verification evidence for all launches |
| Verification | verification-methodology | Boundary labeling and gap declaration for evidence assessment |
Cross-domain entry evidence
Before the gate model runs, entry evidence must exist from every applicable domain. The bundle does not gather this evidence — it requires it. The complete evidence packet specification is in references/evidence-packet.md.
Summary:
- Every evidence domain (readiness, migration, recovery, capacity/cost, incident-learning) has a named source or an explicit gap with an owner and due date.
- The packet is usable for both new services and changes to existing systems — domains irrelevant to the change are explicitly marked "not applicable" with a reason.
- Missing evidence is never silently omitted. Every gap is recorded.
Gate and exception model
The gate model produces exactly one of five outcomes for every production change. Full definitions, conditions, and evidence requirements are in references/gates.md.
| Outcome | Meaning | Key condition |
|---|---|---|
| Go | Authorized to proceed to production | All required evidence domains are sourced; no blocking gaps |
| No-go | Blocked; must not proceed | A required domain has a blocking gap, or an irreversible step has no verified recovery path |
| Defer | Postponed with explicit conditions | A non-blocking gap or dependency has a committed resolution date; re-evaluation is scheduled |
| Exception | Proceeds under an explicit waiver | A human authority (not the agent, not the service owner alone) approves a time-bounded, risk-bounded exception |
| Escalation | Decision escalated to a higher body | Irreconcilable gate conflict, trust-boundary security gap, cross-team authority gap, or regulatory boundary |
Every outcome is anchored to service ownership, risk, evidence, and verification. No gate passes on a bare checklist. Each outcome names the accountable owner and records the evidence that supports it.
Operational handoff and post-launch learning
After a gate outcome is reached, the operational handoff record (references/handoff-record.md) is populated.
Post-launch learning paths
Launch outcomes and post-launch observations feed two learning routes:
-
Incident learning — post-launch incidents (SLO degradations, unexpected failures, capacity breaches) are routed to incident-learning. The handoff record provides the launch context; the incident-learning skill's verified-closure requirement ensures follow-up items are tracked to completion.
-
Lifecycle learning — expected outcomes recorded in the handoff (SLO targets, capacity assumptions, cost projections) are routed to product-lifecycle-learning for expected-vs-observed comparison at the handoff's review cadence. The lifecycle-learning skill's continue/improve/harvest/pivot/pause/retire decisions are informed by the gap between predicted and observed production behavior.
The handoff record is populated for every outcome — not only Go. No-go, Defer, Exception, and Escalation each produce a handoff record with the blocking condition, the follow-up path, and the accountable owner.
Loading and nested-skill behavior
This bundle is the discoverable entry point. It does not contain nested
sub-skills under a skills/ directory. All routed skills are top-level catalog
skills referenced via relative markdown links. Harnesses that support progressive
disclosure will discover this bundle through its SKILL.md frontmatter and load
the referenced specialists on trigger.
See AGENTS.md for agent-specific loading notes.
File map
| Path | Loaded when |
|---|---|
| references/discovery-brief.md | Understanding the bundle's boundary against existing production and release skills |
| references/evidence-packet.md | Assembling cross-domain evidence for a production decision |
| references/gates.md | Running the gate model — go/no-go/defer/exception/escalation |
| references/handoff-record.md | Producing the operational handoff record and post-launch learning path |
| manifest.yaml | Machine-readable composition contract (schema v1): purpose, audience, stages, included skills, prerequisites, outputs, handoffs, conflicts, and eval suite; consumed by the lifecycle capability matrix |