Files
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
9d6bddad61 test: add lifecycle evaluation corpus for new product and production skills (#232)
* test(evals): scope claims to harness model fixtures and revision

Append the neckbeard claims-scoping sentence to one representative
expected_output per per-skill manifest so every corpus member states
VAL-EVL-032 scope (harness, model, fixtures, revision under test).

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(product-lifecycle): upgrade integrated launch trajectory

Add an explicit launch-decision assertion to the new-product lifecycle
case so the integrated product-launch scenario terminates in a launch
decision recorded as a lifecycle evidence-ledger entry (VAL-CRP-010),
and scope its expected_output claims per VAL-EVL-032.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(production-excellence): add integrated migration reconciliation failure case

Add integrated-migration-reconciliation-failure: the production-excellence
gate model returns No-go on a reconciliation mismatch, records the failure
evidence, produces a rollback/roll-forward decision with an accountable
owner, and does not proceed to launch (VAL-CRP-012).

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(agent-production-operations): add privacy boundary escalation case

Add integrated-privacy-boundary-escalation (VAL-CRP-015): the runtime
control plan halts a cross-boundary EU PII trace export before any data
processing, names the privacy boundary, and escalates to jurisdiction-
specific legal review and a human operator. Also add a tool-authority-
health handoff assertion to the read-only contract case (VAL-CRP-016).

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(lifecycle-evals): add lifecycle evaluation corpus

Add the #204 corpus home: run tooling (run-corpus.sh, fake adapter only),
programmatic coverage validator (validate-corpus-coverage.py), machine-
readable coverage index + human-readable coverage matrix, regression-
detection and fixture/source notes, the bounded discovery brief, and a
one-snapshot committed set of fake-adapter per-trial run artifacts with
harness/model/date scoping fields.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 20:13:36 -04:00
..

Incident Learning

Convert operational incidents, near misses, and exercise findings into verified, owned improvements across product, engineering, test, evaluation, and governance — with closure evidence, not just tickets.

Why Install This Skill

Most teams create tickets after incidents. Few verify that the intended change actually happened and had the intended effect. This skill provides a structured method for converting raw incident evidence into durable follow-up work with verified closure — separating what was observed from what was inferred, mapping each finding to the right domain (product, code, tests, evals, operations, governance), and tracking every follow-up through to verified completion.

Install this skill when your agent needs to help teams move from "we filed tickets after the postmortem" to "we verified that the monitoring gap was closed, the regression test was added, and the eval case now catches the failure mode." It composes specialist capabilities from SRE, QA, verification, agent evaluation, product lifecycle learning, implementation planning, resilience-and-recovery, and production-readiness without duplicating their methodology, and it feeds learning records into the production-excellence and agent-production-operations bundles.

What You Get

Path What it provides
SKILL.md Trigger conditions, core principles, the incident learning record structure, loading guide, template index, routing table, and ownership boundaries
references/discovery-brief.md Survey of adjacent skills (SRE, QA, verification, agent evals, product lifecycle learning, implementation planning, resilience-and-recovery, production-readiness) with ownership boundaries and routing decisions
references/evidence-inference-taxonomy.md Full taxonomy for separating observed facts, causal hypotheses (with confidence levels), contributing conditions, and unresolved uncertainty
references/escaped-from-analysis.md Method for mapping incidents to originating gaps: escaped requirements, missing monitoring/observability, unsafe authority/access, migration gaps, adoption consequences
references/follow-up-domains.md Six-domain follow-up taxonomy (product, code, tests, evals, operations, governance) with ownership patterns and verification methods per domain
references/verification-and-closure.md Closure standard requiring implementation evidence, verification evidence, and effect evidence; explicit rejection of ticket-only closure
templates/incident-learning-record.md Structured record template with fields for observed facts, causal hypotheses, contributing conditions, unresolved uncertainty, and escaped-from mapping
templates/causal-evidence-ledger.md Ledger template for tracking each causal claim with supporting evidence, confidence level, and alternative explanations
templates/follow-up-work-map.md Six-domain follow-up work map with ownership, verification method, and status tracking per finding
templates/verification-and-closure-record.md Per-follow-up closure record requiring implementation evidence, verification evidence, and effect evidence
evals/evals.json Five output-quality eval cases covering: noisy incident report, monitoring gap, process failure, agent authority failure, and non-actionable follow-up rejection

Quick Start

Start with the incident-learning record template to structure the raw incident evidence — separating facts from hypotheses from uncertainty, and mapping the escaped-from gap. Then use the follow-up work map to assign each finding to a domain and owner. Track each follow-up through the verification and closure record.

Ask your agent to "convert this incident into a learning record" or "build a follow-up work map from this postmortem" and the skill's triggers will route the work.

Triggers

  • Converting operational incidents, near misses, or exercise findings into structured learning records with follow-up work.
  • Separating observed facts from causal hypotheses and unresolved uncertainty in incident analysis.
  • Mapping incident findings to follow-up work across product, code, tests, evals, operations, and governance domains.
  • Tracking incident follow-up work to verified closure with implementation, verification, and effect evidence.
  • Linking incidents to escaped requirements, missing monitoring, unsafe authority, migration gaps, or adoption consequences.
  • Auditing incident follow-up closure rates or detecting unverified closures.

Requirements

No runtime dependencies. The methodology is host-neutral and requires no specific tools, platforms, or API keys. The skill provides templates and method; incident data and follow-up work tracking must be supplied by the user's environment.