* feat(migration-engineering): add migration-engineering skill Add the migration-engineering skill for safe cross-system migrations: schema, data, API, infrastructure, and service migrations. - SKILL.md: expand/contract pattern, compatibility windows, dual-running, backfills, reconciliation, cutover, deprecation, and cleanup. Four distinct recovery paths (rollback, roll-forward, restore, irreversible). Structured planning fields for reconciliation, correctness evidence, observability, customer impact, and ownership. Four migration types with detailed compatibility/correctness/recovery characteristics. Specialist routing to api-design-and-evolution, data-engineering, platform-engineering, release-engineering, site-reliability-engineering, implementation-planning, secure-software-engineering, qa-methodology, and verification-methodology. Prose routing to production-readiness and production-excellence. - README.md: human-facing overview with all five required sections. - references/discovery-brief.md: bounded survey of migration-adjacent skills and clear ownership boundaries. - references/compatibility-patterns.md: forward/backward compatibility by type. - references/recovery-classification.md: four recovery paths with decision tree. - templates/: migration plan, compatibility matrix, reconciliation plan, cutover and recovery record. - evals/evals.json: 5 output-quality cases covering additive schema change, backfill with reconciliation, API version migration, irreversible cutover, and reconciliation failure. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> * chore(migration-engineering): update catalogs and routing Regenerate catalog files and add migration-engineering entries to root README.md catalog and references/skill-triggers.md. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> --------- Co-authored-by: username <username> Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
5.1 KiB
Recovery Classification
Deep reference on the four recovery paths with decision rules, examples, and anti-patterns. Load this reference when classifying recovery for concrete migration steps.
The four recovery paths
Rollback
Reverse to the prior state by undoing the change.
Decision rule: Rollback is possible when the old state still exists and can be reactivated, OR the change is purely additive and can be removed without side effects.
Examples:
- Undo a feature-flag-controlled code path by turning off the flag.
- Restore read traffic to the old database by reverting the connection string.
- Drop a newly added nullable column that is not yet consumed.
- Cancel a deprecation notice and keep the old API endpoint live.
Anti-patterns:
- Claiming rollback is possible because "we can fix it in the next deploy." That is roll-forward, not rollback.
- Claiming rollback is possible when the old system has been decommissioned. That is restore or irreversible.
Roll-forward
Fix forward in the new state. The old state is no longer reachable, but a fix can be deployed to the new system.
Decision rule: Roll-forward is the correct path when (a) reversal is impossible or more expensive than fixing forward, AND (b) a fix can be deployed within the acceptable recovery window.
Examples:
- A data migration cutover has completed and the old store is read-only; a bug in the new service is discovered. Deploy a fix to the new service.
- An API migration has removed the old endpoint; a consumer reports a regression. Fix the new endpoint.
- A configuration error in the new infrastructure is causing errors. Correct the configuration and redeploy.
Anti-patterns:
- Using roll-forward as the default recovery path without assessing whether rollback is simpler and safer.
- Failing to define the acceptable fix-forward window (how long can the system be degraded before the fix lands?).
Restore
Restore the prior state from a backup or snapshot.
Decision rule: Restore is the path when the old system is no longer operational but a backup exists and a restore procedure is tested and has a known recovery time.
Examples:
- A schema migration dropped the wrong table; restore from the pre-migration backup.
- A data migration corrupted the target store; restore the target from the pre-migration snapshot and re-run the migration.
- An infrastructure migration destroyed the old environment; restore from the infrastructure-as-code state and redeploy.
Anti-patterns:
- Assuming restore is possible because "we have backups." A backup that has not been tested with a restore drill is not a recovery path — it is a hope.
- Failing to define the Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for the restore.
Irreversible
Reversal is impossible. The change cannot be undone at any level.
Decision rule: Irreversible when (a) the old state is physically destroyed, (b) the operation is one-way by design (e.g., cryptographic erasure), or (c) a third-party action cannot be recalled.
Examples:
- Physical hardware decommissioning where the device is shipped back and wiped.
- Cryptographic key rotation where old keys are destroyed after rotation.
- Third-party data export where the receiving party cannot be compelled to delete the data.
- Permanent data deletion to satisfy a regulatory requirement (e.g., GDPR right-to-erasure).
Required for every irreversible step:
- Acceptance criteria — what conditions must be met before the irreversible step is executed (e.g., "reconciliation passed at 100% for 7 consecutive days").
- Stakeholder communication — who must be informed and who must approve before the step executes.
- Contingency plan — what happens if the irreversible step succeeds but the overall migration subsequently fails (e.g., "rebuild from source of truth," "accept data loss within defined scope").
Anti-patterns:
- Treating an irreversible step as if it were reversible — stating "rollback: N/A" without the acceptance, communication, and contingency requirements.
- Claiming "irreversible" for a step that is merely expensive or inconvenient to reverse. Irreversible means physically or logically impossible, not merely costly.
Classification decision tree
Can the old state be reactivated without data loss?
├── YES → Rollback is possible
└── NO:
├── Can the new state be fixed within the acceptable recovery window?
│ └── YES → Roll-forward is possible
└── Can the old state be restored from backup?
├── YES, and restore procedure is tested → Restore is possible
└── NO → Irreversible
When not to claim rollback
Never claim rollback is possible when:
- The old system has been decommissioned and cannot be restarted.
- The old data has been deleted and no backup exists.
- The old API has been removed and cannot be redeployed.
- A third-party action cannot be reversed.
- The rollback procedure has never been tested.
In these cases, classify as roll-forward, restore, or irreversible — not as rollback with caveats.