Files
magnus919_agent-skills/migration-engineering/references/recovery-classification.md
Magnus HedemarkGitHubusername <username>factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
652521a09e feat(migration-engineering): add migration-engineering skill (#222)
* feat(migration-engineering): add migration-engineering skill

Add the migration-engineering skill for safe cross-system migrations:
schema, data, API, infrastructure, and service migrations.

- SKILL.md: expand/contract pattern, compatibility windows, dual-running,
  backfills, reconciliation, cutover, deprecation, and cleanup. Four distinct
  recovery paths (rollback, roll-forward, restore, irreversible). Structured
  planning fields for reconciliation, correctness evidence, observability,
  customer impact, and ownership. Four migration types with detailed
  compatibility/correctness/recovery characteristics. Specialist routing
  to api-design-and-evolution, data-engineering, platform-engineering,
  release-engineering, site-reliability-engineering, implementation-planning,
  secure-software-engineering, qa-methodology, and verification-methodology.
  Prose routing to production-readiness and production-excellence.
- README.md: human-facing overview with all five required sections.
- references/discovery-brief.md: bounded survey of migration-adjacent skills
  and clear ownership boundaries.
- references/compatibility-patterns.md: forward/backward compatibility by type.
- references/recovery-classification.md: four recovery paths with decision tree.
- templates/: migration plan, compatibility matrix, reconciliation plan,
  cutover and recovery record.
- evals/evals.json: 5 output-quality cases covering additive schema change,
  backfill with reconciliation, API version migration, irreversible cutover,
  and reconciliation failure.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* chore(migration-engineering): update catalogs and routing

Regenerate catalog files and add migration-engineering entries to
root README.md catalog and references/skill-triggers.md.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 18:16:04 -04:00

5.1 KiB

Recovery Classification

Deep reference on the four recovery paths with decision rules, examples, and anti-patterns. Load this reference when classifying recovery for concrete migration steps.

The four recovery paths

Rollback

Reverse to the prior state by undoing the change.

Decision rule: Rollback is possible when the old state still exists and can be reactivated, OR the change is purely additive and can be removed without side effects.

Examples:

  • Undo a feature-flag-controlled code path by turning off the flag.
  • Restore read traffic to the old database by reverting the connection string.
  • Drop a newly added nullable column that is not yet consumed.
  • Cancel a deprecation notice and keep the old API endpoint live.

Anti-patterns:

  • Claiming rollback is possible because "we can fix it in the next deploy." That is roll-forward, not rollback.
  • Claiming rollback is possible when the old system has been decommissioned. That is restore or irreversible.

Roll-forward

Fix forward in the new state. The old state is no longer reachable, but a fix can be deployed to the new system.

Decision rule: Roll-forward is the correct path when (a) reversal is impossible or more expensive than fixing forward, AND (b) a fix can be deployed within the acceptable recovery window.

Examples:

  • A data migration cutover has completed and the old store is read-only; a bug in the new service is discovered. Deploy a fix to the new service.
  • An API migration has removed the old endpoint; a consumer reports a regression. Fix the new endpoint.
  • A configuration error in the new infrastructure is causing errors. Correct the configuration and redeploy.

Anti-patterns:

  • Using roll-forward as the default recovery path without assessing whether rollback is simpler and safer.
  • Failing to define the acceptable fix-forward window (how long can the system be degraded before the fix lands?).

Restore

Restore the prior state from a backup or snapshot.

Decision rule: Restore is the path when the old system is no longer operational but a backup exists and a restore procedure is tested and has a known recovery time.

Examples:

  • A schema migration dropped the wrong table; restore from the pre-migration backup.
  • A data migration corrupted the target store; restore the target from the pre-migration snapshot and re-run the migration.
  • An infrastructure migration destroyed the old environment; restore from the infrastructure-as-code state and redeploy.

Anti-patterns:

  • Assuming restore is possible because "we have backups." A backup that has not been tested with a restore drill is not a recovery path — it is a hope.
  • Failing to define the Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for the restore.

Irreversible

Reversal is impossible. The change cannot be undone at any level.

Decision rule: Irreversible when (a) the old state is physically destroyed, (b) the operation is one-way by design (e.g., cryptographic erasure), or (c) a third-party action cannot be recalled.

Examples:

  • Physical hardware decommissioning where the device is shipped back and wiped.
  • Cryptographic key rotation where old keys are destroyed after rotation.
  • Third-party data export where the receiving party cannot be compelled to delete the data.
  • Permanent data deletion to satisfy a regulatory requirement (e.g., GDPR right-to-erasure).

Required for every irreversible step:

  1. Acceptance criteria — what conditions must be met before the irreversible step is executed (e.g., "reconciliation passed at 100% for 7 consecutive days").
  2. Stakeholder communication — who must be informed and who must approve before the step executes.
  3. Contingency plan — what happens if the irreversible step succeeds but the overall migration subsequently fails (e.g., "rebuild from source of truth," "accept data loss within defined scope").

Anti-patterns:

  • Treating an irreversible step as if it were reversible — stating "rollback: N/A" without the acceptance, communication, and contingency requirements.
  • Claiming "irreversible" for a step that is merely expensive or inconvenient to reverse. Irreversible means physically or logically impossible, not merely costly.

Classification decision tree

Can the old state be reactivated without data loss?
  ├── YES → Rollback is possible
  └── NO:
      ├── Can the new state be fixed within the acceptable recovery window?
      │   └── YES → Roll-forward is possible
      └── Can the old state be restored from backup?
          ├── YES, and restore procedure is tested → Restore is possible
          └── NO → Irreversible

When not to claim rollback

Never claim rollback is possible when:

  • The old system has been decommissioned and cannot be restarted.
  • The old data has been deleted and no backup exists.
  • The old API has been removed and cannot be redeployed.
  • A third-party action cannot be reversed.
  • The rollback procedure has never been tested.

In these cases, classify as roll-forward, restore, or irreversible — not as rollback with caveats.