Files
magnus919_agent-skills/digital-twin/templates/evaluation-plan.md
2026-08-19 13:30:25 -04:00

1.9 KiB

Digital Twin Evaluation Plan

Intended use

  • Represented original/process:
  • Decision or action:
  • Evaluation mode: descriptive / predictive / prescriptive / closed-loop
  • Risk tier:
  • Owner and stop authority:
  • Truth/reference sources:
  • Validity domain and exclusions:

Contracts

  • Critical fields and maximum age:
  • Quantities of interest and units:
  • Tolerances:
  • False-positive, false-negative, stale-answer, and overconfidence costs:
  • Allowed tools/actions:
  • Fallback and abstention behavior:

Evaluation planes

  • Represented-system health
  • Synchronization/data health
  • Model credibility: verification
  • Model credibility: validation
  • Model credibility: uncertainty/calibration
  • Twin platform health
  • Agent/action quality

Scenarios and evidence

  • Golden replay manifest:
  • Held-out validation data:
  • Boundary/rare/OOD slices:
  • Adversarial/chaos cases:
  • Shadow or canary plan:
  • Independent oracle or reviewer:

Metrics and gates

Metric Target/tolerance Slice/horizon Evidence artifact Owner

Stop conditions

  • Critical freshness breach
  • Integrity/authenticity failure
  • Missing provenance or version
  • Unavailable/too-wide uncertainty
  • Material predicted/observed divergence
  • Unauthorized or unreconciled side effect
  • Required audit telemetry unavailable

Decision

  • Verdict: approve / conditional approve / hold / block
  • Evidence versions:
  • Expiry/review date:
  • Decision owner:
  • Dissent/open questions: