mirror of
https://github.com/magnus919/agent-skills.git
synced 2026-09-15 05:26:28 +03:00
1.9 KiB
1.9 KiB
Digital Twin Evaluation Plan
Intended use
- Represented original/process:
- Decision or action:
- Evaluation mode: descriptive / predictive / prescriptive / closed-loop
- Risk tier:
- Owner and stop authority:
- Truth/reference sources:
- Validity domain and exclusions:
Contracts
- Critical fields and maximum age:
- Quantities of interest and units:
- Tolerances:
- False-positive, false-negative, stale-answer, and overconfidence costs:
- Allowed tools/actions:
- Fallback and abstention behavior:
Evaluation planes
- Represented-system health
- Synchronization/data health
- Model credibility: verification
- Model credibility: validation
- Model credibility: uncertainty/calibration
- Twin platform health
- Agent/action quality
Scenarios and evidence
- Golden replay manifest:
- Held-out validation data:
- Boundary/rare/OOD slices:
- Adversarial/chaos cases:
- Shadow or canary plan:
- Independent oracle or reviewer:
Metrics and gates
| Metric | Target/tolerance | Slice/horizon | Evidence artifact | Owner |
|---|---|---|---|---|
Stop conditions
- Critical freshness breach
- Integrity/authenticity failure
- Missing provenance or version
- Unavailable/too-wide uncertainty
- Material predicted/observed divergence
- Unauthorized or unreconciled side effect
- Required audit telemetry unavailable
Decision
- Verdict: approve / conditional approve / hold / block
- Evidence versions:
- Expiry/review date:
- Decision owner:
- Dissent/open questions: