Files
magnus919_agent-skills/ai-operating-economics/templates/ai-initiative-evidence-record.md
T
Magnus HedemarkandGitHub 94b7231147 feat: add AI operating economics skill (#368)
* feat: add AI operating economics skill

Add an evidence-led methodology for evaluating AI workflow value, cost, worker effects, quality guardrails, and authority expansion. Includes research references, durable decision templates, and six eval cases. AI assistance: Jasper, on behalf of Magnus Hedemark.

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix: resolve AI economics review findings

Align section numbering, evidence-language examples, and intervention-mode terminology identified by the exact-head review.

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

---------

Signed-off-by: Magnus Hedemark <magnus919@pm.me>
2026-08-21 13:53:49 -04:00

132 lines
3.4 KiB
Markdown

# AI Initiative Evidence Record
## Decision header
- Initiative:
- Workflow:
- Population and scope:
- Intervention mode: assist / recommend / route / execute / replace
- Decision sought: scale / constrain / redesign / hold / retire / exception
- Accountable decision owner:
- Review date or trigger:
- Record version:
## Value hypothesis
> For [population] doing [workflow], [intervention] will change [outcome] by [direction/range] without exceeding [countermetric boundary], at [full cost boundary], compared with [baseline], over [period].
- Hypothesis status: supported / weakened / refuted / unresolved
- Expected benefit:
- Enabled capacity:
- Operational benefit:
- Realized economic or mission benefit:
- Benefit realization mechanism:
- Human-control boundary:
- Authority proposed for next slice:
## Evidence comparison
- Baseline:
- Comparison design:
- Treatment period:
- Comparison period:
- Inclusion and exclusion rules:
- Known selection effects:
- Concurrent changes:
- Quality measurement limitations:
- Statistical or causal analysis owner:
| Claim | Class | Source | Result | Caveat | Permitted interpretation |
|---|---|---|---|---|---|
| | observed / causal / inferred / vendor-reported / asserted / normative | | | | |
## Outcome and countermetrics
| Metric | Type | Definition and denominator | Baseline | Observed | Target/boundary | Owner | Evidence source |
|---|---|---|---:|---:|---|---|---|
| | primary / leading / countermetric / adoption | | | | | | |
## Segment review
| Slice | Adoption or exposure | Outcome | Quality/countermetric | New burden or benefit | Decision implication |
|---|---:|---:|---:|---|---|
| | | | | | |
Slices not available and why:
## Cost boundary
### Billing truth
- Provider/infrastructure source:
- Billing period and version:
- Reconciliation status:
### Allocated cost
- Allocation target:
- Shared-cost rule:
- Allocation owner:
### Economic cost
- Meaningful unit:
- Model/inference:
- Tools and APIs:
- Retrieval/storage/networking:
- Human review and exception handling:
- Incremental capacity:
- Engineering, evaluation, observability, governance, training, support, and exit:
- Fixed, variable, step-function, avoided, transferred, and uncertain costs:
- Calculation and allocation method:
- Range or sensitivity:
## Findings and gaps
### Supported findings
1.
### Unresolved or conflicting findings
1.
### Missing evidence
| Gap | Why it matters | Owner | Next evidence | Due date or trigger |
|---|---|---|---|---|
| | | | | |
## Governance evidence packet
- Intended use and risk tier:
- System/model/prompt/policy/tool/provider/version inventory:
- Acceptable-use, refusal, escalation, and human-oversight rules:
- Pre-deployment evaluation and release threshold:
- Third-party/provider assessment and contractual evidence:
- Incident, override, and near-miss record:
- Change/revalidation trigger:
- Retention, dependency, leakage, user-impact, and decommissioning plan:
## Decision and controls
- Disposition:
- Scope of approval:
- Authority limit:
- Budget or quota limit:
- Human review or escalation rule:
- Stop trigger:
- Rollback, containment, or retirement path:
- Exception approver, if applicable:
- Revisit condition:
## Learning closure
- Expected versus observed outcome:
- Expected versus observed cost:
- Countermetric and subgroup result:
- Incidents, overrides, or near misses:
- Changes since prior record:
- Hypothesis update:
- Follow-up artifact or owner: