Files
magnus919_agent-skills/ai-operating-economics/references/source-index.md
Magnus HedemarkandGitHub 94b7231147 feat: add AI operating economics skill (#368)
* feat: add AI operating economics skill

Add an evidence-led methodology for evaluating AI workflow value, cost, worker effects, quality guardrails, and authority expansion. Includes research references, durable decision templates, and six eval cases. AI assistance: Jasper, on behalf of Magnus Hedemark.

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix: resolve AI economics review findings

Align section numbering, evidence-language examples, and intervention-mode terminology identified by the exact-head review.

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

---------

Signed-off-by: Magnus Hedemark <magnus919@pm.me>
2026-08-21 13:53:49 -04:00

9.5 KiB

Source Index and Evidence Boundaries

This index records the research basis for the skill. Access dates and versions should be refreshed when a decision depends on a time-sensitive provider capability. These sources inform method and guardrails; they do not establish universal ROI.

Workflow outcome and measurement

Source Evidence type Supports Does not support
Brynjolfsson, Li, and Raymond, Generative AI at Work Independent working paper, revised 2023 In one customer-support deployment, AI assistance increased resolved issues per hour by about 14% on average, with much larger gains for novice/lower-skilled workers and minimal gains for experienced/high-skilled workers; the paper also examines quality, sentiment, retention, adherence, and learning General enterprise ROI, current agentic-system performance, or universal productivity claims
NBER digest summary Independent study summary Plain-language description of the measured deployment and heterogeneous effects A substitute for the working paper when methodological detail matters
METR, Uplift Update Independent measurement research Selection effects, task-selection bias, concurrent-agent accounting, and why an attractive developer-speed estimate may not support a strong causal claim A general estimate of enterprise AI productivity
METR, AI Usage Survey Independent survey research The distinction between speed and value, plus caveats on self-reported uplift and counterfactual estimation Audited financial ROI or causal productivity evidence
Noy and Zhang, Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence Randomized experiment manuscript Short writing-task effects on completion time, blinded quality, performer heterogeneity, task composition, and the limits of inferring durable workplace value Firm-level ROI, long-term skill development, or generalization to organization-specific work
OECD Employment Outlook 2023: AI, job quality and inclusiveness Independent policy research Worker, management, working-condition, skill, productivity, wage, employment, and transition dimensions that should accompany a narrow productivity measure A universal prediction of AI's labor-market impact or a substitute for local worker evidence
OECD case studies of AI implementation Independent case research Heterogeneity across worker profiles, sectors, countries, task composition, skill requirements, and job quality A causal benchmark for a specific organization

Cost and usage

Source Evidence type Supports Does not support
FinOps for AI overview Primary foundation guidance Extending FinOps practices to volatile model pricing, token meters, GPU scarcity, allocation, quotas, tagging, and outcome alignment Proof that any organization has realized savings
How to Build a Generative AI Cost and Usage Tracker Primary foundation guidance Token attribution levels, centralized or common interfaces, shared-throughput allocation, and the need to account for more than inference A universal architecture or exact savings formula
GenAI FinOps: How Token Pricing Really Works Primary foundation guidance The warning that advertised token prices do not describe complete application TCO A measured cross-provider cost comparison
Token Economics: The Atomic Unit of AI Value Primary foundation insight Tokens are computation units and require contextual interpretation Tokens as a direct measure of business value

Benefits realization and cost estimation

Source Evidence type Supports Does not support
UK Digital and Data Benefits Framework Government guidance Distinguishing benefits, disbenefits, measures, owners, baselines, dependencies, and realization tracking in a digital business case A universal accounting treatment or proof that a forecast benefit will be realized
UK Magenta Book evaluation guidance Government evaluation guidance Evaluation planning, theory of change, counterfactual reasoning, monitoring, and proportionate evidence design A replacement for domain-specific statistical or financial expertise
FinOps Open Cost and Usage Specification (FOCUS) Primary specification A common structure for normalizing billing data and supporting reconciliation, allocation, chargeback, budgeting, and forecasting Complete TCO, causal ROI, labor cost, or the correct local allocation policy
GAO Cost Estimating and Assessment Guide Government cost-estimation guidance Lifecycle cost categories, documented assumptions, uncertainty, sensitivity, independent review, and updating estimates as evidence changes An AI-specific cost model or a guarantee of estimate accuracy

Governance and worker impact

Source Evidence type Supports Does not support
NIST AI RMF: Generative AI Profile Primary voluntary framework Generative-AI-specific risk considerations, testing and evaluation, monitoring, incident response, human oversight, and lifecycle controls Certification, compliance, safety proof, or realized value
ISO/IEC 42001 International standard description The AI management-system idea: objectives, responsibilities, performance evaluation, corrective action, and continual improvement Certification or conformance without an authorized audit and the applicable standard
ILO, Generative AI and Jobs Independent policy research Job quantity and quality, autonomy, work organization, and worker voice as dimensions beyond productivity A prediction of the impact on a particular employer or occupation

Telemetry and controls

Source Evidence type Supports Does not support
OpenTelemetry GenAI observability Primary technical guidance Agent, model-call, tool-execution, model/version, duration, and token telemetry; prompt and tool content should not be captured by default when sensitive Complete production observability or a guarantee of privacy
OpenTelemetry GenAI attributes Primary technical specification Portable provider, model, operation, token, and agent attributes, with version/movement caveats Stable interoperability across every implementation without version pinning
NIST AI RMF Core Primary voluntary framework Govern, map, measure, and manage lifecycle structure; inventory, roles, monitoring, incident response, recovery, and deactivation expectations Certification, safety proof, or legal compliance
NIST Generative AI Profile Primary voluntary framework Generative-AI-specific risk-management considerations to supplement AI RMF 1.0 A complete implementation design or outcome guarantee
AWS Bedrock prompt routing Vendor documentation An example of provider-side quality/cost routing and its documented limitations Independent savings or quality superiority
Microsoft Foundry cost management Vendor documentation An example of estimation, representative test traffic, cost grouping, budgets, alerts, and dependent-resource TCO Hard spend stops, universal controls, or independent ROI

Research-use rules

  • Preserve the source URL, access date, evidence type, scope, and caveat with every retained claim.
  • Treat working papers as research evidence, not peer-reviewed consensus unless the source says otherwise.
  • Treat vendor documentation as capability evidence only.
  • Treat vendor surveys and case studies as reported claims, not causal outcomes.
  • Do not copy copyrighted source text into public skill content; paraphrase and link.
  • Re-verify provider capabilities and moving OpenTelemetry conventions before using them as implementation requirements.