mirror of
https://github.com/magnus919/agent-skills.git
synced 2026-09-17 06:26:31 +03:00
* feat: add AI operating economics skill Add an evidence-led methodology for evaluating AI workflow value, cost, worker effects, quality guardrails, and authority expansion. Includes research references, durable decision templates, and six eval cases. AI assistance: Jasper, on behalf of Magnus Hedemark. Signed-off-by: Magnus Hedemark <magnus919@pm.me> * fix: resolve AI economics review findings Align section numbering, evidence-language examples, and intervention-mode terminology identified by the exact-head review. Signed-off-by: Magnus Hedemark <magnus919@pm.me> --------- Signed-off-by: Magnus Hedemark <magnus919@pm.me>
9.5 KiB
9.5 KiB
Source Index and Evidence Boundaries
This index records the research basis for the skill. Access dates and versions should be refreshed when a decision depends on a time-sensitive provider capability. These sources inform method and guardrails; they do not establish universal ROI.
Workflow outcome and measurement
| Source | Evidence type | Supports | Does not support |
|---|---|---|---|
| Brynjolfsson, Li, and Raymond, Generative AI at Work | Independent working paper, revised 2023 | In one customer-support deployment, AI assistance increased resolved issues per hour by about 14% on average, with much larger gains for novice/lower-skilled workers and minimal gains for experienced/high-skilled workers; the paper also examines quality, sentiment, retention, adherence, and learning | General enterprise ROI, current agentic-system performance, or universal productivity claims |
| NBER digest summary | Independent study summary | Plain-language description of the measured deployment and heterogeneous effects | A substitute for the working paper when methodological detail matters |
| METR, Uplift Update | Independent measurement research | Selection effects, task-selection bias, concurrent-agent accounting, and why an attractive developer-speed estimate may not support a strong causal claim | A general estimate of enterprise AI productivity |
| METR, AI Usage Survey | Independent survey research | The distinction between speed and value, plus caveats on self-reported uplift and counterfactual estimation | Audited financial ROI or causal productivity evidence |
| Noy and Zhang, Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence | Randomized experiment manuscript | Short writing-task effects on completion time, blinded quality, performer heterogeneity, task composition, and the limits of inferring durable workplace value | Firm-level ROI, long-term skill development, or generalization to organization-specific work |
| OECD Employment Outlook 2023: AI, job quality and inclusiveness | Independent policy research | Worker, management, working-condition, skill, productivity, wage, employment, and transition dimensions that should accompany a narrow productivity measure | A universal prediction of AI's labor-market impact or a substitute for local worker evidence |
| OECD case studies of AI implementation | Independent case research | Heterogeneity across worker profiles, sectors, countries, task composition, skill requirements, and job quality | A causal benchmark for a specific organization |
Cost and usage
| Source | Evidence type | Supports | Does not support |
|---|---|---|---|
| FinOps for AI overview | Primary foundation guidance | Extending FinOps practices to volatile model pricing, token meters, GPU scarcity, allocation, quotas, tagging, and outcome alignment | Proof that any organization has realized savings |
| How to Build a Generative AI Cost and Usage Tracker | Primary foundation guidance | Token attribution levels, centralized or common interfaces, shared-throughput allocation, and the need to account for more than inference | A universal architecture or exact savings formula |
| GenAI FinOps: How Token Pricing Really Works | Primary foundation guidance | The warning that advertised token prices do not describe complete application TCO | A measured cross-provider cost comparison |
| Token Economics: The Atomic Unit of AI Value | Primary foundation insight | Tokens are computation units and require contextual interpretation | Tokens as a direct measure of business value |
Benefits realization and cost estimation
| Source | Evidence type | Supports | Does not support |
|---|---|---|---|
| UK Digital and Data Benefits Framework | Government guidance | Distinguishing benefits, disbenefits, measures, owners, baselines, dependencies, and realization tracking in a digital business case | A universal accounting treatment or proof that a forecast benefit will be realized |
| UK Magenta Book evaluation guidance | Government evaluation guidance | Evaluation planning, theory of change, counterfactual reasoning, monitoring, and proportionate evidence design | A replacement for domain-specific statistical or financial expertise |
| FinOps Open Cost and Usage Specification (FOCUS) | Primary specification | A common structure for normalizing billing data and supporting reconciliation, allocation, chargeback, budgeting, and forecasting | Complete TCO, causal ROI, labor cost, or the correct local allocation policy |
| GAO Cost Estimating and Assessment Guide | Government cost-estimation guidance | Lifecycle cost categories, documented assumptions, uncertainty, sensitivity, independent review, and updating estimates as evidence changes | An AI-specific cost model or a guarantee of estimate accuracy |
Governance and worker impact
| Source | Evidence type | Supports | Does not support |
|---|---|---|---|
| NIST AI RMF: Generative AI Profile | Primary voluntary framework | Generative-AI-specific risk considerations, testing and evaluation, monitoring, incident response, human oversight, and lifecycle controls | Certification, compliance, safety proof, or realized value |
| ISO/IEC 42001 | International standard description | The AI management-system idea: objectives, responsibilities, performance evaluation, corrective action, and continual improvement | Certification or conformance without an authorized audit and the applicable standard |
| ILO, Generative AI and Jobs | Independent policy research | Job quantity and quality, autonomy, work organization, and worker voice as dimensions beyond productivity | A prediction of the impact on a particular employer or occupation |
Telemetry and controls
| Source | Evidence type | Supports | Does not support |
|---|---|---|---|
| OpenTelemetry GenAI observability | Primary technical guidance | Agent, model-call, tool-execution, model/version, duration, and token telemetry; prompt and tool content should not be captured by default when sensitive | Complete production observability or a guarantee of privacy |
| OpenTelemetry GenAI attributes | Primary technical specification | Portable provider, model, operation, token, and agent attributes, with version/movement caveats | Stable interoperability across every implementation without version pinning |
| NIST AI RMF Core | Primary voluntary framework | Govern, map, measure, and manage lifecycle structure; inventory, roles, monitoring, incident response, recovery, and deactivation expectations | Certification, safety proof, or legal compliance |
| NIST Generative AI Profile | Primary voluntary framework | Generative-AI-specific risk-management considerations to supplement AI RMF 1.0 | A complete implementation design or outcome guarantee |
| AWS Bedrock prompt routing | Vendor documentation | An example of provider-side quality/cost routing and its documented limitations | Independent savings or quality superiority |
| Microsoft Foundry cost management | Vendor documentation | An example of estimation, representative test traffic, cost grouping, budgets, alerts, and dependent-resource TCO | Hard spend stops, universal controls, or independent ROI |
Research-use rules
- Preserve the source URL, access date, evidence type, scope, and caveat with every retained claim.
- Treat working papers as research evidence, not peer-reviewed consensus unless the source says otherwise.
- Treat vendor documentation as capability evidence only.
- Treat vendor surveys and case studies as reported claims, not causal outcomes.
- Do not copy copyrighted source text into public skill content; paraphrase and link.
- Re-verify provider capabilities and moving OpenTelemetry conventions before using them as implementation requirements.