diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index c383fcd..88c77f7 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -165,7 +165,7 @@ "./capacity-and-cost-engineering" ], "strict": false, - "description": "Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. Use when projecting capacity from growth forecasts, sizing for peak events, designing cost-aware scaling policies, defining budget thresholds or quota/rate-limit enforcement, running or planning load/soak tests as capacity evidence, or resolving SLO-cost tradeoffs. Do NOT use for financial P&L statements, fundraising scenarios, or SaaS metrics (route to financial-modeling); for infrastructure implementation or cloud-resource provisioning (route to platform-engineering); or for generic cloud-cost tips and universal utilization targets — this skill does not prescribe fixed savings rates or one-size-fits-all thresholds." + "description": "Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. Use when projecting capacity from growth forecasts, sizing for peak events, designing cost-aware scaling policies, defining budget thresholds or quota/rate-limit enforcement, running or planning load/soak tests as capacity evidence, resolving SLO-cost tradeoffs, or modeling multi-tenant demand distributions, hot-tenant skew, pooled or siloed headroom, fairness evidence, and tenant-variable unit cost. Do NOT use for financial P&L statements, fundraising scenarios, or SaaS metrics (route to financial-modeling); for infrastructure implementation or cloud-resource provisioning (route to platform-engineering); or for generic cloud-cost tips and universal utilization targets — this skill does not prescribe fixed savings rates or one-size-fits-all thresholds." }, { "name": "chief-of-staff-methodology", diff --git a/README.md b/README.md index f84d281..289c150 100644 --- a/README.md +++ b/README.md @@ -78,7 +78,7 @@ Make system boundaries, responsibilities, and relationships legible at the archi ### [capacity-and-cost-engineering](capacity-and-cost-engineering/SKILL.md) -Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. Covers growth forecasting, peak sizing, degraded-mode capacity, budget/quota controls, load/soak test evidence standards, and SLO-cost tradeoff records. Routes financial P&L and fundraising to financial-modeling, infrastructure implementation to platform-engineering, degradation-path design to resilience-and-recovery, and feeds capacity/cost evidence to production-readiness. Ships a discovery brief comparing seven adjacent skills, five fillable templates (capacity model, unit-economics record, budget/quota decision, load/soak test plan, SLO-cost tradeoff record), and five evals. +Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. Covers growth forecasting, peak sizing, degraded-mode capacity, multi-tenant demand distributions and skew, pooled/siloed headroom, tier-linked admission and fairness evidence, tenant-variable unit cost, budget/quota controls, representative load/soak evidence, and SLO-cost tradeoff records. Routes SaaS architecture, security, financial outcomes, infrastructure implementation, SRE, analytics, and general architecture to their specialist owners. Ships focused references, six fillable templates, and ten evals. ### [chief-of-staff-methodology](chief-of-staff-methodology/SKILL.md) diff --git a/capacity-and-cost-engineering/README.md b/capacity-and-cost-engineering/README.md index d3f9979..49fb6c3 100644 --- a/capacity-and-cost-engineering/README.md +++ b/capacity-and-cost-engineering/README.md @@ -6,21 +6,24 @@ Connect demand, performance, reliability, and spend into defensible capacity and Every service that serves users has a capacity limit and a cost. When your agent can model capacity, calculate unit cost, define budget controls, and require load-test evidence for capacity claims, it stops treating infrastructure as "someone else's problem" and starts making decisions that respect real-world constraints. This skill fills the gap between the financial team's P&L models (which don't know what a request costs in compute) and the platform team's infrastructure-as-code (which doesn't know why a specific SLO target was chosen or what it costs). -After installing this skill, your agent can: project capacity from growth forecasts with utilization targets and scaling triggers; calculate what one request or one user costs to serve at the infrastructure level; define budget thresholds with operational consequences (alert, throttle, deny); design load and soak tests as mandatory capacity evidence — not optional nice-to-haves; and resolve SLO-cost tradeoffs with explicit evidence, ownership, and accountability. When a product manager asks "what would it cost to serve 2x the users?", your agent has a structured answer instead of a guess. +After installing this skill, your agent can: project capacity from growth forecasts with utilization targets and scaling triggers; model multi-tenant demand distributions, hot tenants, partition skew, and pooled versus siloed headroom; connect tier promises to quotas, admission, and contextual fairness evidence; calculate platform baseline, tenant-variable, and allocated unit cost; define budget thresholds with operational consequences (alert, throttle, deny); design representative per-tenant load and soak tests as mandatory capacity evidence; and resolve SLO-cost tradeoffs with explicit evidence, ownership, and accountability. ## What You Get | Directory | What it provides | |-----------|-----------------| -| `SKILL.md` | Core methodology: connected dimensions (demand/performance/reliability/spend), working method with five steps, four named scenarios (growth, peak, degraded, cost-constrained), routing table to adjacent skills, and guardrails against generic cloud-cost tips and universal utilization targets | +| `SKILL.md` | Core methodology: connected dimensions (demand/performance/reliability/spend), multi-tenant loading route, working method with five steps, four named scenarios, routing table to adjacent skills, and guardrails against generic cloud-cost tips and universal thresholds | | `README.md` | This human-facing overview | -| `references/discovery-brief.md` | Ownership boundary analysis comparing seven adjacent skills (financial-modeling, platform-engineering, site-reliability-engineering, product-analytics-and-measurement, production-readiness, product-roadmapping-and-portfolio, resilience-and-recovery) with explicit routing decisions | +| `references/discovery-brief.md` | Ownership boundary analysis for financial, platform, SRE, analytics, SaaS architecture, security, software architecture, readiness, roadmap, and recovery concerns | +| `references/multi-tenant-capacity-and-unit-cost.md` | Original method for tenant distributions, hot tenants, partition skew, pooled/siloed headroom, tier promises, admission, fairness evidence, shared-cost allocation, and tenant-variable unit cost | +| `references/source-index.md` | Public provenance and transformation boundary | | `templates/capacity-model.md` | Fillable capacity model: demand assumptions, capacity-unit mapping, utilization targets with rationale, scaling triggers, evidence sources, ownership, and tradeoffs | | `templates/unit-economics-record.md` | Fillable unit-economics record: unit definition, cost numerator with allocation method, demand denominator, unit-cost calculation formula, cost-per-SLO comparison, and structured assumptions/evidence/ownership/tradeoffs fields | | `templates/load-soak-test-plan.md` | Fillable load/soak test plan: objective, target throughput, duration, environment requirements, success criteria (latency percentiles, error rate, utilization), data collection, and evidence record | | `templates/budget-quota-decision.md` | Fillable budget/quota decision: budget owner, period, thresholds (alert/soft/hard), quota/rate-limit configuration, enforcement mechanism, operational behavior at each threshold, cost attribution, and approval | | `templates/slo-cost-tradeoff-record.md` | Fillable SLO-cost tradeoff record: SLO under discussion, current and projected cost, alternative SLO comparison, degradation path, error budget impact, accountable owner, and approval | -| `evals/evals.json` | Five output-quality evaluation cases covering growth forecast, peak event, SLO-cost conflict, quota decision, and misleading unit-cost calculation | +| `templates/tenant-capacity-model.md` | Fillable tenant capacity and unit-cost model covering profiles, skew, pooled/siloed comparison, quotas, fairness, cost allocation, and representative per-tenant load/soak evidence | +| `evals/evals.json` | Ten output-quality evaluation cases covering general capacity/cost decisions plus tenant distributions, pooled/siloed headroom, fairness, tenant-variable cost, and routing boundaries | ## Quick Start @@ -39,6 +42,7 @@ Load this skill when the task involves: - Planning or reviewing a load or soak test as capacity evidence - Resolving an SLO-cost tradeoff or cost-constrained reliability decision - Reviewing a cost anomaly or attributing cost to services/teams +- Modeling multi-tenant demand distributions, hot tenants, partition skew, quotas, fairness, pooled/siloed headroom, or tenant-variable unit cost ## Requirements diff --git a/capacity-and-cost-engineering/SKILL.md b/capacity-and-cost-engineering/SKILL.md index 21393b6..e273311 100644 --- a/capacity-and-cost-engineering/SKILL.md +++ b/capacity-and-cost-engineering/SKILL.md @@ -5,18 +5,21 @@ description: >- demand, performance, and reliability decisions. Use when projecting capacity from growth forecasts, sizing for peak events, designing cost-aware scaling policies, defining budget thresholds or quota/rate-limit enforcement, running - or planning load/soak tests as capacity evidence, or resolving SLO-cost - tradeoffs. Do NOT use for financial P&L statements, fundraising scenarios, or - SaaS metrics (route to financial-modeling); for infrastructure implementation - or cloud-resource provisioning (route to platform-engineering); or for - generic cloud-cost tips and universal utilization targets — this skill does - not prescribe fixed savings rates or one-size-fits-all thresholds. + or planning load/soak tests as capacity evidence, resolving SLO-cost tradeoffs, + or modeling multi-tenant demand distributions, hot-tenant skew, pooled or siloed + headroom, fairness evidence, and tenant-variable unit cost. Do NOT use for + financial P&L statements, fundraising scenarios, or SaaS metrics (route to + financial-modeling); for infrastructure implementation or cloud-resource + provisioning (route to platform-engineering); or for generic cloud-cost tips + and universal utilization targets — this skill does not prescribe fixed savings + rates or one-size-fits-all thresholds. license: MIT compatibility: Platform-agnostic methodology. No runtime dependency. metadata: tags: capacity-engineering, cost-engineering, unit-economics, load-testing, capacity-planning, budget-controls, quota-management, rate-limiting, - cost-performance-tradeoffs, scaling-models, utilization-modeling + cost-performance-tradeoffs, scaling-models, utilization-modeling, + multi-tenant-capacity, tenant-unit-cost --- # Capacity and Cost Engineering @@ -50,6 +53,7 @@ Load this skill when the task involves any of: | Plan or review a load/soak test as capacity evidence | `SKILL.md` + `templates/load-soak-test-plan.md` | | Resolve an SLO-cost tradeoff or cost-constrained reliability decision | `SKILL.md` + `templates/slo-cost-tradeoff-record.md` | | Review a cost anomaly or attribute cost to services/teams | `SKILL.md` + `templates/unit-economics-record.md` | +| Model multi-tenant demand, skew, fairness, or tenant-variable unit cost | `SKILL.md` + `references/multi-tenant-capacity-and-unit-cost.md` + `templates/tenant-capacity-model.md` | | Understand ownership boundaries with adjacent skills | `SKILL.md` + `references/discovery-brief.md` | ## Working method @@ -80,6 +84,14 @@ Both numerator and denominator must be measured over the same period, with the s The **unit-economics record template** (`templates/unit-economics-record.md`) requires: the unit definition, the cost numerator with allocation method, the demand denominator with measurement source, the resulting unit cost, a cost-per-SLO comparison (what happens to unit cost at 99.9% vs 99.99%?), and an assumptions/evidence/ownership/tradeoffs section. The template includes a structured field for the unit-cost calculation formula. +For a multi-tenant service, do not use a fleet average as the only unit. Load +`references/multi-tenant-capacity-and-unit-cost.md` and distinguish platform +baseline cost from tenant-variable cost, then report a distribution of tenant +costs or resource consumption. A tenant's variable unit cost may depend on +request mix, storage, background work, burst shape, placement, and tier +entitlements. Shared-cost allocation is an explicit modeling choice, not a +claim that every tenant consumes an equal share. + ### 3. Define budget and quota controls Budget controls are spending limits with operational consequences — spending alerts at thresholds, hard caps that prevent further spend, and the operational behavior when a cap is hit (degrade, throttle, or stop). Quota and rate-limit enforcement are the mechanisms that implement budget controls at the request or resource level. @@ -107,6 +119,12 @@ The **load/soak test plan template** (`templates/load-soak-test-plan.md`) captur Modeling without test evidence, or testing without a model, is incomplete. Both are required. +For multi-tenant claims, the evidence must exercise representative tenant +profiles together, including ordinary tenants, high-demand tenants, bursty +tenants, and relevant tier or placement variants. A single-tenant benchmark or +fleet-average test cannot establish protection against hot tenants, partition +skew, or fairness behavior. + ### 5. Resolve SLO-cost tradeoffs An SLO-cost tradeoff arises when the cost of meeting an SLO at projected demand exceeds the budget, or when a budget constraint forces a lower SLO than the team would otherwise target. This is a structured decision, not an implicit acceptance. @@ -182,6 +200,9 @@ This skill does **not** permit cost optimization to justify degrading reliabilit | Degradation-path design, recovery verification, RTO/RPO decisions, game days | [resilience-and-recovery](../resilience-and-recovery/SKILL.md) | | Statistical modeling of demand, time-series forecasting, causal inference on growth drivers | [data-scientist](../data-scientist/SKILL.md) | | Cost data pipeline implementation, spend-data ETL, cost-dashboard data models | [data-engineering](../data-engineering/SKILL.md) | +| End-to-end tenant semantics, control/application planes, tenancy choice, lifecycle, or billing handoffs | [multi-tenant-saas-architecture](../multi-tenant-saas-architecture/SKILL.md) | +| Tenant isolation threats, authorization, privileged support, or security controls | [secure-software-engineering](../secure-software-engineering/SKILL.md) | +| General architecture boundaries, topology, or decomposition decisions | [software-architecture](../software-architecture/SKILL.md) | ## File map @@ -193,3 +214,6 @@ This skill does **not** permit cost optimization to justify degrading reliabilit | [templates/budget-quota-decision.md](templates/budget-quota-decision.md) | Defining budget thresholds, quota limits, rate-limit enforcement, and operational consequences | | [templates/load-soak-test-plan.md](templates/load-soak-test-plan.md) | Designing or reviewing a load or soak test as capacity evidence | | [templates/slo-cost-tradeoff-record.md](templates/slo-cost-tradeoff-record.md) | Resolving an SLO-cost tradeoff with evidence, accountability, and approval | +| [references/multi-tenant-capacity-and-unit-cost.md](references/multi-tenant-capacity-and-unit-cost.md) | Modeling tenant distributions, skew, pooled/siloed headroom, tier promises, admission, fairness evidence, and cost allocation | +| [templates/tenant-capacity-model.md](templates/tenant-capacity-model.md) | Recording tenant profiles, partition behavior, headroom, quota/admission evidence, and tenant-variable unit cost | +| [references/source-index.md](references/source-index.md) | Public provenance and original-writing boundary for this methodology | diff --git a/capacity-and-cost-engineering/evals/evals.json b/capacity-and-cost-engineering/evals/evals.json index 7ea75dd..9b12c98 100644 --- a/capacity-and-cost-engineering/evals/evals.json +++ b/capacity-and-cost-engineering/evals/evals.json @@ -74,6 +74,70 @@ "The response explicitly states that cost optimization must not justify degrading reliability, privacy, or user outcomes, and that if budget is exceeded the tradeoff is escalated", "The corrected calculation includes an owner and states what evidence is needed to validate it" ] + }, + { + "id": "tenant-demand-distribution", + "prompt": "A shared document-processing SaaS serves 4,000 tenants. The average tenant submits 2 jobs per minute, but 2% of tenants generate 45% of jobs and a few tenants concentrate work on one partition. We need a six-month capacity model and guidance on what evidence is required before promising a new enterprise tier.", + "expected_output": "A tenant-aware capacity model that preserves the demand distribution instead of using the fleet average. It identifies hot tenants and partition skew separately, describes tenant profiles and burst shape, maps the profiles to the saturating resource, and compares ordinary and hot-tenant scenarios. It states that the enterprise promise must be translated into measurable capacity behavior, with a contextual quota/admission and fairness policy. It requires representative concurrent per-tenant load and soak evidence, including hot and skewed profiles, and labels unmeasured assumptions and owners. It routes tenant semantics and tier architecture to multi-tenant-saas-architecture, isolation controls to secure-software-engineering, SLO ownership to site-reliability-engineering, and implementation to platform-engineering.", + "assertions": [ + "The output models tenant demand as a distribution and does not rely on the average tenant", + "Hot-tenant behavior and partition skew are analyzed as distinct capacity risks", + "The enterprise promise is translated into measurable capacity behavior with contextual quota or admission decisions", + "Representative concurrent per-tenant load and soak evidence is required before the capacity promise is accepted", + "Missing data is labeled as an assumption or evidence gap with an owner", + "Neighboring SaaS, security, SRE, and platform ownership boundaries are explicit" + ] + }, + { + "id": "pooled-versus-siloed-headroom", + "prompt": "We are deciding whether to keep a large tenant in a pooled worker fleet, give it a dedicated silo, or use a hybrid placement. The pooled fleet has lower idle cost but shows queue spikes when the tenant bursts. The silo adds fixed capacity and another failure-domain obligation. Produce an evidence-based comparison without prescribing a universal utilization or headroom percentage.", + "expected_output": "A comparison of pooled, siloed, and hybrid choices across ordinary load, burst, tenant growth or movement, and failure-domain loss. It separates baseline, variable demand, shared contention, idle capacity, redundancy, and operational overhead. It explains what headroom protects, the provisioning lead time, observed saturation behavior, and the cost of unused capacity rather than using a universal percentage. It defines the load and soak scenarios needed to compare user outcomes, queueing, recovery, and cost, with a named decision owner and unresolved assumptions.", + "assertions": [ + "Pooled, siloed, and hybrid options are compared across burst and failure-domain scenarios", + "Baseline, variable demand, idle capacity, redundancy, and operational overhead are separated", + "Headroom is justified by workload, saturation evidence, and provisioning lead time rather than a universal percentage", + "The evidence plan includes representative tenant load and soak behavior and cost inputs", + "The decision has an owner and explicit assumptions or gaps" + ] + }, + { + "id": "tenant-admission-fairness", + "prompt": "One premium tenant is allowed large bursts on a shared search service. During those bursts, standard tenants see elevated P99 latency and queue depth. Define quotas, admission control, and fairness evidence while preserving the premium contract and avoiding a universal fairness ratio.", + "expected_output": "A contextual policy that identifies the search resource and affected tenants, translates each tier promise into steady and burst behavior, and chooses an explicit admission response such as weighted scheduling, reserved capacity, bounded queues, or throttling. It states what happens to the constrained tenant and protected tenants, how retry amplification is prevented, and how recovery is verified. It does not claim a universal fairness ratio or threshold; each limit is justified by the contract, workload, SLO impact, and observed evidence. It requires a tenant-distributed load/soak test and routes security controls and infrastructure configuration to their owners.", + "assertions": [ + "The policy identifies the shared resource, affected tenants, and tier-specific steady and burst promises", + "Quota and admission behavior is operationally specific and includes the response for both constrained and protected tenants", + "Retry amplification, queue bounds, and recovery verification are addressed", + "Fairness is contextual and no universal fairness ratio or threshold is prescribed", + "Tenant-distributed load and soak evidence is required", + "Security-control and infrastructure-implementation boundaries are preserved" + ] + }, + { + "id": "tenant-variable-unit-cost", + "prompt": "Our multi-tenant analytics service costs $90,000 per month. $35,000 is the minimum shared platform baseline, $25,000 is storage and transfer that can be measured by tenant, and $30,000 is pooled compute. Enterprise tenants run expensive scheduled queries while self-serve tenants mostly run small interactive queries. Design a tenant-variable unit-cost model that does not pretend every tenant consumes an equal share.", + "expected_output": "A unit-cost model that separates platform baseline and measured tenant-variable storage/transfer from pooled compute; defines a transparent allocation method for the pooled compute; segments interactive and scheduled query demand; and reports tenant or profile-level direct, allocated, and marginal cost with aligned periods and denominators. It explains how idle capacity, shared control-plane work, redundancy, and fixed commitments are treated. It routes pricing and margin decisions to financial-modeling and implementation of cost tags or resource policy to platform-engineering. It names evidence needed to validate attribution, including representative tenant load/soak data, and does not infer commercial profitability from technical cost alone.", + "assertions": [ + "The response separates the $35,000 baseline from tenant-variable storage and transfer costs, with an explicit allocation method for the $30,000 pooled-compute pool", + "Pooled compute allocation is explicit and does not assume equal tenant consumption", + "Interactive and scheduled query units are segmented rather than averaged together", + "Periods and demand denominators align and direct, allocated, and marginal views are distinguished", + "Fixed, idle, shared, and redundancy costs are treated explicitly", + "Financial outcomes and infrastructure implementation are routed to their specialist owners", + "Evidence needed to validate tenant attribution is named" + ] + }, + { + "id": "tenant-boundary-routing", + "prompt": "A product team asks for an end-to-end multi-tenant SaaS architecture, including tenant identity, control-plane boundaries, data partitioning, entitlements, billing, isolation, capacity, and cost. Explain which parts capacity-and-cost-engineering should own and which should be handed to neighboring skills.", + "expected_output": "A routing response that keeps capacity-and-cost-engineering focused on measured tenant demand distributions, hot tenants, partition skew, pooled or siloed headroom, tier-linked quotas and admission, fairness evidence, shared-cost allocation, tenant-variable unit cost, and representative load/soak evidence. It routes tenant semantics, control/application planes, partitioning architecture, lifecycle, entitlements, and billing handoffs to multi-tenant-saas-architecture; threat controls to secure-software-engineering; pricing and financial outcomes to financial-modeling; infrastructure implementation to platform-engineering; SLOs and live operations to site-reliability-engineering; demand instrumentation to product-analytics-and-measurement; and general architecture decisions to software-architecture. It does not duplicate those methods.", + "assertions": [ + "Capacity-and-cost ownership is limited to tenant capacity and cost evidence rather than end-to-end SaaS architecture", + "The response routes tenant semantics, planes, lifecycle, entitlements, and billing to multi-tenant-saas-architecture", + "Security, financial, platform, SRE, analytics, and general architecture boundaries are all named", + "Representative per-tenant load and soak evidence remains owned by capacity-and-cost-engineering", + "The response avoids duplicating neighboring methodologies" + ] } ] } diff --git a/capacity-and-cost-engineering/references/discovery-brief.md b/capacity-and-cost-engineering/references/discovery-brief.md index 4471889..45e7d58 100644 --- a/capacity-and-cost-engineering/references/discovery-brief.md +++ b/capacity-and-cost-engineering/references/discovery-brief.md @@ -75,3 +75,21 @@ This brief surveys adjacent skills in the agent-skills catalog to define the own ## Summary Capacity-and-cost-engineering fills a gap between financial modeling (which owns the business view of cost), platform engineering (which implements infrastructure), SRE (which owns reliability targets), product analytics (which owns demand signals), production-readiness (which consumes capacity/cost evidence), product roadmapping (which allocates portfolio capacity), and resilience-and-recovery (which owns degradation design). It is the method for modeling technical capacity, calculating unit cost, defining budget and quota controls, requiring load/soak evidence, and making cost-performance tradeoffs explicit — producing evidence that feeds production-readiness decisions and constrains or supports roadmap and reliability choices. + +## Multi-tenant ownership boundary + +`multi-tenant-saas-architecture` owns tenant semantics, control/application +planes, tenancy and partitioning choices, lifecycle, entitlements, metering and +billing handoffs, and the end-to-end SaaS architecture. This skill owns the +measured capacity and cost consequences: tenant demand distributions, hot tenants, +partition skew, pooled versus siloed headroom, tier-linked quotas and admission, +fairness evidence, shared-cost allocation, and tenant-variable unit cost. + +`secure-software-engineering` owns isolation threats and enforceable controls; +`financial-modeling` owns pricing, margin, and commercial outcomes; +`platform-engineering` owns resource and telemetry implementation; +`product-analytics-and-measurement` owns demand instrumentation and metric +definitions; `site-reliability-engineering` owns SLOs, error budgets, and live +operations; and `software-architecture` owns general system boundary and +topology decisions. A multi-tenant capacity claim must hand off to each owner +whose evidence or authority it requires, rather than absorbing those workflows. diff --git a/capacity-and-cost-engineering/references/multi-tenant-capacity-and-unit-cost.md b/capacity-and-cost-engineering/references/multi-tenant-capacity-and-unit-cost.md new file mode 100644 index 0000000..acea088 --- /dev/null +++ b/capacity-and-cost-engineering/references/multi-tenant-capacity-and-unit-cost.md @@ -0,0 +1,143 @@ +# Multi-Tenant Capacity and Unit Cost + +Use this reference when a shared service serves multiple tenants whose demand, +promises, or resource footprints differ. The purpose is to produce capacity and +cost evidence, not to choose the SaaS domain model or security controls. + +## Start with distributions + +Represent demand as tenant profiles rather than one aggregate average. For each +profile, record the demand units, request or job mix, payload or object size, +concurrency, burst shape, background work, storage growth, and tier or placement +promise. Use observed tenant cohorts where possible and label synthetic profiles +as assumptions. Preserve the distribution's shape: a small number of very large +tenants can dominate a pooled system even when the mean tenant looks modest. + +Useful views include: + +- per-tenant time series for rate, concurrency, queue depth, storage, and work; +- a distribution across tenants for each resource, with the statistic chosen for + the decision rather than a default percentile; +- joint views that show whether high demand, large objects, and burstiness occur + in the same tenants; +- a scenario for new, growing, dormant, migrating, and unusually hot tenants; +- confidence and provenance for every profile, including sampling bias and + omitted tenants. + +Do not substitute a global average, a single representative tenant, or an +arbitrary percentile for the distribution. If a percentile or cap is selected, +explain the user promise, failure consequence, and evidence that make it useful. + +## Find skew and hot tenants + +Map the demand profile onto the resource boundary that can saturate: CPU, +memory, connections, partitions, IOPS, queue workers, cache capacity, search +shards, network, or an external quota. Partition skew is a separate question from +request skew. A tenant may be moderate overall but overload one partition, key +range, shard, worker pool, or availability zone. + +For each hot-tenant or skew scenario, record: + +1. the detection signal and the identity granularity available to operators; +2. the resource and neighboring tenants exposed to contention; +3. the admission, queueing, scheduling, placement, throttling, or isolation + response; +4. the impact on the hot tenant and on protected tenants; +5. the recovery and rebalancing path, including backlog and data-integrity + checks; and +6. the evidence boundary, workload mix, duration, and unresolved gaps. + +Route isolation mechanisms, authorization, and noisy-neighbor threat analysis to +`secure-software-engineering`. Route the end-to-end tenancy, lifecycle, and +placement architecture to `multi-tenant-saas-architecture`. + +## Compare pooled and siloed headroom + +For a pooled deployment, model shared baseline, aggregate demand distribution, +correlation between tenants, admission behavior, and the headroom needed to +protect the promised service during a hot-tenant or dependency scenario. Pooling +can benefit from imperfectly correlated demand, but its usable headroom is +bounded by the resource with the worst contention or skew, not by a fleet average. + +For a siloed or dedicated deployment, model the per-tenant baseline, reserved +headroom, failure-domain requirement, idle capacity, and operational overhead. +Do not assume that a silo is cheaper or more reliable. Compare the scenarios that +matter: ordinary load, tenant growth, burst, tenant failure, placement movement, +and loss of a resource or failure domain. Hybrid placement should show which +resources are pooled and which are dedicated, with separate evidence for each. + +Headroom is a decision variable. State what it protects, the time horizon and +provisioning lead time, the observed saturation behavior, and the cost of unused +capacity. Never present one utilization or headroom percentage as a general rule. + +## Translate tier promises into capacity controls + +For every tier promise, connect the customer-visible statement to a measurable +capacity behavior: sustained demand, burst allowance, concurrency, storage or +job limit, latency treatment, priority, isolation, or recovery treatment. Record +whether the promise is a contract, a product default, or an operational goal. + +Quotas and rate limits are controls, not proof of capacity. Define the scope, +steady behavior, burst behavior, response to excess, fairness objective, and +backpressure or degradation path. Admission should preserve the critical path +and make rejection or delay explicit rather than allowing unbounded queues and +retry amplification. + +Fairness is contextual. Choose the fairness policy from the promise and resource: +weighted shares, reserved capacity, tier priority, work conservation, isolation, +or another explicit rule. Measure both protected-tenant outcomes and the +consequence for the tenant being constrained. Do not use a universal fairness +ratio, quota, utilization target, or rejection threshold. A threshold is valid +only when its rationale, owner, workload, and review trigger are recorded. + +## Produce tenant-variable unit cost + +Separate the cost model into at least: + +- **platform baseline:** costs that remain for the service or pool when tenant + demand is absent or minimal; +- **tenant-variable cost:** incremental compute, storage, transfer, operations, + or other resource cost attributable to a tenant profile; and +- **allocation of shared cost:** the chosen method for assigning pooled cost, + such as measured consumption, reserved entitlement, capacity reservation, or a + transparent blended allocation. + +Report direct or marginal cost separately from fully allocated cost. Use the same +period, scope, and demand denominator. For tenant `t`, a useful model is: + +```text +tenant cost(t) = allocated baseline(t) + measured variable cost(t) +tenant unit cost(t) = tenant cost(t) / tenant demand units(t) +``` + +The allocation method must explain how idle pooled capacity, shared control-plane +work, replication, support, backups, and failure-domain redundancy are treated. +Keep fixed commitments separate from costs that change with demand. Segment units +by materially different work types rather than averaging cheap metadata work with +expensive jobs. Route revenue, pricing, margin, and commercial packaging to +`financial-modeling`; route cost tags, billing exports, and resource policy +implementation to `platform-engineering`. + +## Evidence standard + +Before approving a tenant capacity or unit-cost claim, require representative +per-tenant load and soak evidence. The test should exercise the relevant tenant +distribution concurrently, include hot and skewed profiles, use production-like +data and placement, and observe tenant-level and shared-resource outcomes. Record +latency, errors, queueing, throttling, admission decisions, resource saturation, +partition balance, backlog recovery, and cost inputs. + +Component benchmarks can explain a mapping, but they do not establish end-to-end +fairness or pooled headroom. A fleet aggregate can show total spend, but it does +not establish tenant-variable cost. If representative evidence is unavailable, +the output is a model with an explicit gap and a required test, not a validated +capacity claim. + +## Completion check + +Stop when the model names the tenant distribution, hot/skew scenario, pooled or +siloed headroom choice, tier promise, quota/admission behavior, fairness policy, +baseline and variable cost allocation, representative per-tenant load/soak +evidence, owners, and unresolved assumptions. Escalate a promise that cannot be +supported without silently weakening reliability, privacy, security, or user +outcomes. diff --git a/capacity-and-cost-engineering/references/source-index.md b/capacity-and-cost-engineering/references/source-index.md new file mode 100644 index 0000000..0901382 --- /dev/null +++ b/capacity-and-cost-engineering/references/source-index.md @@ -0,0 +1,34 @@ +# Source Index + +This skill is an original, task-centered synthesis. Public sources inform +terminology and decision pressures; they are not copied as instructional text. + +| Source | Use in this skill | URL | +|---|---|---| +| Google SRE resources | Capacity, overload, service behavior, and evidence-oriented reliability framing | https://sre.google/sre-book/table-of-contents/ | +| OpenSLO specification | Portable vocabulary for connecting service objectives to evidence; SLO ownership remains with SRE | https://github.com/OpenSLO/OpenSLO | +| OpenTelemetry semantic conventions | Tenant-aware measurement vocabulary and observability handoff; implementation remains with platform/telemetry owners | https://opentelemetry.io/docs/specs/semconv/ | +| FinOps Framework | Shared-cost allocation, unit economics, and accountability vocabulary; financial outcomes remain with financial-modeling | https://www.finops.org/framework/ | +| Kubernetes resource management documentation | Resource requests, limits, and scheduling concepts as implementation context; platform-engineering owns configuration | https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ | +| RFC 6585 | HTTP status vocabulary for rate-limit and overload responses; API contract details remain with api-design-and-evolution | https://www.rfc-editor.org/rfc/rfc6585 | + +## Ownership and transformation boundary + +- End-to-end tenant semantics, control/application planes, lifecycle, and + placement architecture belong to `multi-tenant-saas-architecture`. +- Threat controls and tenant isolation evidence belong to + `secure-software-engineering`. +- Financial statements, pricing, margin, and SaaS outcomes belong to + `financial-modeling`. +- Infrastructure and telemetry implementation belong to + `platform-engineering`; SLOs and live operations belong to + `site-reliability-engineering`. +- The supplied private comparison report informed the gap framing only. No + purchased ebook is a source file for this skill, and no protected prose, + table, diagram, example, taxonomy, or chapter structure is reproduced. + +## Review rule + +If a future edit resembles a source's distinctive expression or presentation, +rewrite it from the issue requirements, public sources, and repository ownership +boundaries before publication. diff --git a/capacity-and-cost-engineering/templates/tenant-capacity-model.md b/capacity-and-cost-engineering/templates/tenant-capacity-model.md new file mode 100644 index 0000000..3899da7 --- /dev/null +++ b/capacity-and-cost-engineering/templates/tenant-capacity-model.md @@ -0,0 +1,92 @@ +# Tenant Capacity and Unit-Cost Model + +Fill this template when tenant demand or tier promises change the capacity and +cost decision. Use measured distributions where available; mark assumptions and +synthetic profiles clearly. + +## Service and promise + +- **Service/resource boundary:** _[fill: API, worker pool, database partition, storage, etc.]_ +- **Decision:** _[fill: sizing, placement, quota, tier promise, or cost allocation decision]_ +- **Tenant tiers or profiles in scope:** _[fill: names and why they are representative]_ +- **Customer promises:** _[fill: contractual promises, product defaults, and operational goals separately]_ +- **SLO/performance dependency:** _[fill: owner and relevant target; route SLO definition to SRE]_ + +## Tenant demand profiles + +| Profile or cohort | Tenant count/weight | Demand and mix | Burst/concurrency | Storage/background work | Evidence/confidence | +|---|---:|---|---|---|---| +| _[fill: ordinary pooled tenant]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | +| _[fill: hot or bursty tenant]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | +| _[fill: dedicated/silo tenant]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | + +- **Distribution view:** _[fill: per-tenant time series, selected statistic(s), correlation/skew analysis]_ +- **Sampling limits:** _[fill: omitted tenants, seasonality, new-tenant uncertainty, synthetic data]_ + +## Skew and hot-tenant analysis + +- **Resource/partition boundary:** _[fill: shard, key range, worker, zone, cache, queue, or other]_ +- **Detection signal:** _[fill: tenant-level and shared-resource signal]_ +- **Contention path:** _[fill: which tenants or tiers are affected and how]_ +- **Response:** _[fill: admission, queue, scheduling, throttling, placement, or isolation behavior]_ +- **Recovery/rebalance:** _[fill: backlog, movement, reconciliation, and verification]_ +- **Security handoff:** _[fill: isolation and authorization evidence owned by secure-software-engineering]_ + +## Pooled, siloed, or hybrid comparison + +| Scenario | Pooled baseline/headroom | Siloed baseline/headroom | Hybrid choice | Evidence and tradeoff | +|---|---|---|---|---| +| Ordinary load | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | +| Hot tenant or burst | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | +| Failure-domain loss | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | +| Tenant growth/movement | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | + +- **Headroom rationale:** _[fill: what it protects, lead time, observed saturation, and cost of idle capacity; no universal percentage]_ + +## Quota, admission, and fairness + +- **Quota scope and units:** _[fill: tenant/tier/resource and steady/burst units]_ +- **Admission rule:** _[fill: accept, queue, prioritize, throttle, or reject and why]_ +- **Excess response:** _[fill: status/backpressure/degradation and customer communication]_ +- **Fairness policy:** _[fill: explicit policy for this promise/resource, not a universal ratio]_ +- **Fairness evidence:** _[fill: protected-tenant outcomes, constrained-tenant outcome, workload mix, duration, owner]_ +- **Review trigger:** _[fill: what observed change causes recalibration]_ +- **Implementation handoff:** _[fill: platform owner; this record does not configure infrastructure]_ + +## Tenant-variable unit cost + +- **Period and scope:** _[fill: same period for cost and demand]_ +- **Platform baseline cost:** _[fill: idle/minimum pool, control plane, shared redundancy]_ +- **Variable cost pool:** _[fill: compute, storage, transfer, jobs, support, or other measured costs]_ +- **Shared-cost allocation:** _[fill: measured use, reserved entitlement, capacity reservation, blended, or other rationale]_ +- **Tenant demand unit:** _[fill: request/job/GB/concurrency unit and measurement source]_ + +```text +tenant cost = allocated baseline + measured variable cost +tenant unit cost = tenant cost / tenant demand units +``` + +| Tenant/profile | Allocated baseline | Variable cost | Total cost | Demand units | Unit cost | Confidence/gap | +|---|---:|---:|---:|---:|---:|---| +| _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | _[fill]_ | + +- **Financial handoff:** _[fill: pricing, margin, or commercial decision routed to financial-modeling]_ + +## Representative load/soak evidence + +- **Test plan:** _[fill: link to `templates/load-soak-test-plan.md`]_ +- **Profiles exercised concurrently:** _[fill: ordinary, hot, skewed, tier, placement]_ +- **Environment parity:** _[fill: data volume, topology, partitions, dependencies]_ +- **Observed tenant outcomes:** _[fill: latency, errors, throttles, admission, backlog]_ +- **Observed shared outcomes:** _[fill: saturation, partition balance, recovery, cost inputs]_ +- **Soak findings:** _[fill: leaks, queue growth, fairness drift, or other trends]_ +- **Verdict:** _[fill: validated / model with gap / failed; list criteria]_ + +## Assumptions, ownership, and decision + +- **Assumptions and gaps:** _[fill]_ +- **Capacity/model owner:** _[fill]_ +- **Demand measurement owner:** _[fill: route instrumentation to product-analytics-and-measurement]_ +- **SLO/reliability owner:** _[fill: route SLO and error budget to site-reliability-engineering]_ +- **Decision owner and review date:** _[fill]_ +- **Chosen option and rationale:** _[fill]_ diff --git a/llms.txt b/llms.txt index 3107c8f..37126dc 100644 --- a/llms.txt +++ b/llms.txt @@ -19,7 +19,7 @@ - [binary-analysis](binary-analysis/SKILL.md): Analyze unknown binary files through a deterministic CLI that wraps Ghidra's static-analysis engine. Use when you need to inspect a PE, ELF, or Mach-O file — triage suspicious binaries, map imported APIs, decompile functions, trace call paths, or produce structured evidence reports. Do not use for runtime analysis (debugging, dynamic tracing, sandbox execution), for modifying or patching binaries, or for binaries you already know everything about. The skill owns planning, hypothesis formation, and evidence synthesis; the CLI owns all deterministic operations. - [brand-designer](brand-designer/SKILL.md): Create comprehensive brand identity documentation for any brand. Guides you through documenting strategy, visual identity (logo, color, typography, imagery), voice and tone, application guidelines, governance, and asset inventory. Produces markdown specs, compiled brand books, and brand-compliant images via reference-image-aware generation. Use when you need to capture a brand's identity in structured, durable form — for vault storage, agency handoff, or press kit distribution. - [c4-diagramming](c4-diagramming/SKILL.md): Create C4 software-architecture diagrams using Mermaid or Structurizr. Use when teams need clear system context, container, component, or code-level views. -- [capacity-and-cost-engineering](capacity-and-cost-engineering/SKILL.md): Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. Use when projecting capacity from growth forecasts, sizing for peak events, designing cost-aware scaling policies, defining budget thresholds or quota/rate-limit enforcement, running or planning load/soak tests as capacity evidence, or resolving SLO-cost tradeoffs. Do NOT use for financial P&L statements, fundraising scenarios, or SaaS metrics (route to financial-modeling); for infrastructure implementation or cloud-resource provisioning (route to platform-engineering); or for generic cloud-cost tips and universal utilization targets — this skill does not prescribe fixed savings rates or one-size-fits-all thresholds. +- [capacity-and-cost-engineering](capacity-and-cost-engineering/SKILL.md): Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. Use when projecting capacity from growth forecasts, sizing for peak events, designing cost-aware scaling policies, defining budget thresholds or quota/rate-limit enforcement, running or planning load/soak tests as capacity evidence, resolving SLO-cost tradeoffs, or modeling multi-tenant demand distributions, hot-tenant skew, pooled or siloed headroom, fairness evidence, and tenant-variable unit cost. Do NOT use for financial P&L statements, fundraising scenarios, or SaaS metrics (route to financial-modeling); for infrastructure implementation or cloud-resource provisioning (route to platform-engineering); or for generic cloud-cost tips and universal utilization targets — this skill does not prescribe fixed savings rates or one-size-fits-all thresholds. - [chief-of-staff-methodology](chief-of-staff-methodology/SKILL.md): Prepare accountable executive decisions, information triage, briefing, calendar choices, organizational sensing, and institutional memory without assuming authority or monitoring people. Use when a chief of staff or CoS, executive office, gatekeeping, decision memo, executive briefing, board materials, organizational sensing, team health, institutional memory, calendar triage, meeting audit, strategic time, or attention allocation is requested. - [cli-builder](cli-builder/SKILL.md): Build or refactor CLI tools designed for AI agent consumption: non-interactive, flag-driven, idempotent, with --json output and --dry-run preview. Use when creating a new script the agent will call, adding agent-friendly flags to an existing tool, or debugging why an agent keeps failing to use your CLI. - [cncf-landscape](cncf-landscape/SKILL.md): Use this skill when discovering and comparing cloud-native technologies from the CNCF Landscape for an architecture or engineering decision. Query the live public Landscape API, filter candidates by capability, category, maturity, license, and repository signals, then produce an evidence-backed shortlist with trade-offs, unknowns, and validation steps. Do not use it as a substitute for project documentation, production-readiness testing, legal review, or general architecture methodology. diff --git a/references/skill-triggers.md b/references/skill-triggers.md index 5dc2ced..db34c92 100644 --- a/references/skill-triggers.md +++ b/references/skill-triggers.md @@ -94,7 +94,7 @@ Each skill's `description` field is the canonical routing contract. This conveni | "mermaid-diagrams", "mermaid diagrams" | [mermaid-diagrams](../mermaid-diagrams/SKILL.md) | | "adr-authoring", "adr authoring", "architecture decision record", "fitness function", "decision confirmation" | [adr-authoring](../adr-authoring/SKILL.md) | | "c4-diagramming", "c4 diagramming" | [c4-diagramming](../c4-diagramming/SKILL.md) | -| "capacity engineering", "cost engineering", "capacity model", "capacity planning", "unit cost", "cost per request", "cost per user", "budget threshold", "spending alert", "spending cap", "rate limit enforcement", "quota management", "load test plan", "soak test plan", "capacity projection", "growth forecast capacity", "peak sizing", "degraded capacity", "SLO cost tradeoff", "cost attribution", "cost anomaly review", "cost-aware architecture", "cost-performance tradeoff" | [capacity-and-cost-engineering](../capacity-and-cost-engineering/SKILL.md) | +| "capacity engineering", "cost engineering", "capacity model", "capacity planning", "unit cost", "cost per request", "cost per user", "budget threshold", "spending alert", "spending cap", "rate limit enforcement", "quota management", "load test plan", "soak test plan", "capacity projection", "growth forecast capacity", "peak sizing", "degraded capacity", "SLO cost tradeoff", "cost attribution", "cost anomaly review", "cost-aware architecture", "cost-performance tradeoff", "tenant capacity", "tenant demand distribution", "hot tenant capacity", "partition skew capacity", "pooled headroom", "siloed headroom", "tenant fairness evidence", "tenant variable cost" | [capacity-and-cost-engineering](../capacity-and-cost-engineering/SKILL.md) | | "technology-radar", "technology radar", "technology portfolio governance", "proportional technology governance", "technology advice process", "federated technology decision", "technology standards exception", "technology standards" | [technology-radar](../technology-radar/SKILL.md) | | "strategy", "strategic planning", "OKRs", "strategic narrative", "Five Forces", "Blue Ocean", "competitive positioning", "moat", "Ansoff", "Three Horizons", "market entry", "capital allocation", "M&A evaluation", "BCG Matrix", "portfolio management" | [strategy-frameworks](../strategy-frameworks/SKILL.md) | | "Supabase", "Supabase CLI", "supabase start", "supabase migration", "Supabase Auth", "Supabase RLS", "Supabase Storage", "Supabase Realtime", "Edge Functions", "self-host Supabase", "self-hosted Supabase", "Supabase Docker", "Supabase backup", "Supabase restore", "Supabase upgrade" | [supabase](../supabase/SKILL.md) |