diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 9a9a09c..fdcc2e6 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -608,6 +608,15 @@ "strict": false, "description": "Use this skill when building or operating internal developer platforms: infrastructure as code, CI/CD, container orchestration, service networking, secrets, and observability. Do not use it to define release process, promotion, rollout, or rollback policy; use release-engineering for that delivery model." }, + { + "name": "product-analytics-and-measurement", + "source": "./", + "skills": [ + "./product-analytics-and-measurement" + ], + "strict": false, + "description": "Define observable, governed evidence for product outcomes through metric trees, event/tracking plans, instrumentation QA, and measurement governance. Use when defining product metrics, designing tracking plans, building event taxonomies, running product funnels or cohort analysis, setting up dashboard contracts, or auditing instrumentation quality. Do NOT use for statistical inference or experiment design (route to data-scientist), data pipeline implementation (route to data-engineering), data architecture decisions (route to data-architect), observability or infrastructure monitoring (route to site-reliability-engineering), or general business intelligence and dashboard building (route to BI tooling)." + }, { "name": "product-design-and-ux", "source": "./", diff --git a/.codex-plugin/plugin.json b/.codex-plugin/plugin.json index 22fd9ba..a44d4fc 100644 --- a/.codex-plugin/plugin.json +++ b/.codex-plugin/plugin.json @@ -88,6 +88,7 @@ "./pace-plan", "./peertube", "./platform-engineering", + "./product-analytics-and-measurement", "./product-design-and-ux", "./product-discovery", "./product-experimentation", diff --git a/README.md b/README.md index 74cb8e6..46cc33f 100644 --- a/README.md +++ b/README.md @@ -287,6 +287,10 @@ PeerTube federated video platform from the terminal. Browse videos and channels, Infrastructure as code, CI/CD, container orchestration, service networking — methodology and reference patterns for building and operating internal developer platforms. +### [product-analytics-and-measurement](product-analytics-and-measurement/SKILL.md) + +Turn intended product outcomes into observable, governed evidence. Covers metric trees (leading/lagging indicators, countermetrics, guardrails), event and tracking plans (identity, session, data quality, ownership), instrumentation QA, funnels, cohorts, retention, adoption, dashboard contracts, privacy-aware measurement, and decision cadence. Routes statistical inference to data-scientist and data pipelines to data-engineering. + ### [product-design-and-ux](product-design-and-ux/SKILL.md) Turn validated evidence and approved product scope into traceable user-facing behavior: information architecture, plain-language content, task flows, applicable state and recovery models, interface contracts, authorized usability evidence, and observable engineering handoffs. Portable and framework-neutral; routes WCAG/ARIA depth to web-accessibility and formal software acceptance to spec-driven-development. Ships 10 focused references and 6 fillable templates. diff --git a/llms.txt b/llms.txt index 8e0eb16..61c1869 100644 --- a/llms.txt +++ b/llms.txt @@ -69,6 +69,7 @@ - [pace-plan](pace-plan/SKILL.md): Build, coordinate, operate, troubleshoot, exercise, and improve an authorized Primary, Alternate, Contingency, and Emergency communications plan. Use for resilient emergency-communications paths and their ownership, triggers, check-ins, tests, and corrective actions. Do not use for generic incident status messaging, frequency or channel planning, radio programming, or unauthorized transmission and activation. - [peertube](peertube/SKILL.md): Browse PeerTube federated video from the terminal: view videos and channels, search across instances, check server stats, and manage your account. Uses OAuth2 authentication with token persistence. Use when the user mentions PeerTube, federated video, decentralized video platforms, or browsing/uploading to a PeerTube instance. - [platform-engineering](platform-engineering/SKILL.md): Use this skill when building or operating internal developer platforms: infrastructure as code, CI/CD, container orchestration, service networking, secrets, and observability. Do not use it to define release process, promotion, rollout, or rollback policy; use release-engineering for that delivery model. +- [product-analytics-and-measurement](product-analytics-and-measurement/SKILL.md): Define observable, governed evidence for product outcomes through metric trees, event/tracking plans, instrumentation QA, and measurement governance. Use when defining product metrics, designing tracking plans, building event taxonomies, running product funnels or cohort analysis, setting up dashboard contracts, or auditing instrumentation quality. Do NOT use for statistical inference or experiment design (route to data-scientist), data pipeline implementation (route to data-engineering), data architecture decisions (route to data-architect), observability or infrastructure monitoring (route to site-reliability-engineering), or general business intelligence and dashboard building (route to BI tooling). - [product-design-and-ux](product-design-and-ux/SKILL.md): Define user-facing product behavior from validated evidence and approved scope. Use for information architecture, task flows, state and recovery models, interface contracts, usability-study plans, interaction-pattern tradeoffs, or engineering UX handoffs. Use after product discovery and product decisions; route WCAG/ARIA conformance work to web-accessibility and formal software specifications to spec-driven-development. - [product-discovery](product-discovery/SKILL.md): Discover product requirements from human stakeholders — map who to talk to, ask questions that surface hidden assumptions, detect gaps in real time, resolve conflicts, and translate conversations into structured SDD specs. Phase 0 upstream of Spec-Driven Development. - [product-experimentation](product-experimentation/SKILL.md): Run end-to-end product experiments from assumption to decision: translate assumptions into testable hypotheses and experiment briefs, select the right method among qualitative interviews, prototypes, concierge tests, fake doors, feature flags, and A/B tests, and produce readouts that update the roadmap and decision record. Do not use when a qualitative or prototype test is the clearly right answer without statistical measurement; do not prescribe A/B testing by default; do not treat statistical significance as the only decision criterion or hide ethical and guardrail considerations. diff --git a/product-analytics-and-measurement/README.md b/product-analytics-and-measurement/README.md new file mode 100644 index 0000000..5856898 --- /dev/null +++ b/product-analytics-and-measurement/README.md @@ -0,0 +1,55 @@ +# Product Analytics and Measurement + +Turn intended product outcomes into observable, governed evidence. + +## Why Install This Skill + +Most product teams measure what is easy, not what matters. They inherit dashboards built by previous teams, track events nobody reviews, and discover too late that a "North Star" metric was never actually measurable. When metrics conflict across teams — marketing counts an activation differently than product — there is no contract to resolve the disagreement. + +After installing this skill, your agent can define a complete measurement system for any product: decompose a North Star into a metric tree with leading and lagging indicators, produce an event tracking plan with identity and data-quality rules, audit instrumentation before trusting the numbers, and establish a dashboard contract and decision cadence so measurement drives action instead of accumulating dashboards. The skill works across SaaS, internal tools, public services, and consumer products — not just one business model. + +## What You Get + +| Directory | What it provides | +|-----------|-----------------| +| [`SKILL.md`](SKILL.md) | Core methodology: metric trees, tracking plans, instrumentation QA, privacy-aware measurement, dashboard contracts, and decision cadence. Includes the Loading Guide, When to Use / When Not to Use, and related-skill routing. | +| [`references/discovery-brief.md`](references/discovery-brief.md) | Bounded discovery brief mapping existing metrics content across product-strategy, go-to-market, financial-modeling, data-engineering, and data-scientist, with ownership boundaries. | +| [`references/metric-tree.md`](references/metric-tree.md) | How to build, validate, and maintain a metric tree: leading/lagging indicators, countermetrics, ownership rules, measurability gates, and examples across product types. | +| [`templates/tracking-plan.md`](templates/tracking-plan.md) | Fillable tracking plan template: event taxonomy, property schema, identity resolution, session definition, data quality rules, and ownership fields. | +| [`templates/instrumentation-qa-checklist.md`](templates/instrumentation-qa-checklist.md) | Checklist for verifying instrumentation quality across client, server, pipeline, and end-to-end layers before trusting a metric on a dashboard. | +| [`templates/outcome-review.md`](templates/outcome-review.md) | Template for a recurring product outcome review: metric tree health, countermetric checks, decision log, and metric deprecation/retirement decisions. | +| [`evals/evals.json`](evals/evals.json) | Output-quality evaluation cases covering new features, internal products, public services, conflicting metrics, and unmeasurable-North-Star rejection. | + +## Quick Start + +Load the skill when your task involves defining product metrics, designing tracking, or auditing instrumentation. The agent will follow the working method in `SKILL.md`: + +1. Define the outcome and decompose into a metric tree (`references/metric-tree.md`). +2. Verify measurability — reject unmeasurable metrics rather than fabricate proxies. +3. Produce a tracking plan contract (`templates/tracking-plan.md`). +4. QA the instrumentation before trusting the dashboard (`templates/instrumentation-qa-checklist.md`). +5. Establish a decision cadence and outcome review rhythm (`templates/outcome-review.md`). + +No API keys, external services, or software dependencies are required. The skill is purely methodological. + +## Triggers + +Load this skill when the user asks about: + +- Defining product metrics, North Star, success metrics, or OKR measurement +- Building a metric tree, leading/lagging indicators, or countermetrics +- Designing an event taxonomy, tracking plan, or analytics specification +- Auditing instrumentation quality or analytics data trustworthiness +- Setting up product funnels, cohort definitions, retention measurement, or adoption tracking +- Creating dashboard contracts with metric definitions, sources, and ownership +- Resolving conflicting metric definitions between teams +- Establishing a product measurement cadence or outcome review process +- Privacy-aware measurement design (consent boundaries, minimization, aggregation thresholds) + +Do **not** load for statistical inference or experiment design (use `data-scientist`), data pipeline implementation (use `data-engineering`), data architecture (use `data-architect`), observability or infrastructure monitoring (use `site-reliability-engineering`), or general BI dashboard building. + +## Requirements + +- No software dependencies, API keys, or external services required. +- Compatible with any Agent Skills harness that supports markdown skill loading. +- Purely methodological — all outputs are documents (metric trees, tracking plans, checklists, review templates). diff --git a/product-analytics-and-measurement/SKILL.md b/product-analytics-and-measurement/SKILL.md new file mode 100644 index 0000000..633ce18 --- /dev/null +++ b/product-analytics-and-measurement/SKILL.md @@ -0,0 +1,166 @@ +--- +name: product-analytics-and-measurement +description: >- + Define observable, governed evidence for product outcomes through metric + trees, event/tracking plans, instrumentation QA, and measurement governance. Use + when defining product metrics, designing tracking plans, building event taxonomies, + running product funnels or cohort analysis, setting up dashboard contracts, or + auditing instrumentation quality. Do NOT use for statistical inference or + experiment design (route to data-scientist), data pipeline implementation (route + to data-engineering), data architecture decisions (route to data-architect), + observability or infrastructure monitoring (route to site-reliability-engineering), + or general business intelligence and dashboard building (route to BI tooling). +license: MIT +metadata: + tags: product-analytics, measurement, metrics, instrumentation, tracking-plan, + event-taxonomy, funnels, cohorts, retention, adoption, dashboard-contracts, + privacy-aware-measurement +--- + +# Product Analytics and Measurement + +A dedicated operating method for turning intended product outcomes into observable, +governed evidence. This skill covers the full measurement lifecycle: from defining +what success looks like through metric trees and leading/lagging indicators, to +specifying how that success is tracked through event taxonomies and tracking plans, +to verifying that the instrumentation actually captures what it claims. + +## Loading Guide + +Load this skill when the task involves any of: + +| Trigger | What to load | +|---------|-------------| +| Define product outcomes, North Star, or success metrics | `SKILL.md` + `references/metric-tree.md` | +| Design an event taxonomy or tracking plan | `SKILL.md` + `templates/tracking-plan.md` | +| Audit or QA instrumentation quality | `SKILL.md` + `templates/instrumentation-qa-checklist.md` | +| Build product funnels, path analysis, or cohort definitions | `SKILL.md` + `references/metric-tree.md` | +| Set up a dashboard contract or measurement governance | `SKILL.md` + `templates/outcome-review.md` | +| Resolve conflicting metric definitions across teams | `SKILL.md` + `references/discovery-brief.md` (ownership boundaries) | +| Design a privacy-aware measurement strategy | `SKILL.md` | + +## When to Use + +- You need to define measurable outcomes, metric trees, and leading/lagging indicators for a product (SaaS, internal tool, public service, or consumer). +- You need a tracking plan that specifies events, properties, identity resolution, session boundaries, data quality rules, and ownership. +- You need to verify that existing instrumentation produces trustworthy data. +- You need a dashboard contract that pins metric definitions, sources, refresh cadences, and ownership. +- You need a product outcome review cadence that ties measurements back to decisions. + +## When Not to Use + +This skill does **not** replace: + +- **Data architecture** — schema design, storage selection, data modeling at the infrastructure level belong to `../data-architect/SKILL.md`. +- **Data engineering** — pipeline implementation, ETL/ELT, dbt models, and data quality monitoring at the operational level belong to `../data-engineering/SKILL.md`. +- **Statistical inference** — experiment design, hypothesis testing, causal inference, and model selection belong to `../data-scientist/SKILL.md`. +- **Observability** — system health, latency, error budgets, and infrastructure monitoring belong to `../site-reliability-engineering/SKILL.md`. +- **Business intelligence** — dashboard building, report generation, and data visualization as an end in itself. This skill governs the *measurement contract* and *metric definitions* that feed BI, not the BI layer itself. + +This skill also does **not** prescribe universal benchmark thresholds. Every metric target, retention curve, or conversion rate cited in outputs is labeled as contextual — tied to the specific product, market, stage, and business model under discussion. + +## Core Concepts + +### Metric Tree + +A hierarchical decomposition of a North Star or primary outcome into its constituent drivers. Each node in the tree is: + +- **Observable** — it can be measured from instrumentation or a defined data source. +- **Actionable** — a change in the metric suggests a specific product or operational response. +- **Owned** — exactly one team or role is accountable for its definition, collection, and interpretation. + +The metric tree distinguishes **leading indicators** (predictive, sensitive to change) from **lagging indicators** (confirmatory, slow to move) and pairs every metric with at least one **countermetric** (a guardrail that detects adverse side effects). See `references/metric-tree.md`. + +### Event and Tracking Plan + +An **event** is a recorded action or state change with a timestamp, actor, and properties. A **tracking plan** is the contract that makes events trustworthy: + +- **Event name and taxonomy** — consistent naming conventions, versioned, with a catalog of standard events. +- **Properties** — typed fields with allowed values, required vs optional, and default behaviors. +- **Identity resolution** — how actors are identified across sessions, devices, and authentication states (anonymous id, logged-in id, merged identities). +- **Session definition** — timeout rules, activity boundaries, and cross-platform session continuity. +- **Data quality rules** — validation checks (null rates, freshness, cardinality bounds, distribution drift). +- **Ownership** — who defines the event, who instruments it, who consumes it, and who is paged when it breaks. + +See `templates/tracking-plan.md`. + +### Instrumentation QA + +Before a metric appears on a dashboard, the instrumentation that produces it must be verified. The QA checklist covers: + +- **Client-side** — event firing on correct triggers, property values match UI state, no duplicate firings, no missing required properties. +- **Server-side** — event ingestion, transformation correctness, identity stitching, timestamp integrity. +- **Pipeline** — schema compatibility, no silent drops, latency within SLA, deduplication behavior. +- **End-to-end** — the metric computed from raw events matches the value observed through a manual product walkthrough. + +See `templates/instrumentation-qa-checklist.md`. + +### Privacy-Aware Measurement + +Every measurement design must address: + +- **Consent boundary** — what is measured before consent vs after; what is never measured. +- **Minimization** — collect only what the metric requires; avoid speculative properties. +- **Aggregation threshold** — metrics reported only when the cohort meets a minimum size. +- **Retention and deletion** — raw event retention period; how measurement continues after data deletion requests. +- **Jurisdictional awareness** — note when metric collection crosses regulatory boundaries (GDPR, CCPA, etc.). + +### Dashboard Contracts + +A dashboard without a contract is an argument waiting to happen. A dashboard contract pins: + +- **Metric definition** — the exact formula, data source, and filter conditions. +- **Refresh cadence** — how often the dashboard updates and the acceptable staleness. +- **Owner** — who is accountable for accuracy and who to contact when numbers look wrong. +- **Audience and decision** — who consumes this dashboard and what decision it informs. +- **Thresholds and alerts** — when a metric value triggers investigation (not a universal benchmark, but a product-specific threshold). + +### Decision Cadence + +Measurement without a decision cadence is measurement theater. This skill guides the rhythm: + +- **Daily** — operational metrics, anomaly detection, automated alerts. +- **Weekly** — leading indicators, experiment readouts, funnel health. +- **Monthly** — outcome reviews, metric tree health, countermetric checks. +- **Quarterly** — North Star reassessment, metric tree restructuring, tracking plan deprecation. + +## Working Method + +1. **Start with the outcome, not the event.** Define what success looks like and decompose it into a metric tree before designing any tracking. +2. **Verify measurability before committing.** If a metric cannot be observed from available or feasibly instrumented data sources, reject it — do not fabricate a proxy without documenting the instrumentation constraint. +3. **Design the tracking plan as a contract.** Every event has a defined owner, schema, quality rule, and consumer. No orphan events. +4. **QA instrumentation before trusting the dashboard.** Run the instrumentation QA checklist on every new or changed event before the metric enters a decision cadence. +5. **Pair every metric with a countermetric.** For every metric that drives a decision, define at least one countermetric that would reveal an adverse side effect. +6. **Label every threshold as contextual.** State the product, market, stage, and business model that the threshold assumes. +7. **Review measurement governance periodically.** Deprecate unused events, retire stale metrics, update metric trees when the product strategy shifts. + +## File Map + +| Path | Loaded when | +|------|------------| +| `references/discovery-brief.md` | Understanding ownership boundaries and how this skill relates to existing catalog skills | +| `references/metric-tree.md` | Building or reviewing a metric tree, defining leading/lagging indicators, selecting countermetrics | +| `templates/tracking-plan.md` | Designing an event taxonomy or writing a tracking plan contract | +| `templates/instrumentation-qa-checklist.md` | Auditing or verifying instrumentation quality before trusting metrics | +| `templates/outcome-review.md` | Running a product outcome review or setting up a measurement decision cadence | +| `evals/evals.json` | Evaluating the skill's output quality across representative scenarios | + +## Related Skills + +### Routes to (resolvable, existing skills) + +- `../data-engineering/SKILL.md` — for pipeline implementation, ETL/ELT, dbt models, and data quality monitoring at the operational level. +- `../data-scientist/SKILL.md` — for statistical inference, experiment design, hypothesis testing, causal inference, and model selection. +- `../product-strategy/SKILL.md` — for North Star definition, product vision, PMF assessment, and roadmap prioritization. +- `../go-to-market/SKILL.md` — for growth modeling, CAC/LTV, cohort analysis by channel, and PLG/SLG funnel metrics. +- `../financial-modeling/SKILL.md` — for unit economics, ARR/MRR, churn, NDR, and SaaS operating metrics definitions. + +### Feeds (prose references to skills not yet landed) + +This skill produces measurement contracts, metric trees, and tracking plans that feed: + +- **product-roadmapping-and-portfolio** — roadmap decisions grounded in measured outcomes. +- **product-experimentation** — experiment design and readout on a trustworthy measurement foundation. +- **product-adoption** — adoption funnels, activation metrics, and time-to-value measurement. +- **conditional-customer-success** — account health scores and customer outcome tracking. +- **product-lifecycle-learning** — post-launch learning loops and metric-informed iteration. diff --git a/product-analytics-and-measurement/evals/evals.json b/product-analytics-and-measurement/evals/evals.json new file mode 100644 index 0000000..0dcfd27 --- /dev/null +++ b/product-analytics-and-measurement/evals/evals.json @@ -0,0 +1 @@ +{"schema_version": 1, "skill_name": "product-analytics-and-measurement", "evals": [{"id": "new-feature-metrics", "prompt": "We are launching a new collaborative editing feature in our SaaS document product. I need to define the success metrics, tracking plan, and countermetrics before we ship. The feature lets multiple users edit a document simultaneously and see each other's changes in real time.", "expected_output": "A metric tree decomposing the collaborative editing outcome into leading indicators (adoption rate, time-to-first-collaboration, concurrent editors per document) and lagging indicators (retention of collaborative users), each paired with a countermetric (e.g., edit conflict rate as guardrail against adoption push, support tickets as guardrail for quality). A tracking plan contract specifying events (e.g., collaboration_session_started, edit_conflict_occurred) with property schemas, identity resolution for anonymous-to-authenticated transitions, data quality rules, and ownership. All thresholds labeled with product context (B2B SaaS, mid-market, growth stage).", "assertions": ["The response defines a metric tree with at least one leading indicator and one lagging indicator", "The response pairs each decision-driving metric with a countermetric", "The response includes a tracking plan with event names, property schemas, and ownership", "The response addresses identity resolution across authentication states", "The response labels thresholds and targets with product context (product type, market, stage, business model)", "The response does not cite universal benchmarks without context"]}, {"id": "internal-product-metrics", "prompt": "Our internal developer platform team wants to measure whether the platform is actually saving developer time. This is an internal tool used by 200 engineers across the company. We cannot use revenue or conversion metrics. Help us define a measurement approach.", "expected_output": "A measurement approach designed for an internal product, using productivity-oriented metrics (e.g., time-to-first-deploy, CI pipeline duration, % services using platform defaults) rather than revenue or conversion. The metric tree includes a countermetric for each primary metric (e.g., deployment failure rate as guardrail against speed optimization, shadow IT incidents as guardrail against forced adoption). The response explicitly acknowledges the internal-product context and does not apply SaaS-specific metrics (MRR, churn, CAC) inappropriately. It addresses identity resolution in an enterprise SSO context and data quality rules for internal telemetry.", "assertions": ["The response designs metrics appropriate for an internal product, not SaaS revenue or conversion metrics", "The response includes at least one countermetric that guards against perverse incentives in internal tool adoption", "The response addresses identity resolution in an enterprise authentication context", "The response does not apply SaaS metrics (MRR, churn, CAC) to the internal-product use case", "The response labels the product context as internal tool or internal platform"]}, {"id": "public-service-measurement", "prompt": "Our government digital service allows citizens to apply for benefits online. We need to measure not just completion rates but also equity of access and outcomes. How should we approach measurement for a public service where the goal is successful citizen outcomes, not revenue?", "expected_output": "A measurement framework for a public service that defines success in terms of citizen outcomes (successful application rate, timeliness, equity across demographic segments) rather than commercial metrics. The metric tree includes access metrics (completion rate by device type, language, accessibility), timeliness metrics (time-to-decision, % within service standard), and equity metrics (outcome rate by demographic segment). Countermetrics include support call volume (indicates access failure), appeals rate (indicates incorrect decisions), and error rate on expedited applications. Privacy-aware measurement addresses consent, minimization, aggregation thresholds for small demographic subgroups, and jurisdictional regulatory considerations (GDPR or equivalent).", "assertions": ["The response defines success in terms of citizen outcomes, not commercial or revenue metrics", "The response includes equity measurement across demographic or accessibility segments", "The response addresses privacy-aware measurement including consent boundaries and aggregation thresholds for small cohorts", "The response includes countermetrics that guard against optimizing for speed at the cost of accuracy or equity", "The response labels the product context as public service or government service"]}, {"id": "conflicting-metrics-resolution", "prompt": "Our marketing team defines 'activated user' as someone who completed onboarding within 7 days. The product team defines it as someone who performed a core action within 30 days. These conflicting definitions are causing arguments in every executive review. How do we resolve this?", "expected_output": "A conflict resolution approach that establishes a single source of truth for the metric definition, with clear ownership assignment. The response recommends: (1) identifying which team owns the metric definition based on who is accountable for the outcome, (2) defining both metrics explicitly with different names (e.g., 'onboarding-complete activation' vs 'core-action activation') if they measure different things, (3) documenting the exact formula, data source, and filter conditions for each in a dashboard contract, (4) establishing a single owner who can change the definition and a process for proposing changes. The response does not pick a winner without analysis — it provides the framework for resolving the conflict.", "assertions": ["The response provides a framework for resolving conflicting metric definitions, not just picking one team's definition", "The response recommends distinct metric names when the definitions measure genuinely different things", "The response requires exact formulas, data sources, and filter conditions to be documented", "The response assigns clear ownership for each metric definition", "The response addresses the organizational dimension of metric conflict, not just the technical one"]}, {"id": "unmeasurable-north-star-rejection", "prompt": "Our CEO wants our North Star to be 'customer delight.' We need to build an instrumentation plan to track this. Can you help us define the events and tracking we need?", "expected_output": "A rejection of the request as stated, with a clear explanation that 'customer delight' is not directly measurable from instrumentation — it is a latent construct, not an observable event. The response names the missing instrumentation constraint: there is no event that fires when a customer experiences delight; delight must be operationalized through observable proxy metrics (e.g., NPS survey response, feature usage patterns correlated with retention, support ticket sentiment). The response does not fabricate a tracking plan for 'delight' events or pretend the construct is directly measurable. It offers a path forward: decompose 'customer delight' into observable, actionable sub-metrics that together approximate the construct, and explicitly label each as a proxy with known gaps.", "assertions": ["The response rejects 'customer delight' as a directly measurable North Star", "The response names the specific instrumentation constraint: delight is a latent construct, not an observable event", "The response does not fabricate event names or tracking for 'delight' directly", "The response offers a path forward that decomposes the construct into observable proxy metrics", "The response labels any proposed proxy metrics with their known gaps compared to the ideal construct"]}, {"id": "privacy-boundary-measurement", "prompt": "We are designing analytics for a health-related consumer app. We need to track user engagement and outcomes, but we also have strict privacy requirements: users must explicitly consent to tracking, and we must be able to continue basic measurement even when users decline or delete their data. How should we design the measurement strategy?", "expected_output": "A privacy-aware measurement design that separates pre-consent (minimal, strictly necessary) events from post-consent events. The response defines consent boundaries, specifies which metrics can be computed from aggregated or anonymized data when users opt out, addresses retention policies for raw events, and explains how measurement continues after data deletion requests (e.g., aggregate metrics unaffected, user-level metrics removed). It includes aggregation thresholds to prevent re-identification from small cohorts and notes jurisdictional regulatory considerations. The tracking plan template includes consent boundary, data retention, deletion handling, and aggregation minimum fields.", "assertions": ["The response defines a consent boundary separating pre-consent from post-consent measurement", "The response specifies which metrics remain computable when users opt out of tracking", "The response addresses data retention and deletion handling in the measurement design", "The response includes aggregation thresholds to prevent re-identification", "The response does not recommend collecting data that requires consent without addressing the consent workflow"]}]} diff --git a/product-analytics-and-measurement/references/discovery-brief.md b/product-analytics-and-measurement/references/discovery-brief.md new file mode 100644 index 0000000..0802f8c --- /dev/null +++ b/product-analytics-and-measurement/references/discovery-brief.md @@ -0,0 +1,93 @@ +# Discovery Brief: Product Analytics and Measurement + +## Purpose + +This brief maps existing metrics-related content across the agent-skills catalog and resolves ownership boundaries for the new `product-analytics-and-measurement` skill. It ensures the skill adds net-new capability rather than duplicating existing methodology. + +## Skills Surveyed + +### product-strategy (`../product-strategy/SKILL.md`) + +**What it owns:** North Star metric definition, product vision, competitive positioning, roadmap prioritization (RICE, Kano, OST), product-market fit assessment (Sean Ellis test, retention curves), market sizing (TAM/SAM/SOM), platform strategy, and product lifecycle management. + +**Boundary:** Product-strategy defines *what* the North Star is at the strategic level. It does not specify *how* to measure it — it names the metric (e.g., "weekly active users") but does not define the event taxonomy, tracking plan, instrumentation QA, or dashboard contract needed to produce that metric reliably. Product-strategy's retention curves are strategic frameworks (e.g., "does the retention curve flatten above 30%?"), not operational cohort definitions with identity resolution and session boundary rules. + +**Handoff:** This skill (product-analytics-and-measurement) receives the North Star and strategic metric choices from product-strategy and operationalizes them into measurable, governed evidence. + +### go-to-market (`../go-to-market/SKILL.md`) + +**What it owns:** Positioning and messaging (April Dunford), acquisition channel strategy (paid, organic, PLG, SLG), brand architecture, growth modeling (CAC/LTV by channel, cohort analysis by acquisition channel), market entry strategy, and competitive response. + +**Boundary:** GTM's growth modeling uses CAC/LTV and cohort analysis as strategic frameworks for channel investment decisions. It does not define how those cohorts are constructed in the analytics system, how identity is resolved across anonymous-to-known-user transitions, or how the event data feeding CAC/LTV calculations is instrumented and QA'd. GTM's funnel metrics are channel-level; this skill defines the underlying event and session model that makes those funnels computable. + +**Handoff:** This skill provides the measurement infrastructure (event definitions, identity resolution, tracking plan) that GTM's growth models consume. The two skills are complementary: GTM asks "which channel delivers the best CAC/LTV?", this skill ensures CAC and LTV are computed from trustworthy, well-instrumented data. + +### financial-modeling (`../financial-modeling/SKILL.md`) + +**What it owns:** Assumptions-led financial models, unit economics (CAC, LTV, contribution margin, payback), pricing strategy, fundraising scenarios, cap tables, and SaaS operating metrics (ARR/MRR, churn, NDR, Rule of 40, Magic Number, burn multiple). + +**Boundary:** Financial-modeling defines *what* these metrics mean for financial analysis and scenario planning. It does not define the event-level instrumentation or tracking plan that produces the raw data feeding those metrics. For example, financial-modeling defines churn as "the percentage of revenue or customers lost over a period," but does not specify how churn events are captured, how the customer lifecycle state machine is modeled in the analytics system, or how to QA the churn calculation end-to-end. + +**Handoff:** This skill defines the measurement contracts (event definitions, data quality rules, metric formulas at the analytics layer) that financial-modeling's SaaS metrics depend on. Financial-modeling is the consumer of well-measured ARR, churn, and NDR; this skill is the producer. + +### data-engineering (`../data-engineering/SKILL.md`) + +**What it owns:** Database operations, ETL/ELT pipeline design (dbt patterns, incremental loading), SQL analytical patterns, data quality monitoring at the pipeline level, schema migration, and storage infrastructure management. + +**Boundary:** Data-engineering implements the pipelines that move and transform data. This skill defines *what* data should exist (the tracking plan, the event schema, the quality rules) and *how to verify* it is correct (instrumentation QA). Data-engineering then implements the pipeline that ingests, transforms, and stores that data. The handoff is: this skill produces a tracking plan contract; data-engineering builds the pipeline to ingest it. + +**Handoff:** Clear and complementary. This skill is the "what and why" of measurement data; data-engineering is the "how" of moving and storing it. Neither duplicates the other. + +### data-scientist (`../data-scientist/SKILL.md`) + +**What it owns:** Statistical inference, experimental design (A/B testing, power analysis), causal inference (DAGs, IV, RDD, DID), regression and Bayesian analysis, model selection, and research methodology. + +**Boundary:** Data-scientist analyzes data that has already been collected and structured. This skill ensures that the data feeding those analyses is well-defined, properly instrumented, and trustworthy. For example, data-scientist runs the statistical test on an A/B experiment; this skill ensures the experiment's success metric is observable, properly tracked, and QA'd before the test begins. + +**Handoff:** Data-scientist is the consumer of well-measured data. This skill ensures the measurement foundation exists before statistical analysis begins. + +## What This Skill Owns + +| Owned capability | Not owned by any existing skill | +|-----------------|--------------------------------| +| Metric tree construction and governance | No skill decomposes a North Star into an owned, observable, actionable tree with leading/lagging indicators and countermetrics. | +| Event taxonomy and tracking plan contracts | No skill defines event naming conventions, property schemas, identity resolution rules, session boundaries, data quality rules, and event ownership in a unified contract. | +| Instrumentation QA methodology | No skill provides a systematic checklist for verifying instrumentation correctness across client, server, pipeline, and end-to-end layers. | +| Dashboard contracts | No skill pins metric definitions, data sources, refresh cadences, and ownership into a dashboard-level contract. | +| Measurement decision cadence | No skill defines the rhythm (daily/weekly/monthly/quarterly) at which metrics are reviewed, deprecated, and updated. | +| Privacy-aware measurement design | No skill addresses consent boundaries, minimization, aggregation thresholds, and retention in the context of product measurement. | +| Cross-team metric conflict resolution | No skill provides a framework for resolving when two teams define the same metric differently. | + +## What This Skill Routes (Does Not Own) + +| Routed capability | Target skill | +|------------------|-------------| +| Statistical inference, experiment design, causal analysis | `data-scientist` | +| Pipeline implementation, ETL/ELT, dbt models | `data-engineering` | +| Data architecture, storage selection, schema design at infrastructure level | `data-architect` | +| North Star and product vision definition | `product-strategy` | +| Growth modeling, CAC/LTV by channel | `go-to-market` | +| Financial SaaS metrics (ARR, NDR, Rule of 40) | `financial-modeling` | +| System observability, error budgets, infrastructure monitoring | `site-reliability-engineering` | +| General BI dashboard building and data visualization | BI tooling (external to this catalog) | + +## Product Scope Beyond SaaS + +This skill is designed to work across product types, not only SaaS: + +- **SaaS products** — subscription metrics, activation funnels, feature adoption, churn measurement. +- **Internal tools** — productivity metrics, workflow completion rates, time-to-task, adoption within the organization. +- **Public services** — service delivery outcomes, accessibility metrics, equity measurement, constituent satisfaction. +- **Consumer products** — engagement loops, retention cohorts, content consumption patterns, network effects. + +The metric tree and tracking plan templates are product-type-agnostic. Examples in reference files are labeled with their product context so readers can adapt, not blindly copy. + +## Key Design Decisions + +1. **Prose routing to not-yet-landed skills.** Five consumer skills (product-roadmapping-and-portfolio, product-experimentation, product-adoption, conditional-customer-success, product-lifecycle-learning) are referenced by name in prose only, not as markdown links, because they do not yet exist in the catalog. When they land, those prose references become resolvable links. + +2. **No universal benchmarks.** Every threshold, target, or benchmark cited in examples is explicitly labeled with its product, market, stage, and business model context. The skill refuses to provide "industry standard" conversion rates or retention benchmarks without context. + +3. **Reject unmeasurable metrics.** If a requested metric cannot be observed from available or feasibly instrumented data sources, the skill rejects it rather than fabricating a proxy. The rejection names the missing instrumentation constraint so the team knows what would need to change to make the metric measurable. + +4. **Minimal but complete.** The skill covers the full measurement lifecycle without duplicating existing specialist skills. It adds net-new capability in the areas no existing skill covers. diff --git a/product-analytics-and-measurement/references/metric-tree.md b/product-analytics-and-measurement/references/metric-tree.md new file mode 100644 index 0000000..49e7505 --- /dev/null +++ b/product-analytics-and-measurement/references/metric-tree.md @@ -0,0 +1,178 @@ +# Metric Tree Reference + +A metric tree decomposes a product's primary outcome (often called the North Star) into a hierarchy of measurable, actionable, and owned sub-metrics. It is the single source of truth for "how we measure success." + +## Structure + +``` +North Star (primary outcome) +├── Driver 1 (leading indicator) +│ ├── Sub-driver 1a +│ │ ├── Metric (owned by team X) +│ │ └── Countermetric (guardrail) +│ └── Sub-driver 1b +├── Driver 2 (lagging indicator) +│ └── ... +└── Driver 3 (leading indicator) + └── ... +``` + +## Node Properties + +Every node in the metric tree must satisfy four properties: + +### 1. Observable + +The metric can be measured from instrumentation or a defined data source. If no data source exists and none can be feasibly created, the metric is **not currently observable** — and must be rejected or marked as aspirational with a named constraint. + +**Observability test:** +- Can you name the specific event(s) or data source that produces this metric? +- Can you write the SQL or analytics query that computes it? +- Can you trace the metric value on a dashboard back to raw events? + +### 2. Actionable + +A change in the metric suggests a specific product or operational response. If the metric moves and nobody knows what to do, it is not actionable. + +**Actionability test:** +- If this metric drops 20%, what specific action does the team take? +- If this metric rises 20%, what does it tell you to do more of? + +### 3. Owned + +Exactly one team or role is accountable for the metric's definition, collection, interpretation, and accuracy. Shared ownership is no ownership. + +**Ownership fields:** +- **Definition owner** — who decides what this metric means and when to change the definition. +- **Instrumentation owner** — who ensures the event fires correctly. +- **Consumer** — who uses this metric to make decisions. +- **On-call** — who gets paged when the data looks wrong. + +### 4. Contextual + +The metric's target, threshold, and interpretation depend on product type, market, stage, and business model. No universal benchmark applies. + +**Context label format:** `[Product: | Market: | Stage: | Model: ]` + +Example: `[Product: B2B SaaS | Market: mid-market | Stage: growth | Model: subscription]` + +## Leading vs Lagging Indicators + +| Property | Leading indicator | Lagging indicator | +|----------|------------------|-------------------| +| Timing | Changes before the outcome moves | Changes after the outcome moves | +| Sensitivity | High — responds quickly to product changes | Low — takes time to reflect changes | +| Use | Day-to-day decision making, experiment readouts | Strategic review, board reporting | +| Example | Feature adoption rate, activation completion | Quarterly revenue, annual retention | +| Risk | Noisy, may false-signal | Slow, may confirm a problem too late | + +A healthy metric tree has both: leading indicators for fast feedback and lagging indicators for confirmation. + +## Countermetrics (Guardrails) + +Every metric that drives a decision must be paired with at least one countermetric — a metric that would reveal an adverse side effect of optimizing for the primary metric. + +**Countermetric selection rules:** +1. The countermetric must detect a real, plausible harm — not a theoretical edge case. +2. The countermetric must have a defined threshold at which the primary metric optimization is paused or re-evaluated. +3. The countermetric must be owned by a different team than the primary metric when the primary metric's optimization creates a conflict of interest. + +**Example pairs:** + +| Primary metric | Countermetric | What it guards against | +|---------------|---------------|----------------------| +| New user signups | 30-day retention | Growth hacking that attracts users who churn | +| Feature adoption rate | Support ticket volume | Pushing adoption before the feature is stable | +| Time-to-value (decrease) | Activation quality score | Rushing onboarding at the expense of comprehension | +| Revenue per user (increase) | NPS or satisfaction | Monetization that degrades experience | +| Messages sent (increase) | Unsubscribe rate | Engagement maximization that drives users away | + +## Measurability Gate + +Before committing a metric to the tree, it must pass a measurability gate. A metric that fails is either rejected (with a named constraint) or marked as aspirational (tracked separately from operational metrics). + +### Measurability assessment + +| Question | Pass condition | +|----------|---------------| +| Is the data source identified? | A specific event, database table, or API is named. | +| Is the data source currently instrumented? | The data exists in production, or a feasible instrumentation plan exists. | +| Are identity and session semantics defined? | For user-level metrics: how users are identified across sessions, devices, and auth states is specified. | +| Can the metric be computed end-to-end? | The full pipeline from raw event to dashboard is traceable. | +| Is the metric stable under reasonable data quality issues? | Late-arriving data, duplicates, and nulls have defined handling. | + +### Unmeasurable Metric Handling + +When a metric fails the measurability gate: + +1. **Name the constraint** — what specifically is missing (e.g., "no event fires when a user completes onboarding step 3," "identity resolution fails for users who sign in with SSO after starting anonymously"). +2. **Do not fabricate a proxy** — do not substitute a different metric that "seems close." Record the gap honestly. +3. **If a proxy is explicitly requested**, define it with: (a) the exact formula, (b) the data source, (c) the known biases and gaps compared to the ideal metric, and (d) a plan to close the gap. + +## Metric Tree Maintenance + +The metric tree is a living artifact, not a one-time exercise: + +- **Quarterly review** — is the North Star still the right primary outcome? Are all metrics still observable and actionable? +- **Metric deprecation** — when a metric is no longer used for decisions, retire it. Remove it from dashboards and stop instrumenting it. +- **Ownership refresh** — when teams reorganize, reassign metric ownership. An unowned metric is untrustworthy. +- **Countermetric audit** — verify that every decision-driving metric still has a valid countermetric and that countermetric thresholds have not been silently ignored. + +## Examples (Contextual) + +### Example 1: B2B SaaS Collaboration Tool +**Context:** `[Product: B2B SaaS | Market: mid-market | Stage: growth | Model: subscription]` + +``` +North Star: Weekly Active Teams +├── Team Activation (leading) +│ ├── % teams with ≥3 members completing core action within 7 days +│ ├── Time-to-first-collaboration (days) +│ └── Countermetric: Support tickets per activated team +├── Team Engagement Depth (leading) +│ ├── Avg collaborative actions per team per week +│ ├── % teams using ≥2 features +│ └── Countermetric: Feature confusion rate (clicks-to-undo ratio) +└── Team Retention (lagging) + ├── 90-day team retention rate + ├── Expansion revenue per retained team + └── Countermetric: Contraction rate (% teams downgrading) +``` + +### Example 2: Internal Developer Platform +**Context:** `[Product: Internal tool | Market: enterprise-internal | Stage: adoption | Model: productivity]` + +``` +North Star: Developer Time Saved +├── Onboarding Speed (leading) +│ ├── Time-to-first-deploy for new service (minutes) +│ ├── % services using platform defaults vs custom config +│ └── Countermetric: Deployment failure rate +├── Daily Workflow Efficiency (leading) +│ ├── Median CI pipeline duration (minutes) +│ ├── % builds that pass on first attempt +│ └── Countermetric: Build queue wait time +└── Platform Satisfaction (lagging) + ├── Developer NPS (quarterly survey) + ├── % teams voluntarily on platform (not mandated) + └── Countermetric: Shadow IT incidents (services deployed outside platform) +``` + +### Example 3: Public Digital Service +**Context:** `[Product: Public service | Market: government-citizen | Stage: mature | Model: service-delivery]` + +``` +North Star: Successful Outcome Rate +├── Access (leading) +│ ├── % applications started that are completed +│ ├── Completion rate by device type (mobile vs desktop) +│ └── Countermetric: Support call volume (calls indicate access failure) +├── Timeliness (leading) +│ ├── Median time-to-decision (days) +│ ├── % decisions within service standard +│ └── Countermetric: Error rate on expedited applications +└── Equity (lagging) + ├── Outcome rate by demographic segment + ├── % eligible population using the service + └── Countermetric: Appeals rate (indicates incorrect decisions) +``` diff --git a/product-analytics-and-measurement/templates/instrumentation-qa-checklist.md b/product-analytics-and-measurement/templates/instrumentation-qa-checklist.md new file mode 100644 index 0000000..1b8c773 --- /dev/null +++ b/product-analytics-and-measurement/templates/instrumentation-qa-checklist.md @@ -0,0 +1,76 @@ +# Instrumentation QA Checklist + +Use this checklist to verify that event instrumentation produces trustworthy data before any metric computed from those events appears on a dashboard or feeds a decision. Run the checklist for every new event and for existing events after any change to the instrumentation, pipeline, or metric definition. + +## Checklist Header + +| Field | Value | +|-------|-------| +| **Event(s) under test** | _[fill: event name(s)]_ | +| **QA owner** | _[fill: person running the QA]_ | +| **Date** | _[fill: date]_ | +| **App version** | _[fill: version that includes this instrumentation]_ | +| **Environment** | _[fill: staging / production / both]_ | + +## 1. Client-Side Verification + +Verify that events fire correctly at the point of instrumentation. + +- [ ] **Trigger correctness:** Event fires on the intended user action or system condition, and only on that condition. No false positives (firing when it shouldn't) or false negatives (not firing when it should). +- [ ] **Property completeness:** All required properties are present on every event occurrence. Optional properties are present when their condition is met. +- [ ] **Property accuracy:** Property values reflect the actual application state at the time of the event. No stale values from previous screen state, no default values where a real value should be. +- [ ] **No duplicate firing:** The event does not fire multiple times for a single user action. Verify across: double-click, rapid navigation, page refresh, back-button, and app background/foreground transitions. +- [ ] **Timing accuracy:** The event timestamp reflects when the action occurred, not when the event was queued or flushed. Skew between client time and server time is within acceptable bounds. +- [ ] **Offline handling:** Events generated while offline are queued and sent when connectivity returns. Order is preserved or explicitly documented as not guaranteed. + +## 2. Server-Side Verification + +Verify that events are correctly received and processed. + +- [ ] **Ingestion:** Event reaches the ingestion endpoint. HTTP 200 for valid events. Appropriate error codes for malformed events (4xx, not 5xx for client errors). +- [ ] **Schema validation:** Event structure matches the tracking plan schema. Unknown properties are handled per policy (dropped, logged, or passed through). Missing required properties trigger an alert. +- [ ] **Identity stitching:** Events from the same user across sessions, devices, and authentication states are attributed to the correct user identity. Anonymous-to-known-user transitions are handled correctly. Identity merge rules produce the expected attribution. +- [ ] **Timestamp integrity:** Server timestamp is recorded at ingestion time. Client timestamp is preserved. Late-arriving events (beyond acceptable delay) are flagged. +- [ ] **Deduplication:** Duplicate events (same event_id) are detected and handled. Exactly-once semantics or at-least-once with idempotent consumers is confirmed. + +## 3. Pipeline Verification + +Verify that events survive the data pipeline intact. + +- [ ] **Schema compatibility:** Event schema is compatible with downstream consumers (data warehouse tables, analytics models). No silent column drops or type coercion. +- [ ] **No silent drops:** Events are not dropped by pipeline filters, sampling, or rate limiting without explicit configuration. Drop rate is monitored and within SLA. +- [ ] **Latency:** End-to-end latency from event fire to availability in the analytics system is within SLA. Pipeline backlog is monitored. +- [ ] **Transformation correctness:** Any pipeline transformations (enrichment, filtering, aggregation) produce correct results. Test with known input and verify output. + +## 4. End-to-End Verification + +Verify that the metric computed from events matches reality. + +- [ ] **Manual walkthrough:** Perform a known set of actions that should produce a predictable metric value. Verify that the metric on the dashboard matches the expected value within acceptable tolerance. +- [ ] **Metric formula verification:** The SQL or analytics query that produces the metric is reviewed and confirmed to match the metric definition in the tracking plan. No off-by-one errors, incorrect join conditions, or misunderstood filter semantics. +- [ ] **Edge case coverage:** Test edge cases: new user (no history), returning user after long absence, user with very high activity volume, user who authenticates mid-session, user who opts out of tracking. +- [ ] **Dashboard reconciliation:** For metrics that appear on multiple dashboards, verify that values are consistent across dashboards or that any differences are explained by documented filter or timing differences. + +## 5. Data Quality Monitoring Setup + +Verify that ongoing data quality monitoring is configured. + +- [ ] **Null rate alert:** Alert fires when required property null rate exceeds threshold. +- [ ] **Freshness alert:** Alert fires when events are delayed beyond SLA. +- [ ] **Volume anomaly alert:** Alert fires when event volume deviates significantly from baseline. +- [ ] **Cardinality alert:** Alert fires when property cardinality exceeds expected bounds. +- [ ] **Distribution drift alert:** Alert fires when property value distribution shifts significantly (for categorical properties). + +## Summary + +| Layer | Status | Notes | +|-------|--------|-------| +| Client-side | _[fill: pass / fail / partial]_ | _[fill: issues found]_ | +| Server-side | _[fill: pass / fail / partial]_ | _[fill: issues found]_ | +| Pipeline | _[fill: pass / fail / partial]_ | _[fill: issues found]_ | +| End-to-end | _[fill: pass / fail / partial]_ | _[fill: issues found]_ | +| Monitoring | _[fill: pass / fail / partial]_ | _[fill: issues found]_ | + +**QA verdict:** _[fill: approved / conditionally approved (list conditions) / rejected]_ + +**Conditions for re-QA:** _[fill: what changes would require re-running this checklist]_ diff --git a/product-analytics-and-measurement/templates/outcome-review.md b/product-analytics-and-measurement/templates/outcome-review.md new file mode 100644 index 0000000..eabb4aa --- /dev/null +++ b/product-analytics-and-measurement/templates/outcome-review.md @@ -0,0 +1,98 @@ +# Product Outcome Review Template + +A recurring review that ties product measurements back to decisions. Run monthly (recommended) or at whatever cadence matches the product's decision rhythm. This is not a status update — it is a deliberate inspection of whether the metrics are telling the truth and whether the team is acting on what they say. + +## Review Header + +| Field | Value | +|-------|-------| +| **Product** | _[fill: product name]_ | +| **Review date** | _[fill: date]_ | +| **Review period** | _[fill: e.g. January 2026]_ | +| **Participants** | _[fill: names and roles]_ | +| **Facilitator** | _[fill: name]_ | + +## 1. North Star Health + +_[fill: Current North Star value and trend. Is it moving in the intended direction? At what rate?]_ + +| Metric | Current value | Previous period | Change | Trend | Context label | +|--------|--------------|----------------|--------|-------|---------------| +| _[North Star]_ | _[fill]_ | _[fill]_ | _[fill: absolute and %]_ | _[up/down/flat]_ | _[Product/Market/Stage/Model]_ | + +**Assessment:** _[fill: Is the North Star healthy? Is it still the right North Star?]_ + +## 2. Metric Tree Health + +_[fill: For each node in the metric tree, record current value, target, and trend.]_ + +| Metric | Current | Target | Trend | Owner | Action status | +|--------|---------|--------|-------|-------|--------------| +| _[Driver 1]_ | _[fill]_ | _[fill]_ | _[up/down/flat]_ | _[team]_ | _[on-track / needs-attention / critical]_ | +| _[Sub-driver 1a]_ | _[fill]_ | _[fill]_ | _[up/down/flat]_ | _[team]_ | _[on-track / needs-attention / critical]_ | +| ... | ... | ... | ... | ... | ... | + +**Largest positive mover:** _[fill: which metric improved most, and why?]_ + +**Largest negative mover:** _[fill: which metric declined most, and why?]_ + +## 3. Countermetric Audit + +_[fill: For each primary metric that drives a decision, verify that its countermetric has not crossed the guardrail threshold.]_ + +| Primary metric | Countermetric | Countermetric value | Threshold | Status | +|---------------|---------------|-------------------|-----------|--------| +| _[metric]_ | _[counter]_ | _[value]_ | _[threshold]_ | _[ok / warning / breached]_ | + +**Breached guardrails:** _[fill: Any countermetric that crossed its threshold. What action is taken?]_ + +## 4. Decision Log + +_[fill: What decisions were made in this review period based on metric signals? What was the outcome of decisions made in the previous review period?]_ + +| Decision | Date | Trigger metric(s) | Expected outcome | Actual outcome | Learning | +|----------|------|-------------------|-----------------|----------------|---------| +| _[decision]_ | _[date]_ | _[metrics that informed it]_ | _[expected]_ | _[actual]_ | _[what we learned]_ | + +## 5. Instrumentation Health + +_[fill: Summary of instrumentation QA status. Any events with data quality issues? Any new events added? Any events deprecated?]_ + +| Check | Status | +|-------|--------| +| New events QA'd and approved this period | _[fill: count and list]_ | +| Events with data quality alerts this period | _[fill: count and list]_ | +| Deprecated events pending removal | _[fill: count and list]_ | +| Tracking plan version | _[fill: current version]_ | + +## 6. Metric Deprecation and Retirement + +_[fill: Any metrics that are no longer used for decisions? Any dashboards that should be retired?]_ + +| Metric / Dashboard | Reason for deprecation | Retirement date | Consumer notification | +|--------------------|-----------------------|-----------------|----------------------| +| _[name]_ | _[why]_ | _[date]_ | _[teams notified]_ | + +## 7. Measurement Gaps and Requests + +_[fill: What would the team like to measure but cannot currently? What instrumentation constraints need investment to resolve?]_ + +| Gap | Constraint | Estimated effort | Priority | +|-----|-----------|-----------------|---------| +| _[desired measurement]_ | _[missing instrumentation, data source, or identity resolution]_ | _[effort]_ | _[P0-P3]_ | + +## 8. Next Period Actions + +_[fill: Concrete actions for the next review period, with owners.]_ + +| Action | Owner | Due | +|--------|-------|-----| +| _[action]_ | _[person/team]_ | _[date]_ | + +## Review Sign-Off + +| Role | Name | Date | +|------|------|------| +| Product owner | _[name]_ | _[date]_ | +| Engineering lead | _[name]_ | _[date]_ | +| Data / analytics lead | _[name]_ | _[date]_ | diff --git a/product-analytics-and-measurement/templates/tracking-plan.md b/product-analytics-and-measurement/templates/tracking-plan.md new file mode 100644 index 0000000..8c3bbed --- /dev/null +++ b/product-analytics-and-measurement/templates/tracking-plan.md @@ -0,0 +1,117 @@ +# Tracking Plan Template + +A tracking plan is the contract that makes product events trustworthy. Fill one section per event or event group. This template is product-type-agnostic: use it for SaaS, internal tools, public services, or consumer products. + +## Document Header + +| Field | Value | +|-------|-------| +| **Product** | _[fill: product name]_ | +| **Version** | _[fill: semantic version, e.g. 1.0.0]_ | +| **Last updated** | _[fill: date]_ | +| **Owner** | _[fill: team or person accountable for this plan]_ | +| **Review cadence** | _[fill: e.g. quarterly, per-release]_ | + +## Event Taxonomy + +### Naming Convention + +_[fill: Describe the naming pattern, e.g. `category_action_detail` (all lowercase, snake_case). Example: `search_query_submitted`, `checkout_payment_completed`.]_ + +### Standard Properties + +These properties are included on every event unless otherwise noted: + +| Property | Type | Required | Description | +|----------|------|----------|-------------| +| `event_id` | UUID | yes | Unique identifier for this event occurrence | +| `timestamp` | ISO 8601 | yes | When the event occurred (client time) | +| `server_timestamp` | ISO 8601 | yes | When the server received the event | +| `user_id` | string | conditional | Authenticated user identifier (null if anonymous) | +| `anonymous_id` | string | conditional | Anonymous session identifier (null if authenticated) | +| `session_id` | UUID | yes | Session identifier (see Session Definition) | +| `platform` | string | yes | `web`, `ios`, `android`, `server`, `api` | +| `app_version` | string | yes | Application version that generated the event | + +### Custom Property Rules + +_[fill: Rules for custom properties — naming, types, allowed values, required vs optional, default behaviors, deprecation process.]_ + +## Identity Resolution + +### User Identity Model + +_[fill: Describe how users are identified across sessions, devices, and authentication states.]_ + +| State | Identifier | How assigned | +|-------|-----------|-------------| +| Anonymous (pre-signup) | `anonymous_id` | Generated client-side on first visit, stored in cookie/local storage | +| Authenticated (post-login) | `user_id` | Assigned by auth system, stable across sessions | +| Merged (multiple anonymous sessions) | `user_id` + `anonymous_id` history | Server-side identity graph merges anonymous IDs when user authenticates | + +### Identity Merge Rules + +_[fill: When a user authenticates, how are past anonymous events attributed? What happens on logout? What about shared devices?]_ + +## Session Definition + +| Parameter | Value | +|-----------|-------| +| **Activity timeout** | _[fill: e.g. 30 minutes of inactivity ends a session]_ | +| **Absolute timeout** | _[fill: e.g. session ends after 4 hours regardless of activity]_ | +| **Midnight rollover** | _[fill: yes/no — does a new session start at midnight?]_ | +| **Cross-platform continuity** | _[fill: are web and mobile sessions linked? How?]_ | +| **Background/foreground** | _[fill: for mobile — does backgrounding end a session? After how long?]_ | + +## Event Definitions + +_[fill: One table per event. Create as many as needed.]_ + +### Event: _[fill: event name]_ + +| Field | Value | +|-------|-------| +| **Event name** | _[fill: exact event name per naming convention]_ | +| **Description** | _[fill: what user action or system event triggers this]_ | +| **Trigger** | _[fill: exact condition — e.g. "when user clicks 'Submit' and form validation passes"]_ | +| **Frequency** | _[fill: expected volume — e.g. "~10K/day at peak"]_ | +| **Owner (definition)** | _[fill: team or person]_ | +| **Owner (instrumentation)** | _[fill: team or person]_ | +| **Consumer(s)** | _[fill: which teams/dashboards/models depend on this event]_ | + +#### Properties + +| Property | Type | Required | Allowed values / constraints | Description | +|----------|------|----------|------------------------------|-------------| +| _[fill: property name]_ | _[fill: string, integer, float, boolean, enum, JSON]_ | _[yes/no/conditional]_ | _[fill: constraints]_ | _[fill: description]_ | + +#### Data Quality Rules + +| Rule | Threshold | Action on violation | +|------|-----------|-------------------| +| Null rate | _[fill: e.g. <1% for required properties]_ | _[fill: alert, investigation]_ | +| Freshness | _[fill: e.g. events received within 5 minutes of timestamp]_ | _[fill: alert]_ | +| Cardinality | _[fill: e.g. property X has <100 distinct values]_ | _[fill: alert if exceeds]_ | +| Volume | _[fill: e.g. ±50% of 7-day rolling average]_ | _[fill: alert]_ | +| Distribution | _[fill: e.g. no value >90% of events for enum properties]_ | _[fill: investigation]_ | + +## Event Versioning and Deprecation + +| Rule | Description | +|------|-------------| +| **Additive changes** | Adding new optional properties: minor version bump, no breaking change. | +| **Breaking changes** | Removing or renaming properties, changing types: major version bump, coordinate with consumers. | +| **Deprecation notice** | Events to be removed are marked deprecated for at least one review cycle before removal. | +| **Deprecation log** | _[fill: table of deprecated events with deprecation date, sunset date, and migration path]_ | + +## Privacy and Consent + +| Field | Value | +|-------|-------| +| **Consent boundary** | _[fill: which events require user consent before firing]_ | +| **Pre-consent events** | _[fill: which events are permitted before consent (must be minimal)]_ | +| **Opt-out handling** | _[fill: how events are handled when a user opts out of tracking]_ | +| **Data retention** | _[fill: how long raw events are retained]_ | +| **Deletion handling** | _[fill: how measurement continues when a user requests data deletion]_ | +| **Aggregation minimum** | _[fill: minimum cohort size for reporting — metrics with smaller cohorts are suppressed]_ | +| **Jurisdictional notes** | _[fill: GDPR, CCPA, or other regulatory considerations]_ | diff --git a/references/skill-triggers.md b/references/skill-triggers.md index dd1052c..9a6009b 100644 --- a/references/skill-triggers.md +++ b/references/skill-triggers.md @@ -51,6 +51,7 @@ Each skill's `description` field is the canonical routing contract. This conveni | "agent eval", "agent evaluation", "LLM eval", "LLM evals", "evaluation dataset", "grader calibration", "model judge", "trajectory review", "agent observability", "agent traces", "agent telemetry", "agent regression", "agent release gate", "privacy-aware telemetry", "prompt evaluation", "model testing" | [agent-evals-and-observability](../agent-evals-and-observability/SKILL.md) | | "pydanticai", "pydantic AI", "pydantic graph", "AI agent", "LLM agent", "agent framework", "function tool", "tool-using agent", "agent with tools", "agent with dependencies", "structured output", "streaming agent", "agent graph", "state machine graph", "GraphBuilder", "BaseNode", "multi-agent", "agent delegation", "TestModel", "FunctionModel", "capabilities" | [pydanticai](../pydanticai/SKILL.md) | | "product discovery", "stakeholder interview", "requirements discovery", "user research", "customer interview", "requirements gathering", "discovery phase", "stakeholder mapping", "interview guide", "discovery conversation", "transcript to spec", "requirements conflict", "what would have to be true", "pre-mortem", "laddering", "assumption busting" | [product-discovery](../product-discovery/SKILL.md) | +| "product analytics", "product measurement", "metric tree", "leading indicator", "lagging indicator", "countermetric", "tracking plan", "event taxonomy", "instrumentation QA", "analytics QA", "dashboard contract", "product outcome review", "measurement governance", "privacy-aware measurement", "unmeasurable North Star", "North Star measurement", "product funnel", "cohort analysis product", "retention measurement", "adoption measurement" | [product-analytics-and-measurement](../product-analytics-and-measurement/SKILL.md) | | "product experiment", "experiment design", "A/B test design", "hypothesis test product", "feature experiment", "fake door test", "concierge test", "assumption test", "experiment brief", "experiment guardrail", "experiment readout", "ship decision", "no-ship decision", "experiment ethics", "prototype test experiment", "qualitative experiment", "stopping rule experiment", "method selection experiment", "underpowered experiment" | [product-experimentation](../product-experimentation/SKILL.md) | | "product UX", "product design", "interaction design", "information architecture", "task flow", "user flow", "state model", "recovery path", "interface contract", "UX handoff", "usability study plan" | [product-design-and-ux](../product-design-and-ux/SKILL.md) | | "prioritize", "RICE", "MoSCoW", "opportunity solution tree", "decision log", "write a spec", "product spec", "PRD", "stakeholder communication", "executive brief", "feature prioritization", "backlog ranking", "release scope", "build vs buy", "product decision" | [product-methodology](../product-methodology/SKILL.md) |