Add the product-experimentation skill (issue #190): end-to-end experiment workflow from assumption mapping through method selection, guardrail definition, and decision-readout that updates the roadmap. Includes: - SKILL.md with full 8-step workflow, method ladder (interviews through A/B tests), multi-criteria decision framework, and routing to data-scientist and release-engineering - README.md with 5 required human-facing sections - references/discovery-brief.md mapping existing experimentation guidance across product-methodology, data-scientist, release-engineering, product-design-and-ux, and financial-modeling - references/method-selection.md with decision tree and anti-patterns - references/guardrails-and-ethics.md with guardrail design, ethical boundaries, and stopping rules - references/experiment-readout.md with decision-impact field types - 4 templates: assumption-map, experiment-brief, guardrail-and-decision-rule, readout-learning-entry - evals/evals.json with 5 output-quality cases covering prototype test, feature-flag rollout, underpowered experiment, guardrail omission, and significant-but-no-ship boundary Shared updates: root README catalog entry, skill-triggers.md entry, regenerated marketplace/codex/llms catalogs. Co-authored-by: username <username> Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
4.6 KiB
Method Selection
How to choose the right experiment method for a given hypothesis. The core principle: start at the lightest-weight method that can falsify the hypothesis, and only move down the ladder when the question cannot be answered at the current level.
The Method Ladder
| Method | What it tests | Sample size | Statistical rigor | When to use | When NOT to use |
|---|---|---|---|---|---|
| Qualitative interviews | Mental models, problem validation, unknown unknowns | 5-15 | None (descriptive) | You do not understand the problem space yet; the riskiest assumption is about user behavior or motivation | You need a precise effect size estimate; the hypothesis is about a quantitative change |
| Prototype tests | Interaction flow, usability, concept desirability | 5-20 | None (observational) | The question is "can users figure this out?" or "does this concept resonate?" | You need causal attribution to a business metric; the prototype cannot simulate the real experience |
| Concierge tests | Value delivery, willingness to pay, operational feasibility | 1-20 | None (manual) | You need to test whether anyone will pay for or use the service before building automation | The manual delivery cannot approximate the automated experience closely enough |
| Fake doors | Demand signals, willingness to click/commit | 100+ | Low (conversion rate difference) | You need to size demand before building; the cost of the real feature is high | The fake door deceives users and the deception risk is unacceptable; the feature is trivial to build |
| Feature flags | Operational safety, incremental rollout, kill-switch | Configurable | Medium (controlled rollout) | You have a feature ready and need to validate it safely in production with the ability to turn it off | You have not validated the underlying assumption; a flag controls risk but does not test value |
| A/B tests | Causal attribution of a specific change to a metric | Statistical minimum (power analysis) | High (randomized controlled) | You need causal evidence that a change moves a metric; the sample size is achievable | The question can be answered with a lighter method; the sample is too small for adequate power; the ethics of randomization are questionable |
Selection Decision Tree
Can you learn enough from talking to 5-10 users?
└─ YES → Qualitative interviews
└─ NO → Can you build a clickable prototype in a day and test it with 5-20 people?
└─ YES → Prototype test
└─ NO → Can you manually deliver the value for 1-20 users?
└─ YES → Concierge test
└─ NO → Can you measure demand with a button that does not yet work?
└─ YES → Fake door
└─ NO → Do you have the feature built and need safe rollout?
└─ YES → Feature flag
└─ NO → Do you need causal attribution with statistical rigor?
└─ YES → A/B test
└─ NO → Revisit the hypothesis — is it testable?
Anti-Patterns
Defaulting to A/B testing
The most common experimentation failure. A/B testing is expensive: it requires engineering time to build the variant, statistical design (power analysis, sample-size calculation), enough traffic to detect the effect, and time to run. Before committing to an A/B test, ask: "Could I learn enough from 5 interviews to make this decision?" If yes, do the interviews.
Testing the wrong thing
If the hypothesis is "users want this feature" and you A/B test whether a blue button outperforms a green button, you are testing execution, not value. Execution testing is only useful after value is confirmed.
Ignoring method limitations
Every method has blind spots. Qualitative interviews cannot tell you how many users will convert. A/B tests cannot tell you why users behave differently. Use multiple methods together when the decision is consequential.
Measuring Method Adequacy
Before committing to a method, verify:
- Falsifiability: Can this method produce evidence that would convince you the assumption is wrong?
- Timeliness: Can you get the evidence fast enough to affect the decision?
- Cost proportionality: Is the cost of the method proportionate to the cost of being wrong?
- Ethical fit: Does the method respect user autonomy, consent, and dignity?