The v2.1 ablation sweep (n=10 × 4 brand niches × 3 providers, anchored to
commit 54c3a502, ~544 cells) confirmed these four rules carry no weight in
the skill:
- skill-typo-no-all-caps-body — duplicate of brand-ban-all-caps-body; brand
version is more specific (reserves caps for labels + headings)
- skill-typo-codex-hero-ceiling-repeat — the codex-block restatement of
skill-typo-hero-ceiling didn't add reinforcement on top of the universal
rule
- skill-typo-scale-ratio — duplicate of brand-typo-modular-scale; same
signal, brand version carries the clamp() / fluid implementation detail
- skill-typo-font-count — models don't reach for ≥4 font families in any
niche we test, so the rule has no measurable effect
Each deletion is the Agent A / B / C / D Phase-2 audit recommendation;
none of the four ever validated under either prose state.
Adds EMPIRICAL_VALIDATION.md naming the seven cross-provider winners as the
trustworthy core, and documents the systemic findings (self-priming, detector
saturation, vocabulary anchoring) so future skill edits can avoid the same
traps.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>