Journal and stream dedicated worker phases so Live distinguishes first-variant design and validation from remaining-direction work without adding pollable events.\n\nAI-assisted: OpenAI Codex.
Compare direct Sol execution with cold and persistent app-server paths using identical full-task quality gates, lifecycle timings, and token metrics.\n\nAI-assisted: OpenAI Codex.
Update Live Lab and the Live reference with the default Sol worker, full-task quality gate, Spark control, cold readiness, and production architecture.\n\nAI-assisted: OpenAI Codex.
Tab-strip MEMBERSHIP no longer exempts chromatic top/bottom stripes —
only a real selection marker does: aria-selected="true", aria-current
(any non-false value), or an active/current/selected class hint. A
stripe repeated on every tab in the group ([role=tab], .tabs items,
aria-selected="false" tabs) is decoration and flags as side-tab; the
selected tab's own underline — including the reserved-space
transparent-border pattern — stays legal. Applied consistently across
the element border path (isTabContextElement), the pseudo-element
stripe scan, and the inset box-shadow stripe scan.
Also replaces a stray NUL byte in the marquee scanner's dedupe key
that made tools treat checks.mjs as binary.
Browser bundle regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Route the full-context benchmark through production worker inputs and preserve established shared-control visual roles during variant amplification.\n\nAI-assisted: OpenAI Codex.
Default Codex to a dedicated Sol/medium app-server worker with native skill and image inputs, inherited project context, bounded source neighborhood evidence, and progressive context refresh. Other harnesses retain the portable foreground path.\n\nAI-assisted: OpenAI Codex.
Benchmark realistic bolder and polish tasks across fast, full-model, and full-context worker profiles with deterministic and independent quality gates.\n\nAI-assisted: OpenAI Codex.
Introduce a Live-owned app-server supervisor with progressive fenced publishing, partitioned control polling, cancellation and recovery safety, and measured integration coverage.
AI-assisted implementation under maintainer direction.
Four gaps from human review of gpt-5.6 eval artifacts:
1. codex-grid-background variants: the block scan now also matches the
inverted end-of-tile hairline form (transparent calc(100% - Npx))
and reads the tile cell from the background shorthand's `/ Npx Npx`
slot, not just background-size declarations. A single hairline layer
qualifies when tiled by a px pair cell (page-scale line field);
percent-tiled single hairlines (background-size: 25% 100% rules on
data-viz tracks/graphs) stay legal.
2. hero-eyebrow-chip branch C (dash-prefix): sentence-case, regular-
weight microlabels above the h1 announced by a short chromatic
::before/::after bar (8-80px x 1-6px, accent fill). Static cascade
marks dash-pseudo targets during rule collection; the browser path
reads getComputedStyle(el, '::before'/'::after').
3. New `marquee` slop rule: <marquee> elements, and infinite animations
bound to keyframes with >= 20 percentage points of X travel. Percent
travel only — px-travel loops are bespoke product animations
(waveform playheads, progress sweeps). Centered elements animating
other properties (constant -50% X), non-infinite slide-ins, rotations,
and pulses never qualify.
4. side-tab inset box-shadow variant: single-edge inset shadows
(3-12px offset on one axis, no blur/spread, chromatic) drawn as
stripes on cards/badges/menu items. Selection-state indicators
([aria-current], [aria-selected], [role=tab], active/current/selected
hints, interaction states) stay exempt; the same stripe repeated
unconditionally on every item flags. Narrow fixed-width glyphs
(logo marks) are exempt. isTabContextElement narrowed to match:
bare nav ancestry no longer blanket-exempts top/bottom border
stripes — only explicit tab semantics or state markers do.
Browser bundle regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript evidence (a12 01-observability): plans commit and deliver on
the axes with contract-strength language (palette, type, even theme
inversion) and stay default on the axis without one (layout gets a
single conventional breath). And plans living in invisible reasoning
means nothing can hold a build to its intent. The direction is now
written as a comment block at the top of the artifact answering: the
concept, the hour-later memory, why not the modal competitor page, the
signature, the first viewport's move. Critics and evals can score
delivery-against-contract; a mood is not an answer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four changes driven by human design review of eval artifacts:
1. Static engine contrast fidelity (nav-CTA cascade miss):
- parseAnyColor evaluates color-mix() (premultiplied sRGB mix; exact
for the dominant `color-mix(in oklab, C n%, transparent)` chip form)
- extractStaticColor captures color-mix() balanced instead of plucking
"transparent" out of the expression
- resolveBackground composites translucent layers over the opaque base
in both engines instead of skipping (static) or returning them
as-if-opaque (browser)
- NEW hover pass in the static cascade: :hover rules are matched via
state-stripped selectors, merged per-property against the resting
cascade with real specificity, and checked for WCAG contrast on
styled controls (checkHoverContrast). Catches the recurring miss
where a broader selector (.nav-links a:hover) beats the CTA's own
hover color and drops the pair below AA.
2. New `radial-halo` slop rule: chromatic radial-gradient wash (visible
saturated center -> transparent) as a decorative background on a dark
page. Exempts achromatic vignettes, opaque-end sheens, px-stop dot
textures, url() photo layers, and translucent (<0.7 alpha) staged-
light washes. Separate id from dark-glow so dashboards track the
gradient-drawn variant independently.
3. side-tab horizontal variant: 3-12px chromatic border-top/bottom (and
top/bottom-anchored full-width pseudo stripes) on cards/badges flag as
side-tab. Exempt: tablist/nav/aria-selected underlines, link/button
affordances, table cells, hr, state-conditional pseudo stripes, and
>12px bands. Badge-shaped spans (own visible background) participate.
4. CLI: file:// URLs route to the Puppeteer browser engine (~2s on a
50KB page), and detect --json findings now carry the registry
`category` field so downstream QA loops can separate mechanical slop
tells from judgment calls.
Fixture policy update: flat 3px top-accent cards moved from should-pass
to flag columns; tablist-underline and 16px-band pass cases added.
Browser bundle regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measure framework and provider latency, enforce fidelity and cleanup gates, and publish reproducible results on the dev-only Live Lab.\n\nAI-assisted: OpenAI Codex.
Paul's a10 review: palettes are refreshed (the palette-exclusivity
line's fingerprint) while layouts stay boring in every version. Same
cure, same shape: the layout has exactly two legitimate sources, the
concept or the content's own structure; the category's habitual
skeleton is neither.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's a11 review: heroes are safe SaaS viewports, everything
predictable; mobile Operate ships dark despite a brief that specifies
outdoors-in-motion use. Decide-then-build now opens with three
one-line directions differing in concept (the instinctive pick that
any studio would reach for is the default wearing your name); the
Operate mode adds: the usage scene is part of the spec, the theme
follows the scene, not the category's habit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per Paul: rather than gating a second file, fold what made the craft
path superior into the file both models already read 21/21 through the
gate. new-work.md gains 'Decide, then build' (direction as one
confirmable paragraph; attended pauses, unattended records-and-goes;
codex.md mock flow when image generation exists) and 'Finish like a
studio' (inspect, honest critique, patch, detector). craft becomes a
deprecated alias like teach: invoking it forces attended checkpoints,
nothing else differs; the reference is a redirect stub. codex.md
retargeted. Existing-world feature builds remain governed by the core
floor (unmeasured path, noted in the plan doc).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Invocation A/B on Fable (a9 craft-path vs a9-direct plain): the plain
path scored 38% vs the competitor against the craft path's 50%, and
brief fidelity collapsed to 14% vs bare — the direct path drops asked-
for features that craft's direction step and engineering bar preserve.
Routing now sends any build request through the craft orchestration
unprompted (its gates pause only when a user can respond), and the
craft floor gains a brief-coverage recheck: every requirement the brief
names must exist on the page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two gaps surfaced by human eval review of real artifacts:
1. side-tab missed the pseudo-element variant. The accent stripe drawn as
an absolutely-positioned ::before/::after (left/right: 0, top+bottom: 0
or height: 100%, narrow width, colored background) uses no border
property at all, so neither the element-level border checks (pseudo
elements never enter the static cascade or DOM walk) nor the
border-left/right regexes could see it. New scanCssTextForPseudoStripe
scans stylesheet text for that shape, mirroring the border rule's
gates: >= 3px thick (<= 12px), chromatic fill (var()-resolved, neutral
dividers skipped), full height against a side edge, with the
blockquote/prose exemptions preserved.
2. New pulsing-dot rule (slop): small circular "live" indicator dots
(<= 16px, border-radius >= 40% or pill values) bound to an infinite
animation whose keyframes vary opacity, scale, or box-shadow — or
pulse/blink/ping names when the keyframes aren't in the scanned text —
plus the Tailwind animate-ping/pulse + rounded-full + tiny-size utility
combo. Rotation-only keyframes (spinners) never flag, including when
they hide behind a pulse-like name.
Both scanners live in checkHtmlPatterns, so the static-html engine and
the browser bundle share the same detection path. Browser/extension
bundles regenerated; docs rule count bumped to 47.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
a7 transcript evidence: 01-observability samples drew orange-honey and
green seeds, recited the color-strategy menu, and shipped dark
category-reflex palettes anyway; the model applied the subject's
workmanlike grammar to its own landing page. Two generic lines: the
mode belongs to the surface, not the subject (a landing page for a
dense tool is still Persuade; deciding a page can be plain because its
subject is workmanlike is the category error in reverse), and the
palette has exactly two legitimate sources (seed or the subject's
world; the category's habitual palette is neither).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fix batch from the visitor-mode bias audit. The skill's four modes
(Persuade / Operate / Read / Experience) now reach the places that were
still hard-coded to a SaaS-marketing default:
- palette.mjs: rewrote 45 seed blurbs in material/world terms. The 29
tech-tool-world moods (13 Linear-indigo variants, 6 Figma-era, 5
climate-tech, 3 fintech, 2 Glossier DTC, incl. seed-201's docs-page
CTA red) lose all company names and product-category words; Aesop
trimmed from 17 blurbs to 4 and Klim from 7 to 4, excess rewritten
as unnamed material terms. Also carries the earlier bg-block rewrite
(brand refs out of the composition doc).
- init.md: register explainer now names the four modes and the family
each belongs to (stored value stays brand/product for compatibility);
Conversion & proof interview + PRODUCT.md section gated to Persuade
surfaces only (Experience/Read get no CTA/belief-ladder/proof).
- critique.md: Nielsen heuristics 7 and 10 may score n/a on Persuade
and Experience surfaces, total renormalized to the applicable max,
snapshot records which were n/a; working-memory examples diversified
beyond dashboard/pricing anatomy.
- Register headers in bolder/delight/quieter/colorize/layout/animate/
typeset renamed from Brand:/Product: to Persuade + Experience: /
Operate + Read:; typeset and layout gain one Read-specific sentence
(steady reading measure; navigable linearity).
- animate.md: plan checklist and implementation order lead with
feedback and transitions; the single entrance moment comes after,
scoped to modes that invite it.
- codex.md: mock inventory says "primary-action treatment (when the
surface has one)" instead of assuming a CTA.
- delight.md: loading/empty-state/console-egg examples diversified
beyond SaaS; streaks/badges scoped to Operate surfaces with
recurring tasks.
- distill.md: step-removal and next-action lines neutralized away
from signup/checkout/CTA vocabulary.
- document.md: canonical button label GET STARTED -> SAVE CHANGES;
signature components gain a non-marketing example.
- antipatterns registry: single-font rule renamed to "Single font
without hierarchy" with a description that permits one family when
weight/size contrast carries hierarchy.
Staged provider copies regenerated via build:skills:release for the
touched files only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul: the mode-governs section was de-biasing a persuade-tinted
playbook rather than writing neutral prose, the exact compensating-
paragraph anti-pattern. Rewritten: the corrective section is gone; the
first-viewport thesis speaks of the concept doing its job (the work,
the product, the content, the task); everything-bold's form list
includes the exact-system form natively; prove-don't-claim covers
content delivering; type guidance is parameterized by mode in one
sentence. Net shorter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's gallery check of the a8 docs run: skill-on still SaaS-ified the
documentation page. The playbook was persuade-flavored end to end, so
gating a greenfield Read surface through it risked amplifying exactly
that. New leading section: on Operate and Read surfaces boldness means
a committed system (typographic voice, spacing rhythm, one owned
accent, inevitable structure), the thesis is the content or the task
itself, and nothing invented may stand between the visitor and what
they came to do.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The when-to-choose guidance sat inside the file that only loads after
the choice is made. SKILL.md's routing now says it: bare build requests
build directly through the gate and floor; craft is routed only when
named or when the user asks for a guided, checkpointed build. The
Commands row describes craft by its checkpoints. craft.md's intro just
describes the supervised flow it orchestrates.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Stubs removed per Paul (register: values remain harmless family hints;
nothing points at the files anymore). craft.md now opens by defining
itself against plain invocation: a bare build request goes straight
through the gate and the craft floor; craft is the supervised path with
guaranteed checkpoints and the mock pipeline. One shipping-discipline
line joins the core floor (real content, interaction states, respect
the build pipeline) so one-shots inherit the bar that previously lived
only in craft's Step 4.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Answering the obvious question the family-depth framing dodged: with
modes derived per task, files named for the old two-register taxonomy
had no architectural reason to exist. brand.md's surviving depth (lane
test + inverse test, reflex-reject lanes, color discipline, layout
moves, permissions) folds into new-work.md, where all of it belonged:
it is new-identity Persuade/Experience guidance. product.md's content
moves unchanged to operate.md, its true name. Both old files remain as
one-line redirect stubs because register: brand|product in existing
PRODUCT.md files and older links point there. All cross-references
retargeted (SKILL.md modes intro, context.mjs REGISTER hint, live.md,
typeset.md); 85 tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Overlap audit after the new-work split. brand.md slims to family depth
that exists nowhere else (aesthetic-lane tests, named-reference
discipline, brand layout moves and permissions); everything it
duplicated against new-work.md and the core (font procedure, reject
list, color strategy, imagery, scale/leading) is deleted, killing the
two-copies-drift hazard. product.md keeps its Operate depth nearly
intact (it was not duplicated) and gains a scope note covering Read
surfaces. craft.md becomes pure orchestration: gates, foundation,
shape handoff, image-gen flow, engineering bar, iterate, present;
its duplicated design guidance (imagery rules, visual-craft bullets,
mandatory reference reads) is replaced by pointers to SKILL.md's
craft floor and new-work.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Architecture per Paul: impeccable is primarily a daily driver on
existing codebases; the always-loaded core should serve that 90% path,
not carry the full generative arsenal on every invocation. SKILL.md now
holds brief-wins, existing-worlds (the headline path), the four visitor
modes, the full craft floor, and a hard gate: new identity work
(greenfield, or a redesign discarding the current look) MUST read
reference/new-work.md before any design decision. That file carries the
generative playbook (seed, subject grounding, plan/self-check/signature,
hero-thesis, everything-bold, prove-don't-claim, color commitment,
calibration, persuade type/imagery). context.mjs enforces the gate
mechanically: NEW_WORK directive when no PRODUCT.md/DESIGN.md exists,
and the old mandatory register-file read is replaced by a REGISTER
family hint. No surfaces: map anywhere; mode is derived per task.
Gate compliance is measurable via skillEvidence.directSkillFileReads.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Field report: impeccable SaaS-ified a developer docs page; the Opus
galleries showed the same on an album page. Root cause: two registers
force every surface into persuade-or-operate grammar. The register
section now names the visitor's mode first (Persuade / Operate / Read /
Experience) with mode-borrowing called out as the canonical failure,
and PRODUCT.md's register field maps as family (brand = Persuade +
Experience, product = Operate + Read) for compatibility. Read mode:
comprehension deliverable, navigable structure, chrome out of the way.
Experience mode: the artifact leads at every screen size.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's Opus gallery observation: every impeccable 05-experimental-album
generation reads decidedly SaaS while frontend-design's open with the
art itself, especially at narrow viewports. Cause: the brand register
prescribed stop-the-scroll/earn-the-click/convert for ALL brand
surfaces. Split the register's deliverable by surface: product/service
pages convert; cultural surfaces (album, portfolio, publication, body
of work) lead with the artifact, recede the interface, and treat
conversion grammar as a category error — the visitor meets the work in
the first viewport at every screen size.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul: everything should be bold, nothing bland; bold is neither
decoration nor clutter but commitment to the concept, whose form the
concept chooses (maximal or severely clean, drenched or monochrome,
piercing copy, the product demonstrating itself). Replaces the
'spend your boldness in one place' rule imported from frontend-design,
whose one-bold-element-on-a-quiet-page framing pulled pages toward the
tasteful softness the galleries showed losing. The signature becomes
where the concept peaks rather than the only place it lives.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's spot-check of the Fable validation galleries: frontend-design's
lektor generations read vastly more distinctive and subject-faithful
despite losing the overall pairwise verdict on craft. The arena agrees
on the axis (distinctiveness 8-31 at n=5). Two additions to the core:
the opening viewport is a thesis (open with the most characteristic
thing in the subject's world, with a concrete memory test), and an
explicit polish-is-the-floor counterweight so the craft floor stops
reading as a mandate for quiet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Eval evidence showed the per-edit PostToolUse stream fires overwhelmingly
on copy-level rules (em-dash-overuse ~97x/session) and measurably makes
models more conservative, while a full-detector pass at completion is what
actually fixes contrast/padding/glow. Split the hook accordingly:
- Per-edit (PostToolUse) now surfaces only IMMEDIATE_TIER_RULES: broken
output (broken-image, text-overflow, clipped-overflow-container,
body-text-viewport-edge), objective contrast/legibility failures
(low-contrast, gray-on-color, tiny-text), single-property mechanical
slop (gradient-text, dark-glow), and design-system drift (the four
design-system-* rules, which compound if left uncorrected). Everything
else defers. Override with hook.perEditRules: "all" in
.impeccable/config.json. Tiering is off for Cursor/Copilot harnesses,
which have no Stop pass wired, so nothing gets silently dropped there.
- Stop deep pass (runStopHook): runs the FULL rule set over every UI file
touched this session (tracked via the existing hook.cache.json session
state; deferred-only edits now mark the file touched), dedupes against
everything already surfaced per-edit, honors ignore-rule/file/value and
inline disables, reuses the [impeccable@1] envelope, and no-ops fast
when no UI files were touched. Emits hookSpecificOutput
{ hookEventName: "Stop", additionalContext } per the Claude Code SDK
Stop contract (conversation continues so the model can act on it).
Second Stop fire is silent - deep-pass findings are remembered.
- Wiring: Stop entries (timeout 30) in plugin/hooks/hooks.json, the
.claude settings + .codex hooks manifests (transformers + hook-admin
repair path). Claude Code and Codex both dispatch a native Stop event;
Cursor's stop hook is inconsistently dispatched (pre-write gate stays)
and Copilot's agentStop/sessionEnd don't inject model context, so
neither gets a Stop entry - documented in reference/hooks.md.
- Tests: tiering split/override/harness gating, Stop dedupe + silent
no-touched-files + ignore machinery + kill switches; existing per-edit
tests moved to immediate-tier rule ids. 181 tests green; smoke-tested
the built dist skill end to end (glow surfaced per-edit, em-dash only
at Stop, second Stop silent).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Mechanical build:skills:release output; the source change was already
committed but the tracked staged copy had not been re-synced.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
gpt-5.6-sol evals: skill-on lost craft 0-25 to bare gpt-5.6; removing
the enumerated codex ban block recovered it to 4-16, confirming the
block's literal CSS patterns self-prime the defects they ban (the same
mechanism the v2.1 ablation sweep documented). Replaced with three
shape-level calibration lines: tracking floor (kept, it's a numeric
ceiling), elevation-declared-once + modest container radius, and
material honesty (real assets, surfaces not decoration, specific
claims). Detector rules continue to enforce the mechanical patterns.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Pairwise evals on Fable one-shot (6-task regression set, opus-4-8 judge,
position-bias-cancelled): the hand-distilled ~55-line lean core beat the
heavy v4 core 66% overall / 67% craft head-to-head, and moved the
decisive win-rate vs frontend-design from 13% to 27% (40% with the
completion-time QA scan; craft went positive 6-5 for the first time).
18/18 lean samples ran context.mjs + palette.mjs vs a minority under the
heavy core: shorter instructions get followed. Context weight itself was
suppressing both compliance and boldness.
Structure: persona + brief-wins + existing-worlds + subject-grounding +
plan/self-check + boldness + prove-don't-claim + commit + calibration +
compressed craft floor + two-paragraph registers. Commands table kept;
the no-arg context-aware menu logic moved to reference/routing.md (read
on demand in the only case that is inherently interactive). Provider
blocks and rule anchors preserved.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Judge rationales across cand-v4a2 arenas: competitor wins by showing
the product working (mix panels, comparison tables, live demos) and by
signatures big enough to organize the page; our samples claim, decorate,
and sometimes stop at the hero.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Eval evidence (cand-v4a1-prose): palette.mjs handed a random violet seed
to the Polish-TV lektor brief and the model anchored on it, overriding
subject-grounding; craft/shape user gates can't fire in one-shot runs
and each model improvises around them. Seed is now a reflex-check that
yields to a subject-dictated palette; craft/shape gain an explicit
unattended mode (same bar, no waiting); init interview is skipped when
no user can respond.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per Paul's guidance: (1) existing committed design systems are the
bread-and-butter case and get a first-class core rule (work inside the
world, no parallel colors/fonts/styles, no perf regressions); (2) a
redesign that discards the current look is new identity work and runs
the full concept/tokens/signature process instead of anchoring to the
incumbent skeleton (the lektor failure); (3) the reflex-reject font list
and physical-object font procedure are brand-register rules, moved out
of the universal Commit section — system stacks and workhorse UI faces
are legitimate, often correct, for product UI, stated positively in the
product register.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>