Add repeated production annotation timing, real app-server cache telemetry, the context-delta decision, and correct --judge=false handling.\n\nAI-assisted: OpenAI Codex.
Keep short crash-recovery leases without allowing a healthy worker to queue its own generation twice, and surface non-monotonic benchmark journals as errors.\n\nAI-assisted: OpenAI Codex.
Replaces the a14/a15 attempts (both deleted). Diagnosis: incentive
stacking; the placeholder-completion MUST plus the image tool turned
'bolder' into full-bleed photo insertion. Scope preservation is the
missing rule, not imagery policy.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Return after a durable starting record, overlap app-server initialization with page startup, dynamically reclaim generation after worker failure, and cap hard-crash leases at 15 seconds.\n\nAI-assisted: OpenAI Codex.
x02 a14 rerun: 3/3 samples still imported photos — the unscoped MUST in
the Persuade mode block overrode the existing-worlds principle. Scoping
keeps the greenfield ablation win, frees iteration asks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Remove stale and contradictory experiments, order the surviving decisions by impact, and replace synthetic claims with current production evidence.\n\nAI-assisted: OpenAI Codex.
Replace stale startup and synthetic claims with production browser timing, matched architecture comparisons, Accept latency, and honest run counts.\n\nAI-assisted: OpenAI Codex.
x02-tidewater-bolder eval: 3/3 skill-on samples imported photography into
a photo-free seed system (0% arena vs competitor, which amplified the
seed's own vocabulary instead). One sentence, shape-level, no examples.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Count variants only inside the active generation wrapper so deferred carbonize markers cannot create false fast-path results.\n\nAI-assisted: OpenAI Codex.
Exercise accepting the first progressive variant, immediately preparing another task, and receiving its first result through the independent Codex worker.\n\nAI-assisted: OpenAI Codex.
Let the browser E2E harness launch an independent production worker, exercise real sub-command selection, and carry realistic product/design context.\n\nAI-assisted: OpenAI Codex.
Validate and transactionally publish complete structured agent messages as soon as they arrive while retaining turn-completion serialization for subsequent phases.\n\nAI-assisted: OpenAI Codex.
Journal and stream dedicated worker phases so Live distinguishes first-variant design and validation from remaining-direction work without adding pollable events.\n\nAI-assisted: OpenAI Codex.
Compare direct Sol execution with cold and persistent app-server paths using identical full-task quality gates, lifecycle timings, and token metrics.\n\nAI-assisted: OpenAI Codex.
Update Live Lab and the Live reference with the default Sol worker, full-task quality gate, Spark control, cold readiness, and production architecture.\n\nAI-assisted: OpenAI Codex.
Tab-strip MEMBERSHIP no longer exempts chromatic top/bottom stripes —
only a real selection marker does: aria-selected="true", aria-current
(any non-false value), or an active/current/selected class hint. A
stripe repeated on every tab in the group ([role=tab], .tabs items,
aria-selected="false" tabs) is decoration and flags as side-tab; the
selected tab's own underline — including the reserved-space
transparent-border pattern — stays legal. Applied consistently across
the element border path (isTabContextElement), the pseudo-element
stripe scan, and the inset box-shadow stripe scan.
Also replaces a stray NUL byte in the marquee scanner's dedupe key
that made tools treat checks.mjs as binary.
Browser bundle regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Route the full-context benchmark through production worker inputs and preserve established shared-control visual roles during variant amplification.\n\nAI-assisted: OpenAI Codex.
Default Codex to a dedicated Sol/medium app-server worker with native skill and image inputs, inherited project context, bounded source neighborhood evidence, and progressive context refresh. Other harnesses retain the portable foreground path.\n\nAI-assisted: OpenAI Codex.
Benchmark realistic bolder and polish tasks across fast, full-model, and full-context worker profiles with deterministic and independent quality gates.\n\nAI-assisted: OpenAI Codex.
Introduce a Live-owned app-server supervisor with progressive fenced publishing, partitioned control polling, cancellation and recovery safety, and measured integration coverage.
AI-assisted implementation under maintainer direction.
Four gaps from human review of gpt-5.6 eval artifacts:
1. codex-grid-background variants: the block scan now also matches the
inverted end-of-tile hairline form (transparent calc(100% - Npx))
and reads the tile cell from the background shorthand's `/ Npx Npx`
slot, not just background-size declarations. A single hairline layer
qualifies when tiled by a px pair cell (page-scale line field);
percent-tiled single hairlines (background-size: 25% 100% rules on
data-viz tracks/graphs) stay legal.
2. hero-eyebrow-chip branch C (dash-prefix): sentence-case, regular-
weight microlabels above the h1 announced by a short chromatic
::before/::after bar (8-80px x 1-6px, accent fill). Static cascade
marks dash-pseudo targets during rule collection; the browser path
reads getComputedStyle(el, '::before'/'::after').
3. New `marquee` slop rule: <marquee> elements, and infinite animations
bound to keyframes with >= 20 percentage points of X travel. Percent
travel only — px-travel loops are bespoke product animations
(waveform playheads, progress sweeps). Centered elements animating
other properties (constant -50% X), non-infinite slide-ins, rotations,
and pulses never qualify.
4. side-tab inset box-shadow variant: single-edge inset shadows
(3-12px offset on one axis, no blur/spread, chromatic) drawn as
stripes on cards/badges/menu items. Selection-state indicators
([aria-current], [aria-selected], [role=tab], active/current/selected
hints, interaction states) stay exempt; the same stripe repeated
unconditionally on every item flags. Narrow fixed-width glyphs
(logo marks) are exempt. isTabContextElement narrowed to match:
bare nav ancestry no longer blanket-exempts top/bottom border
stripes — only explicit tab semantics or state markers do.
Browser bundle regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Transcript evidence (a12 01-observability): plans commit and deliver on
the axes with contract-strength language (palette, type, even theme
inversion) and stay default on the axis without one (layout gets a
single conventional breath). And plans living in invisible reasoning
means nothing can hold a build to its intent. The direction is now
written as a comment block at the top of the artifact answering: the
concept, the hour-later memory, why not the modal competitor page, the
signature, the first viewport's move. Critics and evals can score
delivery-against-contract; a mood is not an answer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four changes driven by human design review of eval artifacts:
1. Static engine contrast fidelity (nav-CTA cascade miss):
- parseAnyColor evaluates color-mix() (premultiplied sRGB mix; exact
for the dominant `color-mix(in oklab, C n%, transparent)` chip form)
- extractStaticColor captures color-mix() balanced instead of plucking
"transparent" out of the expression
- resolveBackground composites translucent layers over the opaque base
in both engines instead of skipping (static) or returning them
as-if-opaque (browser)
- NEW hover pass in the static cascade: :hover rules are matched via
state-stripped selectors, merged per-property against the resting
cascade with real specificity, and checked for WCAG contrast on
styled controls (checkHoverContrast). Catches the recurring miss
where a broader selector (.nav-links a:hover) beats the CTA's own
hover color and drops the pair below AA.
2. New `radial-halo` slop rule: chromatic radial-gradient wash (visible
saturated center -> transparent) as a decorative background on a dark
page. Exempts achromatic vignettes, opaque-end sheens, px-stop dot
textures, url() photo layers, and translucent (<0.7 alpha) staged-
light washes. Separate id from dark-glow so dashboards track the
gradient-drawn variant independently.
3. side-tab horizontal variant: 3-12px chromatic border-top/bottom (and
top/bottom-anchored full-width pseudo stripes) on cards/badges flag as
side-tab. Exempt: tablist/nav/aria-selected underlines, link/button
affordances, table cells, hr, state-conditional pseudo stripes, and
>12px bands. Badge-shaped spans (own visible background) participate.
4. CLI: file:// URLs route to the Puppeteer browser engine (~2s on a
50KB page), and detect --json findings now carry the registry
`category` field so downstream QA loops can separate mechanical slop
tells from judgment calls.
Fixture policy update: flat 3px top-accent cards moved from should-pass
to flag columns; tablist-underline and 16px-band pass cases added.
Browser bundle regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measure framework and provider latency, enforce fidelity and cleanup gates, and publish reproducible results on the dev-only Live Lab.\n\nAI-assisted: OpenAI Codex.
Paul's a10 review: palettes are refreshed (the palette-exclusivity
line's fingerprint) while layouts stay boring in every version. Same
cure, same shape: the layout has exactly two legitimate sources, the
concept or the content's own structure; the category's habitual
skeleton is neither.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's a11 review: heroes are safe SaaS viewports, everything
predictable; mobile Operate ships dark despite a brief that specifies
outdoors-in-motion use. Decide-then-build now opens with three
one-line directions differing in concept (the instinctive pick that
any studio would reach for is the default wearing your name); the
Operate mode adds: the usage scene is part of the spec, the theme
follows the scene, not the category's habit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per Paul: rather than gating a second file, fold what made the craft
path superior into the file both models already read 21/21 through the
gate. new-work.md gains 'Decide, then build' (direction as one
confirmable paragraph; attended pauses, unattended records-and-goes;
codex.md mock flow when image generation exists) and 'Finish like a
studio' (inspect, honest critique, patch, detector). craft becomes a
deprecated alias like teach: invoking it forces attended checkpoints,
nothing else differs; the reference is a redirect stub. codex.md
retargeted. Existing-world feature builds remain governed by the core
floor (unmeasured path, noted in the plan doc).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Invocation A/B on Fable (a9 craft-path vs a9-direct plain): the plain
path scored 38% vs the competitor against the craft path's 50%, and
brief fidelity collapsed to 14% vs bare — the direct path drops asked-
for features that craft's direction step and engineering bar preserve.
Routing now sends any build request through the craft orchestration
unprompted (its gates pause only when a user can respond), and the
craft floor gains a brief-coverage recheck: every requirement the brief
names must exist on the page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two gaps surfaced by human eval review of real artifacts:
1. side-tab missed the pseudo-element variant. The accent stripe drawn as
an absolutely-positioned ::before/::after (left/right: 0, top+bottom: 0
or height: 100%, narrow width, colored background) uses no border
property at all, so neither the element-level border checks (pseudo
elements never enter the static cascade or DOM walk) nor the
border-left/right regexes could see it. New scanCssTextForPseudoStripe
scans stylesheet text for that shape, mirroring the border rule's
gates: >= 3px thick (<= 12px), chromatic fill (var()-resolved, neutral
dividers skipped), full height against a side edge, with the
blockquote/prose exemptions preserved.
2. New pulsing-dot rule (slop): small circular "live" indicator dots
(<= 16px, border-radius >= 40% or pill values) bound to an infinite
animation whose keyframes vary opacity, scale, or box-shadow — or
pulse/blink/ping names when the keyframes aren't in the scanned text —
plus the Tailwind animate-ping/pulse + rounded-full + tiny-size utility
combo. Rotation-only keyframes (spinners) never flag, including when
they hide behind a pulse-like name.
Both scanners live in checkHtmlPatterns, so the static-html engine and
the browser bundle share the same detection path. Browser/extension
bundles regenerated; docs rule count bumped to 47.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
a7 transcript evidence: 01-observability samples drew orange-honey and
green seeds, recited the color-strategy menu, and shipped dark
category-reflex palettes anyway; the model applied the subject's
workmanlike grammar to its own landing page. Two generic lines: the
mode belongs to the surface, not the subject (a landing page for a
dense tool is still Persuade; deciding a page can be plain because its
subject is workmanlike is the category error in reverse), and the
palette has exactly two legitimate sources (seed or the subject's
world; the category's habitual palette is neither).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fix batch from the visitor-mode bias audit. The skill's four modes
(Persuade / Operate / Read / Experience) now reach the places that were
still hard-coded to a SaaS-marketing default:
- palette.mjs: rewrote 45 seed blurbs in material/world terms. The 29
tech-tool-world moods (13 Linear-indigo variants, 6 Figma-era, 5
climate-tech, 3 fintech, 2 Glossier DTC, incl. seed-201's docs-page
CTA red) lose all company names and product-category words; Aesop
trimmed from 17 blurbs to 4 and Klim from 7 to 4, excess rewritten
as unnamed material terms. Also carries the earlier bg-block rewrite
(brand refs out of the composition doc).
- init.md: register explainer now names the four modes and the family
each belongs to (stored value stays brand/product for compatibility);
Conversion & proof interview + PRODUCT.md section gated to Persuade
surfaces only (Experience/Read get no CTA/belief-ladder/proof).
- critique.md: Nielsen heuristics 7 and 10 may score n/a on Persuade
and Experience surfaces, total renormalized to the applicable max,
snapshot records which were n/a; working-memory examples diversified
beyond dashboard/pricing anatomy.
- Register headers in bolder/delight/quieter/colorize/layout/animate/
typeset renamed from Brand:/Product: to Persuade + Experience: /
Operate + Read:; typeset and layout gain one Read-specific sentence
(steady reading measure; navigable linearity).
- animate.md: plan checklist and implementation order lead with
feedback and transitions; the single entrance moment comes after,
scoped to modes that invite it.
- codex.md: mock inventory says "primary-action treatment (when the
surface has one)" instead of assuming a CTA.
- delight.md: loading/empty-state/console-egg examples diversified
beyond SaaS; streaks/badges scoped to Operate surfaces with
recurring tasks.
- distill.md: step-removal and next-action lines neutralized away
from signup/checkout/CTA vocabulary.
- document.md: canonical button label GET STARTED -> SAVE CHANGES;
signature components gain a non-marketing example.
- antipatterns registry: single-font rule renamed to "Single font
without hierarchy" with a description that permits one family when
weight/size contrast carries hierarchy.
Staged provider copies regenerated via build:skills:release for the
touched files only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul: the mode-governs section was de-biasing a persuade-tinted
playbook rather than writing neutral prose, the exact compensating-
paragraph anti-pattern. Rewritten: the corrective section is gone; the
first-viewport thesis speaks of the concept doing its job (the work,
the product, the content, the task); everything-bold's form list
includes the exact-system form natively; prove-don't-claim covers
content delivering; type guidance is parameterized by mode in one
sentence. Net shorter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's gallery check of the a8 docs run: skill-on still SaaS-ified the
documentation page. The playbook was persuade-flavored end to end, so
gating a greenfield Read surface through it risked amplifying exactly
that. New leading section: on Operate and Read surfaces boldness means
a committed system (typographic voice, spacing rhythm, one owned
accent, inevitable structure), the thesis is the content or the task
itself, and nothing invented may stand between the visitor and what
they came to do.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The when-to-choose guidance sat inside the file that only loads after
the choice is made. SKILL.md's routing now says it: bare build requests
build directly through the gate and floor; craft is routed only when
named or when the user asks for a guided, checkpointed build. The
Commands row describes craft by its checkpoints. craft.md's intro just
describes the supervised flow it orchestrates.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>