Rewrite the routed bolder reference around what wins scoped
"make this section bolder" asks vs the frontend-design competitor.
Old prose was all visual levers and treated copy as secondary, so
the model kept flat placeholder copy verbatim and reached for a
decorative import for heft. New prose: scope stays sovereign;
diagnose flatness as opting out of the system's own moves; amplify
the system's own vocabulary; let content carry the weight; commit
then clarify; give the section its own scroll rhythm; a skeleton
test scoped to the section; a placeholder is a job, not a photo cue.
Drops the opening named-slop enumeration (self-priming) and the
120-line checklist (ceremony tax); now 31 lines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Detect a missing CLI before worker startup, keep Live usable through the foreground poller, and surface actionable status in Live and Live Lab.\n\nAI-assisted implementation.
Regular-use guard ahead of the realistic-lane validation: forced
creativity must never supersede user input or product context.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Opus probe (r10-opus-wireframe): control articulates loose skeletons,
wireframe arm names them (dubbing script sheet, timecode gutter spine)
and diffs against the standard stack every time. Targets Paul's
layout-diversity question: concept-atom commitment with template
skeletons underneath.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul: 'way better to have the first iteration land fully committed to
the concept, because that's the genuinely hard part. the next pass can
make sure it is clear and effective.' The check selected against the
original lektor site itself, the campaign's 10/10 reference.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The inverse-probe extracted this from Paul's NewRelic review and it won
in the batch4 forward test; it was never ported. Craft diagnosis: pages
lose on rhythm monotony (one treatment uniformly applied), not defects.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The skill's decide-then-build step opens the built HTML artifact with a
DIRECTION CONTRACT comment (UNIQUE / NOT-TEMPLATE / OWN-WORLD / STORY /
FIRST VIEWPORT / FORM). Until now nothing ever judged the finished build
against that contract; the eval harness proved sample contracts promised
radical compositions while the build shipped the standard template anyway.
The Stop deep pass now extracts the leading contract comment from each
session-touched HTML file (marker match in the first 200 chars, body
capped at 1800 chars) and appends a contract-audit section after the
detector findings: audit the render promise by promise, naming the two
observed failure shapes (a promise not in the pixels; a contract whose
own plan is the standard template wearing the concept's nouns). Zero
extra API calls; the audit rides the existing single Stop emission and
fires at most once per file per session via a contractAudited flag on
the same session cache entry the finding dedupe uses.
Ported from the eval harness reference implementation
(extractDirectionContract / composeContractAuditMessage in
impeccable-evals runner/workers/anthropic-native.ts). No hooks.json
changes needed: Claude Code and Codex both already dispatch Stop to
hook.mjs.
Tests: 163 -> 179 in tests/hook.test.mjs (extraction unit coverage plus
Stop-pass integration: present/absent/once-per-session/non-HTML/
malformed/oversized). hook-build 18/18, build:skills prose gate clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul: no stochastic challenger-assignment mode (unreproducible bad draws
= undebuggable bug reports); keep the weigh-off. Every roll now prints
its key so any field report can be replayed with --from.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contract-probe campaign findings (evals repo, notes/fable-oneshot-craft-plan.md):
a single model's resonance ranking is deterministic (30/35 identical
concepts across 16 framings); dice must come from the script, mirroring
the palette-seed result. Derived candidates stay grounded in the
audience's world + subject's cultural home; challengers win only on
identification x clarity; incumbent-with-deliberate-idea overrides the
roll. Validated at contract level on 01-observability + r10-lektor
(teletext ranks #3 for lektor; assigned index 3 produced it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four lines from the r10 dual consultation (codex gpt-5.6-sol + gemini
3.5-pro on the actual HTMLs) and the hero-probe micro-eval: the probe
isolated a first-viewport monoculture (same split template in every
sample, control and skill alike) and showed these lines break it while
codex's raw 15-liner alone does not. The incumbent sentence swap fixes
the r10 root cause both consultants independently identified.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Keep short crash-recovery leases without allowing a healthy worker to queue its own generation twice, and surface non-monotonic benchmark journals as errors.\n\nAI-assisted: OpenAI Codex.
Replaces the a14/a15 attempts (both deleted). Diagnosis: incentive
stacking; the placeholder-completion MUST plus the image tool turned
'bolder' into full-bleed photo insertion. Scope preservation is the
missing rule, not imagery policy.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Return after a durable starting record, overlap app-server initialization with page startup, dynamically reclaim generation after worker failure, and cap hard-crash leases at 15 seconds.\n\nAI-assisted: OpenAI Codex.
x02 a14 rerun: 3/3 samples still imported photos — the unscoped MUST in
the Persuade mode block overrode the existing-worlds principle. Scoping
keeps the greenfield ablation win, frees iteration asks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
x02-tidewater-bolder eval: 3/3 skill-on samples imported photography into
a photo-free seed system (0% arena vs competitor, which amplified the
seed's own vocabulary instead). One sentence, shape-level, no examples.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Validate and transactionally publish complete structured agent messages as soon as they arrive while retaining turn-completion serialization for subsequent phases.\n\nAI-assisted: OpenAI Codex.
Journal and stream dedicated worker phases so Live distinguishes first-variant design and validation from remaining-direction work without adding pollable events.\n\nAI-assisted: OpenAI Codex.
Update Live Lab and the Live reference with the default Sol worker, full-task quality gate, Spark control, cold readiness, and production architecture.\n\nAI-assisted: OpenAI Codex.
Route the full-context benchmark through production worker inputs and preserve established shared-control visual roles during variant amplification.\n\nAI-assisted: OpenAI Codex.
Default Codex to a dedicated Sol/medium app-server worker with native skill and image inputs, inherited project context, bounded source neighborhood evidence, and progressive context refresh. Other harnesses retain the portable foreground path.\n\nAI-assisted: OpenAI Codex.
Introduce a Live-owned app-server supervisor with progressive fenced publishing, partitioned control polling, cancellation and recovery safety, and measured integration coverage.
AI-assisted implementation under maintainer direction.
Transcript evidence (a12 01-observability): plans commit and deliver on
the axes with contract-strength language (palette, type, even theme
inversion) and stay default on the axis without one (layout gets a
single conventional breath). And plans living in invisible reasoning
means nothing can hold a build to its intent. The direction is now
written as a comment block at the top of the artifact answering: the
concept, the hour-later memory, why not the modal competitor page, the
signature, the first viewport's move. Critics and evals can score
delivery-against-contract; a mood is not an answer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's a10 review: palettes are refreshed (the palette-exclusivity
line's fingerprint) while layouts stay boring in every version. Same
cure, same shape: the layout has exactly two legitimate sources, the
concept or the content's own structure; the category's habitual
skeleton is neither.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Paul's a11 review: heroes are safe SaaS viewports, everything
predictable; mobile Operate ships dark despite a brief that specifies
outdoors-in-motion use. Decide-then-build now opens with three
one-line directions differing in concept (the instinctive pick that
any studio would reach for is the default wearing your name); the
Operate mode adds: the usage scene is part of the spec, the theme
follows the scene, not the category's habit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per Paul: rather than gating a second file, fold what made the craft
path superior into the file both models already read 21/21 through the
gate. new-work.md gains 'Decide, then build' (direction as one
confirmable paragraph; attended pauses, unattended records-and-goes;
codex.md mock flow when image generation exists) and 'Finish like a
studio' (inspect, honest critique, patch, detector). craft becomes a
deprecated alias like teach: invoking it forces attended checkpoints,
nothing else differs; the reference is a redirect stub. codex.md
retargeted. Existing-world feature builds remain governed by the core
floor (unmeasured path, noted in the plan doc).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Invocation A/B on Fable (a9 craft-path vs a9-direct plain): the plain
path scored 38% vs the competitor against the craft path's 50%, and
brief fidelity collapsed to 14% vs bare — the direct path drops asked-
for features that craft's direction step and engineering bar preserve.
Routing now sends any build request through the craft orchestration
unprompted (its gates pause only when a user can respond), and the
craft floor gains a brief-coverage recheck: every requirement the brief
names must exist on the page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
a7 transcript evidence: 01-observability samples drew orange-honey and
green seeds, recited the color-strategy menu, and shipped dark
category-reflex palettes anyway; the model applied the subject's
workmanlike grammar to its own landing page. Two generic lines: the
mode belongs to the surface, not the subject (a landing page for a
dense tool is still Persuade; deciding a page can be plain because its
subject is workmanlike is the category error in reverse), and the
palette has exactly two legitimate sources (seed or the subject's
world; the category's habitual palette is neither).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>