Commit Graph
122 Commits
Author SHA1 Message Date
Paul Bakaus c6dfd22329 Harden Live worker recovery
AI-assisted: Codex
2026-07-13 10:44:37 -07:00
Paul Bakaus a274f93c4e Add atomic Live benchmark controls
AI-assisted: Codex
2026-07-13 10:43:21 -07:00
Paul Bakaus e9121b26fe Harden Live production benchmarks and turn failures
AI-assisted: Codex
2026-07-13 10:19:42 -07:00
Paul Bakaus db63d08168 Detach canceled Live generation tails
AI-assisted: Codex
2026-07-13 10:08:27 -07:00
Paul Bakaus 2008381e91 Improve Live worker recovery and Accept latency
AI-assisted: Codex
2026-07-13 09:56:48 -07:00
Paul Bakaus 8a6c4ab486 Fix early Live choice queue ownership
AI-assisted: Codex
2026-07-13 09:54:23 -07:00
Paul Bakaus 1e4927d86a Prevent duplicate long-running Live turns
Keep short crash-recovery leases without allowing a healthy worker to queue its own generation twice, and surface non-monotonic benchmark journals as errors.\n\nAI-assisted: OpenAI Codex.
2026-07-12 21:18:20 -07:00
Paul BakausandClaude Fable 5 79573ce55b refinement scope: keep content + media footprint; recompose for emphasis (codex+gemini consult)
Replaces the a14/a15 attempts (both deleted). Diagnosis: incentive
stacking; the placeholder-completion MUST plus the image tool turned
'bolder' into full-bleed photo insertion. Scope preservation is the
missing rule, not imagery policy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 21:13:58 -07:00
Paul Bakaus 6ed43ca682 Prewarm Codex Live with safe fallback
Return after a durable starting record, overlap app-server initialization with page startup, dynamically reclaim generation after worker failure, and cap hard-crash leases at 15 seconds.\n\nAI-assisted: OpenAI Codex.
2026-07-12 21:11:50 -07:00
Paul BakausandClaude Fable 5 6f3076051f persuade: scope the imagery MUST to new surfaces; existing systems decide their own vocabulary
x02 a14 rerun: 3/3 samples still imported photos — the unscoped MUST in
the Persuade mode block overrode the existing-worlds principle. Scoping
keeps the greenfield ablation win, frees iteration asks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 21:09:58 -07:00
Paul BakausandClaude Fable 5 5674a94114 existing-worlds: boldness from committed materials; new medium = redesign, not refinement
x02-tidewater-bolder eval: 3/3 skill-on samples imported photography into
a photo-free seed system (0% arena vs competitor, which amplified the
seed's own vocabulary instead). One sentence, shape-level, no examples.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 20:59:07 -07:00
Paul Bakaus edde928736 Capture Codex worker token usage
Record per-turn app-server token notifications so Live architecture benchmarks can compare context and cache costs.\n\nAI-assisted: OpenAI Codex.
2026-07-12 20:47:01 -07:00
Paul Bakaus 395953ae1d Publish app-server output before turn completion
Validate and transactionally publish complete structured agent messages as soon as they arrive while retaining turn-completion serialization for subsequent phases.\n\nAI-assisted: OpenAI Codex.
2026-07-12 20:31:30 -07:00
Paul Bakaus 0aa6fc56a0 Show Codex Live generation progress
Journal and stream dedicated worker phases so Live distinguishes first-variant design and validation from remaining-direction work without adding pollable events.\n\nAI-assisted: OpenAI Codex.
2026-07-12 20:21:22 -07:00
Paul Bakaus 2a6f8c3f53 Publish Codex Live quality evidence
Update Live Lab and the Live reference with the default Sol worker, full-task quality gate, Spark control, cold readiness, and production architecture.\n\nAI-assisted: OpenAI Codex.
2026-07-12 20:05:28 -07:00
Paul Bakaus 5e1925f9d2 Handle empty Codex worker shutdown
Treat a never-used app-server thread with no rollout file as already archived while retaining real archive failures.\n\nAI-assisted: OpenAI Codex.
2026-07-12 20:04:19 -07:00
Paul Bakaus 421c1a93f5 Harden Live design-system fidelity
Route the full-context benchmark through production worker inputs and preserve established shared-control visual roles during variant amplification.\n\nAI-assisted: OpenAI Codex.
2026-07-12 20:02:44 -07:00
Paul Bakaus f814dd329e Enable the Codex Live quality worker
Default Codex to a dedicated Sol/medium app-server worker with native skill and image inputs, inherited project context, bounded source neighborhood evidence, and progressive context refresh. Other harnesses retain the portable foreground path.\n\nAI-assisted: OpenAI Codex.
2026-07-12 19:57:47 -07:00
Paul Bakaus e89645e69d Add experimental Codex Live worker
Introduce a Live-owned app-server supervisor with progressive fenced publishing, partitioned control polling, cancellation and recovery safety, and measured integration coverage.

AI-assisted implementation under maintainer direction.
2026-07-12 19:14:05 -07:00
Paul BakausandClaude Fable 5 5635b1d484 the direction becomes a visible contract in the artifact
Transcript evidence (a12 01-observability): plans commit and deliver on
the axes with contract-strength language (palette, type, even theme
inversion) and stay default on the axis without one (layout gets a
single conventional breath). And plans living in invisible reasoning
means nothing can hold a build to its intent. The direction is now
written as a comment block at the top of the artifact answering: the
concept, the hour-later memory, why not the modal competitor page, the
signature, the first viewport's move. Critics and evals can score
delivery-against-contract; a mood is not an answer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 18:38:16 -07:00
Paul Bakaus 099aacee99 Add a compact Codex Live generator
AI-assisted: OpenAI Codex.
2026-07-12 18:35:29 -07:00
Paul Bakaus 2106a2881f Improve Live progressive responsiveness
Add transactional progressive publication, durable cancellation, responsive accept cleanup, and framework-safe Svelte and Nuxt previews.\n\nAI-assisted: OpenAI Codex.
2026-07-12 17:54:50 -07:00
Paul BakausandClaude Fable 5 ff67ad359e layout gets the source-exclusivity construction that fixed palettes
Paul's a10 review: palettes are refreshed (the palette-exclusivity
line's fingerprint) while layouts stay boring in every version. Same
cure, same shape: the layout has exactly two legitimate sources, the
concept or the content's own structure; the category's habitual
skeleton is neither.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 17:40:54 -07:00
Paul BakausandClaude Fable 5 073e17e180 three-directions sketch + the scene decides the theme
Paul's a11 review: heroes are safe SaaS viewports, everything
predictable; mobile Operate ships dark despite a brief that specifies
outdoors-in-motion use. Decide-then-build now opens with three
one-line directions differing in concept (the instinctive pick that
any studio would reach for is the default wearing your name); the
Operate mode adds: the usage scene is part of the spec, the theme
follows the scene, not the category's habit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 17:36:45 -07:00
Paul BakausandClaude Fable 5 cfbac54440 deprecate craft: the build flow lives in new-work.md, checkpoints are a mode
Per Paul: rather than gating a second file, fold what made the craft
path superior into the file both models already read 21/21 through the
gate. new-work.md gains 'Decide, then build' (direction as one
confirmable paragraph; attended pauses, unattended records-and-goes;
codex.md mock flow when image generation exists) and 'Finish like a
studio' (inspect, honest critique, patch, detector). craft becomes a
deprecated alias like teach: invoking it forces attended checkpoints,
nothing else differs; the reference is a redirect stub. codex.md
retargeted. Existing-world feature builds remain governed by the core
floor (unmeasured path, noted in the plan doc).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 17:08:02 -07:00
Paul BakausandClaude Fable 5 139d69f2b7 bare build requests follow the craft orchestration; brief-coverage joins the floor
Invocation A/B on Fable (a9 craft-path vs a9-direct plain): the plain
path scored 38% vs the competitor against the craft path's 50%, and
brief fidelity collapsed to 14% vs bare — the direct path drops asked-
for features that craft's direction step and engineering bar preserve.
Routing now sends any build request through the craft orchestration
unprompted (its gates pause only when a user can respond), and the
craft floor gains a brief-coverage recheck: every requirement the brief
names must exist on the page.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 21:49:05 -07:00
Paul BakausandClaude Fable 5 eeff485c20 mode belongs to the surface; palette sources are exclusive
a7 transcript evidence: 01-observability samples drew orange-honey and
green seeds, recited the color-strategy menu, and shipped dark
category-reflex palettes anyway; the model applied the subject's
workmanlike grammar to its own landing page. Two generic lines: the
mode belongs to the surface, not the subject (a landing page for a
dense tool is still Persuade; deciding a page can be plain because its
subject is workmanlike is the category error in reverse), and the
palette has exactly two legitimate sources (seed or the subject's
world; the category's habitual palette is neither).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 18:46:42 -07:00
Paul BakausandClaude Fable 5 f8180b2027 palette: neutralize the last brand roll-call (text-on-color convention)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 18:37:51 -07:00
Paul BakausandClaude Fable 5 2c62f0f4f9 de-SaaS the skill: mode-aware rules, neutral runtime injections, diversified examples
Fix batch from the visitor-mode bias audit. The skill's four modes
(Persuade / Operate / Read / Experience) now reach the places that were
still hard-coded to a SaaS-marketing default:

- palette.mjs: rewrote 45 seed blurbs in material/world terms. The 29
  tech-tool-world moods (13 Linear-indigo variants, 6 Figma-era, 5
  climate-tech, 3 fintech, 2 Glossier DTC, incl. seed-201's docs-page
  CTA red) lose all company names and product-category words; Aesop
  trimmed from 17 blurbs to 4 and Klim from 7 to 4, excess rewritten
  as unnamed material terms. Also carries the earlier bg-block rewrite
  (brand refs out of the composition doc).
- init.md: register explainer now names the four modes and the family
  each belongs to (stored value stays brand/product for compatibility);
  Conversion & proof interview + PRODUCT.md section gated to Persuade
  surfaces only (Experience/Read get no CTA/belief-ladder/proof).
- critique.md: Nielsen heuristics 7 and 10 may score n/a on Persuade
  and Experience surfaces, total renormalized to the applicable max,
  snapshot records which were n/a; working-memory examples diversified
  beyond dashboard/pricing anatomy.
- Register headers in bolder/delight/quieter/colorize/layout/animate/
  typeset renamed from Brand:/Product: to Persuade + Experience: /
  Operate + Read:; typeset and layout gain one Read-specific sentence
  (steady reading measure; navigable linearity).
- animate.md: plan checklist and implementation order lead with
  feedback and transitions; the single entrance moment comes after,
  scoped to modes that invite it.
- codex.md: mock inventory says "primary-action treatment (when the
  surface has one)" instead of assuming a CTA.
- delight.md: loading/empty-state/console-egg examples diversified
  beyond SaaS; streaks/badges scoped to Operate surfaces with
  recurring tasks.
- distill.md: step-removal and next-action lines neutralized away
  from signup/checkout/CTA vocabulary.
- document.md: canonical button label GET STARTED -> SAVE CHANGES;
  signature components gain a non-marketing example.
- antipatterns registry: single-font rule renamed to "Single font
  without hierarchy" with a description that permits one family when
  weight/size contrast carries hierarchy.

Staged provider copies regenerated via build:skills:release for the
touched files only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 18:35:53 -07:00
Paul BakausandClaude Fable 5 c32fcb47ef new-work: mode-neutral spine instead of a corrective lens
Paul: the mode-governs section was de-biasing a persuade-tinted
playbook rather than writing neutral prose, the exact compensating-
paragraph anti-pattern. Rewritten: the corrective section is gone; the
first-viewport thesis speaks of the concept doing its job (the work,
the product, the content, the task); everything-bold's form list
includes the exact-system form natively; prove-don't-claim covers
content delivering; type guidance is parameterized by mode in one
sentence. Net shorter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 18:09:14 -07:00
Paul BakausandClaude Fable 5 4e4b72c22b new-work: the mode still governs the playbook's energy
Paul's gallery check of the a8 docs run: skill-on still SaaS-ified the
documentation page. The playbook was persuade-flavored end to end, so
gating a greenfield Read surface through it risked amplifying exactly
that. New leading section: on Operate and Read surfaces boldness means
a committed system (typographic voice, spacing rhythm, one owned
accent, inevitable structure), the thesis is the content or the task
itself, and nothing invented may stand between the visitor and what
they came to do.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 18:06:08 -07:00
Paul BakausandClaude Fable 5 a55b162a24 routing owns the craft-vs-direct decision; craft.md stops advising its own loading
The when-to-choose guidance sat inside the file that only loads after
the choice is made. SKILL.md's routing now says it: bare build requests
build directly through the gate and floor; craft is routed only when
named or when the user asks for a guided, checkpointed build. The
Commands row describes craft by its checkpoints. craft.md's intro just
describes the supervised flow it orchestrates.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 17:46:26 -07:00
Paul BakausandClaude Fable 5 b409bedf5d drop brand.md/product.md stubs; craft repositioned as the collaborative build
Stubs removed per Paul (register: values remain harmless family hints;
nothing points at the files anymore). craft.md now opens by defining
itself against plain invocation: a bare build request goes straight
through the gate and the craft floor; craft is the supervised path with
guaranteed checkpoints and the mock pipeline. One shipping-discipline
line joins the core floor (real content, interaction states, respect
the build pipeline) so one-shots inherit the bar that previously lived
only in craft's Step 4.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 17:40:32 -07:00
Paul BakausandClaude Fable 5 051f856113 finish the mode migration: brand.md/product.md become redirect stubs
Answering the obvious question the family-depth framing dodged: with
modes derived per task, files named for the old two-register taxonomy
had no architectural reason to exist. brand.md's surviving depth (lane
test + inverse test, reflex-reject lanes, color discipline, layout
moves, permissions) folds into new-work.md, where all of it belonged:
it is new-identity Persuade/Experience guidance. product.md's content
moves unchanged to operate.md, its true name. Both old files remain as
one-line redirect stubs because register: brand|product in existing
PRODUCT.md files and older links point there. All cross-references
retargeted (SKILL.md modes intro, context.mjs REGISTER hint, live.md,
typeset.md); 85 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 17:35:28 -07:00
Paul BakausandClaude Fable 5 af8b814c99 reference consolidation: single source of truth across core, new-work, brand, product, craft
Overlap audit after the new-work split. brand.md slims to family depth
that exists nowhere else (aesthetic-lane tests, named-reference
discipline, brand layout moves and permissions); everything it
duplicated against new-work.md and the core (font procedure, reject
list, color strategy, imagery, scale/leading) is deleted, killing the
two-copies-drift hazard. product.md keeps its Operate depth nearly
intact (it was not duplicated) and gains a scope note covering Read
surfaces. craft.md becomes pure orchestration: gates, foundation,
shape handoff, image-gen flow, engineering bar, iterate, present;
its duplicated design guidance (imagery rules, visual-craft bullets,
mandatory reference reads) is replaced by pointers to SKILL.md's
craft floor and new-work.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 17:30:30 -07:00
Paul BakausandClaude Fable 5 0fde0850cf skill v4.0.0-alpha.9: daily-driver core + mandatory new-work playbook
Architecture per Paul: impeccable is primarily a daily driver on
existing codebases; the always-loaded core should serve that 90% path,
not carry the full generative arsenal on every invocation. SKILL.md now
holds brief-wins, existing-worlds (the headline path), the four visitor
modes, the full craft floor, and a hard gate: new identity work
(greenfield, or a redesign discarding the current look) MUST read
reference/new-work.md before any design decision. That file carries the
generative playbook (seed, subject grounding, plan/self-check/signature,
hero-thesis, everything-bold, prove-don't-claim, color commitment,
calibration, persuade type/imagery). context.mjs enforces the gate
mechanically: NEW_WORK directive when no PRODUCT.md/DESIGN.md exists,
and the old mandatory register-file read is replaced by a REGISTER
family hint. No surfaces: map anywhere; mode is derived per task.
Gate compliance is measurable via skillEvidence.directSkillFileReads.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 17:22:15 -07:00
Paul BakausandClaude Fable 5 bf2dd7ec13 skill v4.0.0-alpha.8: four visitor modes replace the brand/product bifurcation
Field report: impeccable SaaS-ified a developer docs page; the Opus
galleries showed the same on an album page. Root cause: two registers
force every surface into persuade-or-operate grammar. The register
section now names the visitor's mode first (Persuade / Operate / Read /
Experience) with mode-borrowing called out as the canonical failure,
and PRODUCT.md's register field maps as family (brand = Persuade +
Experience, product = Operate + Read) for compatibility. Read mode:
comprehension deliverable, navigable structure, chrome out of the way.
Experience mode: the artifact leads at every screen size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 17:06:14 -07:00
Paul BakausandClaude Fable 5 2e71facb59 skill v4.0.0-alpha.7: cultural surfaces are the work, not a funnel
Paul's Opus gallery observation: every impeccable 05-experimental-album
generation reads decidedly SaaS while frontend-design's open with the
art itself, especially at narrow viewports. Cause: the brand register
prescribed stop-the-scroll/earn-the-click/convert for ALL brand
surfaces. Split the register's deliverable by surface: product/service
pages convert; cultural surfaces (album, portfolio, publication, body
of work) lead with the artifact, recede the interface, and treat
conversion grammar as a category error — the visitor meets the work in
the first viewport at every screen size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 17:00:54 -07:00
Paul BakausandClaude Fable 5 02f760fbad skill v4.0.0-alpha.6: boldness is page-level commitment, not an element budget
Paul: everything should be bold, nothing bland; bold is neither
decoration nor clutter but commitment to the concept, whose form the
concept chooses (maximal or severely clean, drenched or monochrome,
piercing copy, the product demonstrating itself). Replaces the
'spend your boldness in one place' rule imported from frontend-design,
whose one-bold-element-on-a-quiet-page framing pulled pages toward the
tasteful softness the galleries showed losing. The signature becomes
where the concept peaks rather than the only place it lives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 16:55:13 -07:00
Paul BakausandClaude Fable 5 56edce955a skill v4.0.0-alpha.5: hero-as-thesis + commit-over-refined (distinctiveness push)
Paul's spot-check of the Fable validation galleries: frontend-design's
lektor generations read vastly more distinctive and subject-faithful
despite losing the overall pairwise verdict on craft. The arena agrees
on the axis (distinctiveness 8-31 at n=5). Two additions to the core:
the opening viewport is a thesis (open with the most characteristic
thing in the subject's world, with a concrete memory test), and an
explicit polish-is-the-floor counterweight so the craft floor stops
reading as a mandate for quiet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 16:48:23 -07:00
Paul BakausandClaude Fable 5 c3aba1e343 hooks: two-tier design hook — immediate per-edit rules + full-set Stop deep pass
Eval evidence showed the per-edit PostToolUse stream fires overwhelmingly
on copy-level rules (em-dash-overuse ~97x/session) and measurably makes
models more conservative, while a full-detector pass at completion is what
actually fixes contrast/padding/glow. Split the hook accordingly:

- Per-edit (PostToolUse) now surfaces only IMMEDIATE_TIER_RULES: broken
  output (broken-image, text-overflow, clipped-overflow-container,
  body-text-viewport-edge), objective contrast/legibility failures
  (low-contrast, gray-on-color, tiny-text), single-property mechanical
  slop (gradient-text, dark-glow), and design-system drift (the four
  design-system-* rules, which compound if left uncorrected). Everything
  else defers. Override with hook.perEditRules: "all" in
  .impeccable/config.json. Tiering is off for Cursor/Copilot harnesses,
  which have no Stop pass wired, so nothing gets silently dropped there.

- Stop deep pass (runStopHook): runs the FULL rule set over every UI file
  touched this session (tracked via the existing hook.cache.json session
  state; deferred-only edits now mark the file touched), dedupes against
  everything already surfaced per-edit, honors ignore-rule/file/value and
  inline disables, reuses the [impeccable@1] envelope, and no-ops fast
  when no UI files were touched. Emits hookSpecificOutput
  { hookEventName: "Stop", additionalContext } per the Claude Code SDK
  Stop contract (conversation continues so the model can act on it).
  Second Stop fire is silent - deep-pass findings are remembered.

- Wiring: Stop entries (timeout 30) in plugin/hooks/hooks.json, the
  .claude settings + .codex hooks manifests (transformers + hook-admin
  repair path). Claude Code and Codex both dispatch a native Stop event;
  Cursor's stop hook is inconsistently dispatched (pre-write gate stays)
  and Copilot's agentStop/sessionEnd don't inject model context, so
  neither gets a Stop entry - documented in reference/hooks.md.

- Tests: tiering split/override/harness gating, Stop dedupe + silent
  no-touched-files + ignore machinery + kill switches; existing per-edit
  tests moved to immediate-tier rule ids. 181 tests green; smoke-tested
  the built dist skill end to end (glow surfaced per-edit, em-dash only
  at Stop, second Stop silent).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 16:46:00 -07:00
Paul BakausandClaude Fable 5 50002a9e05 skill: rewrite codex block as positive calibration (self-priming fix)
gpt-5.6-sol evals: skill-on lost craft 0-25 to bare gpt-5.6; removing
the enumerated codex ban block recovered it to 4-16, confirming the
block's literal CSS patterns self-prime the defects they ban (the same
mechanism the v2.1 ablation sweep documented). Replaced with three
shape-level calibration lines: tracking floor (kept, it's a numeric
ceiling), elevation-declared-once + modest container radius, and
material honesty (real assets, surfaces not decoration, specific
claims). Detector rules continue to enforce the mechanical patterns.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 16:28:09 -07:00
Paul BakausandClaude Fable 5 6b3d174e93 skill v4.0.0-alpha.4: the lean core — full design guidance at a quarter the length
Pairwise evals on Fable one-shot (6-task regression set, opus-4-8 judge,
position-bias-cancelled): the hand-distilled ~55-line lean core beat the
heavy v4 core 66% overall / 67% craft head-to-head, and moved the
decisive win-rate vs frontend-design from 13% to 27% (40% with the
completion-time QA scan; craft went positive 6-5 for the first time).
18/18 lean samples ran context.mjs + palette.mjs vs a minority under the
heavy core: shorter instructions get followed. Context weight itself was
suppressing both compliance and boldness.

Structure: persona + brief-wins + existing-worlds + subject-grounding +
plan/self-check + boldness + prove-don't-claim + commit + calibration +
compressed craft floor + two-paragraph registers. Commands table kept;
the no-arg context-aware menu logic moved to reference/routing.md (read
on demand in the only case that is inherently interactive). Provider
blocks and rule anchors preserved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 18:19:35 -07:00
Paul BakausandClaude Fable 5 1172898020 skill v4a3: prove-don't-claim + load-bearing signature
Judge rationales across cand-v4a2 arenas: competitor wins by showing
the product working (mix panels, comparison tables, live demos) and by
signatures big enough to organize the page; our samples claim, decorate,
and sometimes stop at the hero.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 14:09:49 -07:00
Paul BakausandClaude Fable 5 f1078c59b0 skill v4a2: seed defers to subject's world; unattended-mode gates for craft/shape; init skip when no user
Eval evidence (cand-v4a1-prose): palette.mjs handed a random violet seed
to the Polish-TV lektor brief and the model anchored on it, overriding
subject-grounding; craft/shape user gates can't fire in one-shot runs
and each model improvises around them. Seed is now a reflex-check that
yields to a subject-dictated palette; craft/shape gain an explicit
unattended mode (same bar, no waiting); init interview is skipped when
no user can respond.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 13:05:47 -07:00
Paul BakausandClaude Fable 5 440348a498 skill v4 core: existing-world/new-work gate, register-scoped type rules
Per Paul's guidance: (1) existing committed design systems are the
bread-and-butter case and get a first-class core rule (work inside the
world, no parallel colors/fonts/styles, no perf regressions); (2) a
redesign that discards the current look is new identity work and runs
the full concept/tokens/signature process instead of anchoring to the
incumbent skeleton (the lektor failure); (3) the reflex-reject font list
and physical-object font procedure are brand-register rules, moved out
of the universal Commit section — system stacks and workhorse UI faces
are legitimate, often correct, for product UI, stated positively in the
product register.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 12:04:46 -07:00
Paul BakausandClaude Fable 5 a38a0765a9 skill v4.0.0-alpha.1: always-loaded core — brief-wins, subject grounding, token/self-check process, inline craft floor + registers
One-shot evals on Fable 5 (impeccable-evals notes/fable-oneshot-craft-plan.md)
showed the reference-file architecture failing: models skip the register
reads, so most design guidance never reaches them, and skill-on collapses
toward bare-model output (0/9 pairwise wins vs frontend-design on r10).

SKILL.md is now self-contained for one-shot work: persona, the-brief-wins
rule, ground-it-in-the-subject, a plan/tokens/signature/self-check process
gate, commitment guidance, a compact inline craft floor, and distilled
brand/product registers. Reference files remain as sub-command flows and
optional depth. The enumerated absolute-bans list is retired from prose;
mechanical slop enforcement moves to the detector/hook.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 10:46:10 -07:00
Paul BakausandGitHub da99645a58 Add OpenAI plugin submission bundle (#363)
* Add OpenAI plugin submission bundle

Build a Codex-native OpenAI plugin with bundled hooks, public listing metadata, submission guidance, privacy coverage, and regression tests.

AI assistance: OpenAI Codex prepared and validated these changes under maintainer direction.

* Fix provider script command rendering

Replace heuristic rewrites across executable scripts with one explicit provider marker, render pinned shortcuts per target harness, and remove the personal email from the public publisher manifest.

Addresses automated review feedback on PR #363.

AI assistance: OpenAI Codex prepared and validated these changes under maintainer direction.
2026-07-09 17:09:13 -07:00
51e5af258e Expand init to capture positioning, conversion, and proof context (#315)
* Add positioning and conversion questions to init flow

Expand init.md so PRODUCT.md captures audience splits, positioning,
and brand-register conversion/proof context before design work starts.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix init over-inference by raising the evidence bar for skipping questions.

Sparse repos were letting the model treat weak guesses as settled answers; Step 3 now asks unless the codebase provides strong, explicit evidence.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Improve init interview order and PRODUCT.md proof output shape.

Ask positioning in round 1, actively collect proof assets, and give Proof & conversion a plain bullet skeleton so generated PRODUCT.md stays lean.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix init interview bundling and write-time padding, verified via harness runs

Co-authored-by: Cursor <cursoragent@cursor.com>

* Revert init reference follow-up rule to advisory wording on line 88

Co-authored-by: Cursor <cursoragent@cursor.com>

* Tighten init interview rules after harness runs: split register, options, prose

Settle split register before brand-only questions, require standalone emotions
and confirmed secondary audiences, forbid compound options, and keep PRODUCT.md
bold minimal.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Ask brand-register init questions in magazine-editor voice, no skill jargon

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix init chat fallback to ask one question at a time

When no structured question tool exists, init should ask in chat with
lettered options and wait for each answer instead of dumping a list.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Resolve init review comments: split purpose question, gate template section

Purpose and success are now separate questions, and docs-stated purpose
is framed as a hypothesis below the strong-evidence bar rather than a
competing always-ask rule. The PRODUCT.md template now tells product
register to omit the Conversion & proof section including its heading.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Keep belief-sequence question out of skill jargon

Ask what visitors must believe in plain words; map the answer to the
template belief ladder in a parenthetical instead of leading with the term.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Abdul Wahab <abdulwahab@Abduls-MacBook-Pro-2.local>
2026-07-09 16:20:20 -07:00
f40e2f8f0a Add mechanical pre-scan for typeset and layout (#345)
* Add mechanical pre-scan for typeset and layout commands.

Introduce --scope filtering, layout/type rule scopes, DESIGN.md font-size validation, and pre-scan steps in the skill references so agents run detect before LLM judgment.

Fixes #149

Co-authored-by: Cursor <cursoragent@cursor.com>

* Add isolated sub-agent orchestration for typeset and layout pre-scans.

Run the mechanical detector and visual assessment in parallel sub-agents so deterministic findings cannot anchor LLM judgment, matching the critique pattern Paul requested on PR #345.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix: reject bare --scope so detect never scans unscoped by mistake.

When --scope had no value, the CLI dropped the flag and ran a full scan instead of failing, which could silently use the wrong rule set during typeset/layout pre-scans.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix: require both typeset and layout assessments in sub-agents.

Close a loophole where agents ran only the mechanical pre-scan inline by interpreting "running both" as permitting one inline assessment.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Abdul Wahab <abdulwahab@Abduls-MacBook-Pro-2.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-09 08:29:21 -07:00