Commit Graph
11 Commits
Author SHA1 Message Date
Paul BakausandGitHub 49d8cbff16 Comp-fidelity review discipline + conciseness pass on core references (#586)
* Comp-fidelity review discipline + conciseness pass on core references

Process fixes derived from a real Codex session (Hanasaku landing page)
where a build drifted wholesale from the approved comp and still shipped
under a reviewer pass:

- finish reviewer: new Evidence check (check 0) with a fourth
  disposition, recapture, for malformed screenshots; a review on invalid
  evidence binds nothing and owes a full re-review, not a verdict pass
- finish reviewer: verdict passes exit scoring mode when recaptures fail
  check 0 or when the packet carries user-supplied screenshots that
  contradict a prior verdict (those force a fresh full review); a ship
  earned in a verdict pass covers the scored fixes, not the whole surface
- new-work: capture-validity rules (settle entrance motion, capture from
  document top, comp comparison at comp dimensions, open every file once
  before sending); user's actual viewport joins the inspected sizes
- new-work: hero checkpoint now writes .impeccable/review/hero-repro.png
  and the reviewer verifies it exists under Persistence
- new-work: comp authority is explicit (only the user can downgrade it);
  handoff reports the verdict at its actual scope; user evidence reopens
  a full review; documenter re-runs when fixes land after documentation
- craft-floor: Refuse entry for geometric masks approximating organic
  photographic contours (the circular-cutout failure)
- editorial conciseness pass over new-work.md, visualize.md, and both
  agent files: tighter sentences, no dropped rules, all rule markers and
  mechanical tokens preserved

Assisted-by: Claude Code

* fix: define the ship disposition in new-work's action paragraph

Copilot review finding: the paragraph claimed exactly four disposition
words but defined only recapture, rebuild, and fix.

Assisted-by: Claude Code

* fix: rebuild returns get a full review; recapture return shape in preamble

Cursor Bugbot findings:
- a return following a rebuild directive is now a fresh full review on
  both sides of the contract, never a verdict pass, so a wholesale
  rebuild cannot earn a scoped ship on the directive alone
- the turn-ceiling preamble now names the recapture return shape instead
  of contradicting it with "the five sections"

Assisted-by: Claude Code

* fix: absent required captures fail the evidence check

Greptile finding: a packet with no desktop.png/mobile.png (or missing
native device-class captures) routed to the missing-input notice and
could still reach ship. A required capture that is absent now fails
check 0 exactly like a malformed one and forces recapture; the
missing-input allowance in the preamble excludes captures.

Assisted-by: Claude Code

* fix: user-viewport capture is a required, named input to the review

Greptile finding: the evidence gate hard-coded web requirements to
desktop.png and mobile.png, so a reported user viewport could join the
inspected set and still ship uncaptured. The parent now saves it as
user-<width>.png and names every inspected viewport required in the
packet; check 0's required set includes every brief-named capture.

Assisted-by: Claude Code
2026-08-14 05:44:26 -07:00
Paul BakausandClaude Opus 5 d417ff1f01 Craft floor: theme the surfaces you did not draw
A well-made site was audited for what separates it from a competent one, and the
answer was not its ingredients. It runs the default stack, Next and Tailwind and
Geist, with no world and no unusual technique. What it has is attention to the
surfaces a browser renders for you: 29 focus-visible rules, 15 scrollbar rules,
and styled text selection, caret, underline offset and scroll behaviour.

Those are the cheapest signal that a page was built rather than assembled, and
the ones a model skips most reliably, because nobody asks for them and nothing
looks broken without them. The floor already covers contrast, depth, spacing,
measure, motion and states; this is the layer under all of it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 13:32:08 -07:00
Paul BakausandClaude Fable 5 28af30eff0 Condense the grown skill files and harden the reviewer's verdict
Three subagent audits reviewed the files that grew through the last
rounds of patches. Their honest verdict: dense, not bloated; roughly
430 words of true redundancy came out with no rule lost, and every cut
they flagged as removing compliance pressure was skipped. Highlights:
approval recording now has one owner in visualize.md, the asset
producer's crop ban went from three statements to the deliberate pair,
its two transparency passages carried contradictory defaults (resolved
toward true alpha first), and the 450-word medium-gate wall split into
three paragraphs at zero cost. The producer also gained a mode-seam
sentence so a sketch run cannot return an asset manifest.

The reviewer's verdict is no longer soft: a derived disposition line
(rebuild / fix / ship) opens every return, computed from the matrix
rather than felt, recomputed after the verdict pass, and never
softenable by the parent, who must report it verbatim. The second
hamster-wheel run showed the parent inventing 'PASS WITH FIXES' over a
matrix with MATERIAL contradicted on the focal element.

Two additions from the same session's evidence: hard offset shadows
outside a neobrutalist world join the craft floor's refusals (codex
invents them without fail), and hookless harnesses must run detect.mjs
once before the finish review, because codex has no hooks and the
detector otherwise never sees the build at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 17:45:45 -07:00
Paul BakausandClaude Fable 5 e54fd13a33 Close the line-art loophole and widen the rebuild directive
The second codex hamster-wheel run read the medium gate and still
assigned a shaded, perspectived technical illustration to 'Authored SVG
geometry': the world was an instruction booklet, so the affinity
clause's 'diagrams' blessed the downgrade, and the page shipped as flat
clipart against an illustration-grade comp. The gate now says style
does not move the boundary: perspective, shading, figure drawing, or
dense mechanical detail is illustration however line-drawn it looks,
and authored SVG ends where drawing skill begins. The craft floor's
sketchy-SVG rule carries the same sentence.

The reviewer in that run built an honest matrix, MATERIAL contradicted
on the focal element, and still emitted it as a fixable item the parent
answered with CSS. The rebuild directive now fires when MATERIAL is
contradicted on the focal element, not only when TYPE falls with it,
and every asset-requiring fix must say 'produce: <region>' so it cannot
be answered as a style tweak.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 17:45:45 -07:00
Paul BakausandClaude Fable 5 4432b92bbb Harden the comp-to-build translation after the hamster-wheel failure
A codex greenfield build produced an excellent approved comp and then an
abysmal page, and the reviewer approved it. The failure chain: the
implementation inventory downgraded a photographic hero to 'silhouette
in SVG' and sculpted panels to 'material finish: CSS'; the builder read
'no photography on hand' as a license to avoid photographic rendering;
QA looked at one full-page thumbnail; the reviewer was spawned with the
builder's forked history and then scored fix claims instead of pixels;
and the output contract had no way to say 'rejected'.

The fixes, stage by stage:

- The inventory's medium column gets a gate: a human figure, product
  object, machinery, or lit material is raster whatever the stack, and
  such regions are regenerated cleanly at asset resolution with the
  comp and its embedded prompt as reference. Never cropped from the
  comp, whose effective resolution is reference grade; the asset
  producer's direct bucket closes the same hole. Dropping an
  image-native region is a user decision at the approval point.
- Generated imagery is a material, not a claim: evidence rules bind
  assertions, never render fidelity.
- The build thread's inspection becomes a region-by-region side-by-side
  against the comp at legible scale, never one full-page thumbnail.
- The reviewer spawns fresh, never with forked history (fork_turns: 0
  in codex), and gains a rejection lane: when TYPE, MATERIAL, and the
  focal element are all contradicted, the first material fix is a
  rebuild directive the parent surfaces to the user instead of
  patching. Verdict passes score recaptures only; the parent's fix
  narration is not evidence.
- The verdict-loop ceiling softens: two rounds ends an unattended run,
  but an attended session puts the open-items table in front of the
  user and lets them fund another round; any round that resolves
  nothing stops the loop.
- Comp approval joins the roll as skip-proof: question-tool errors fall
  back to the decision page, delegation is recorded in the brief and
  the sidecar and disclosed up front, and the reviewer treats comps
  with no recorded pick as a material finding.
- Craft floor: system display faces (Impact, Arial Black) as an
  own-world display voice and unicode glyphs standing in for icon
  systems are named failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 17:45:45 -07:00
Paul BakausandClaude Fable 5 430d74a12b The stack is the user's decision, and code is a medium of ambition
Two field observations. Greenfield projects with no framework never got
asked what to build on: the interview covered product truth and banned
aesthetic questions, and the model silently picked a scaffold the user
never chose. Init now asks once, static HTML, a named framework, or a
delegated choice plus any deploy constraint, and records the outcome
under a new optional Stack section, including the delegation itself, so
later work knows the choice was offered.

And the medium guidance named raster a dozen times while naming WebGL
once, so models never reached for vector or GPU code unprompted. The
affinity now runs both ways at the decision point: precise geometry,
shape systems, diagrams, expressive motion, shaders, and anything
interactive are vector and GPU territory, where a raster flattens what
should move, scale, and respond. The sketchy-SVG ban states its own
scope: it bans SVG imitating pictures, never SVG doing geometry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 16:42:56 -07:00
Paul BakausandClaude Fable 5 33a1c5fcae Ban kickers outright: one eyebrow above a heading is one too many
The detector's repeated-section-kickers rule waited for three tracked
labels before calling the pattern; generated pages earn the finding on
the first one. Retire that id and replace it with kicker-above-heading,
which flags any tracked-caps or small-caps label block sitting directly
above an h1-h4 or heading-role element, at full warning severity.

The candidate gate absorbs the false-positive shapes the repetition
count used to paper over: editorial category-and-date meta lines,
breadcrumbs with separators, legal and chapter numbering, application
panel context labels, nav landmarks before page titles, and stat
callouts with the label below the number. Hero-scale h1 eyebrows stay
with hero-eyebrow-chip so one element gets one finding, and the static
cascade now carries font-variant so small-caps kickers register.

The craft floor entry moves from caution to ban in the same breath.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:55:52 -07:00
Paul BakausandClaude Fable 5 8634c538fb Verification is two bounded rounds, never a loop
Opus 5 turned the iterate-with-screenshots-until-it-meets-the-bar
instruction into 42 screenshot trips and 150 tool calls per build,
about forty dollars of cache churn a page, before ever reaching the
reviewer. Verification now batches: one desktop-and-mobile round after
the full build, fixes applied together, one confirming round, ceiling
two. Craft-floor's checks share those renders instead of earning
separate trips; per-tweak iteration is live mode's channel.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 18:43:04 -07:00
Paul BakausandClaude Fable 5 36e3c05ca7 Restore dice assignment, fusion, and the commitment counterweights
The ship40 concept pipeline had reversed the proven a-series mechanisms:
the seed's roll decayed into a shortlist nomination that taste functions
(model ranking, candidate floor, simulated user) then argmaxed into the
safest card; the costume check returned as the Translation veto and
carrier-removal test; and the 07-15 rewrite deleted the calibration,
reflex-font lanes, color strategies, and commit-every-atom language that
had held off the cream-editorial default since the alpha era. Five of six
frozen craft directions converged on the same warm-paper family and both
builders obeyed them.

This lands the repair on top of the in-progress simplification:

- new-work.md: the script assigns the build index again on both scopes;
  catalog challengers are fused (challenger supplies form and grammar,
  product supplies every fact, clarity wins conflicts) and weighed on the
  two proven axes only; attended runs present one fully committed
  direction with re-roll and an optional steer instead of a ranked
  lineup; the color-strategy picker, reflex-face list, saturated-look
  calibration, first-viewport thesis and memory test, commit-every-atom,
  scroll pacing, and prove-don't-claim return; the direction contract
  returns as five lean blocks audited by the separate-agent finish.
- concept-seed.mjs: PROMOTED INDEX becomes ASSIGNED INDEX with
  build-assignment semantics; self re-roll only on named factual grounds.
- craft-floor.md: hook-active sessions act on findings instead of
  re-auditing; the Refuse list is framed as category defaults the brief
  can earn; a closing commitment line keeps a ban list from being the
  last word before code.
- codex.md / shape.md: contract references restored for flow coherence.

Adopts the concurrent session's ceremony cuts, softened challenger
instruction, seed SOURCE IDs and --candidate-count, detector-ownership
fix, and the removal of the hook-side contract audit (the audit now
belongs to the separate reviewer at finish).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 10:43:44 -07:00
Paul BakausandClaude 68eeb93b3d Tighten the craft floor
910 words to 682, same 35 rule markers, no guidance dropped.

- Two sections instead of three. "Absolute bans" and "detector-blind
  reflexes" split the same list by whether our scanner happens to catch
  it, which is a fact about our tooling and tells the model nothing about
  the design. Merged into one Refuse list, grouped as page scaffolds and
  surface habits, which is a distinction the model can act on.
- Folded three duplicates: text-overflow was already in the Type check,
  the uniform section reveal was the other half of the Motion check, and
  card-everything was already inside the card-grid ban.
- Cut explanation the model does not need. It knows what gradient text
  is and what group-hover does; it needs the refusal, not the mechanism.
  The gemini block goes from four sentences to three short ones, and the
  motion palette line drops the CSS tutorial for "reach past transform
  and opacity."
- The authority note moves to the header so no item has to hedge.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-21 02:14:22 -07:00
Paul BakausandClaude 153b416f2e Move the slop defects back into the craft floor
The detector-blind slop review existed because the AI-tell rules had been
stripped out of SKILL.md and nothing carried them. The floor is a better
home: it loads after concept ideation and immediately before editing UI,
which is the placement that made stripping them necessary in the first
place. Models tread lightly when a ban list is present during ideation;
by the time the floor loads, the direction is already committed.

- Rename build-floor.md to craft-floor.md and restore the absolute bans
  (side-stripes, gradient text, glassmorphism, hero-metric, identical card
  grids, eyebrow-on-every-section, numbered markers, text overflow), the
  codex and gemini defect lists, and the reflexes no scanner catches.
  Rule ids match the ones the ablation catalog already knows.
- Delete lib/slop-review.mjs and both injections. The Stop hook is now
  purely a mechanical pass and stays silent with nothing to report.
- context.mjs replaces AI_SLOP_REVIEW_REQUIRED with the narrower
  MANUAL_DETECTOR_REQUIRED, emitted only when a session has no hook at
  all. A per-edit hook already covers the mechanical gap, and the floor
  covers the judgment one either way.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-21 02:04:36 -07:00