Commit Graph
1252 Commits
Author SHA1 Message Date
Paul BakausandClaude Fable 5 fa1177ed9c Ship native subagent definitions for GitHub Copilot and Cursor
The github and cursor providers previously received only the generated
degraded/ inline fallbacks. Both harnesses support real custom subagents,
so the build now emits them from the same skill/agents/ source:

- GitHub Copilot: .github/agents/impeccable-<role>.agent.md with portable
  frontmatter only (name + description; omitting tools grants all tools,
  and Copilot has no documented model/effort/max-turns equivalents).
- Cursor: .cursor/agents/impeccable-<role>.md with name, description,
  model: inherit, is_background: false, and readonly derived from the
  agent's tool list (true only for the finish reviewer, which declares
  neither Write nor Edit). effort/max-turns are skipped because Cursor's
  effort option requires an explicit model id.

Agent bodies now also resolve {{scripts_path}} and strip rule markers in
the shared agentFormat pipeline, which fixes the previously unresolved
placeholder in the emitted Claude asset-producer agent.

The CLI installer places agents per scope: project installs write
<repo>/.github/agents/ and <repo>/.cursor/agents/; user-level installs
write ~/.copilot/agents/ (Copilot's user dir, not ~/.github/) and
~/.cursor/agents/, overwriting stale impeccable-* copies. Because
Copilot lets user-level agents shadow same-named project ones, a project
install warns when shadowing copies exist; Cursor gives project agents
precedence, so no warning there.

new-work.md and visualize.md extend their harness-naming clauses with
the Cursor and Copilot invocations. The degraded/ fallbacks keep
shipping for surfaces where the model still fails to delegate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 16:16:57 -07:00
github-actions[bot] dedb8a1df2 Sync generated provider output 2026-07-28 20:41:53 +00:00
Paul BakausandClaude Fable 5 47b875a7e3 One prompt carrier across every harness: embed-prompt.mjs
The prompt behind a generated image was recorded three different ways,
a sidecar in the eval harness, nothing in the skill's API tool, nothing
for native tools, so intent survived or vanished depending on where you
ran. One dependency-free script now embeds the prompt inside the image
itself, PNG tEXt or JPEG COM with a sidecar fallback for other formats,
idempotent, and reads it back from any impeccable-generated file. The
API tool embeds automatically; the prose directs every native-tool
generation through it; copies between machines and harnesses keep their
intent. Comps meanwhile are declared the build thread's own work, never
delegated, and the comp-skeleton guidance now asks for the surface's
actual regions instead of prescribing navs onto pages that have none.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 13:39:48 -07:00
github-actions[bot] abe722d105 Sync generated provider output 2026-07-28 20:27:43 +00:00
Paul BakausandClaude Fable 5 39532a65a2 Comps are pages not vignettes, and the prompt travels with the asset
Two findings from the first human-validated probe. The comps rendered
as scene vignettes because the generation prompts led with the world's
atmosphere; the model painted the fish market instead of the fish
market's website. The comp guidance now demands the page's literal
skeleton in the prompt, nav and its items, headline block, sections in
order, footer, with a self-check: a render that could hang as a poster
is not a comp. And generation context is part of the asset: the thread
that wrote a prompt knows what the image contains and why, so build-
critical imagery prefers the build thread, and subagent-produced assets
must carry their prompts, via the tool's new sidecar or the manifest,
read by the builder before composing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 13:27:03 -07:00
github-actions[bot] 14c27e43af Sync generated provider output 2026-07-28 16:48:59 +00:00
Paul BakausandClaude Fable 5 c4d22bb9dc TYPE and MATERIAL do not lapse when no comp exists
The failed gallery batch bound its seed, ran the reviewer, and still
shipped CSS bevels imitating enamel: the matrix's material row was
defined against the approved comp, and comp-less runs left it with no
reference. The rows now fall back to the contract's OWN-WORLD and the
world's real materials, with faked physicality contradicted on its
face; imitation material is the single most reliable mark of
machine-made design.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 09:48:22 -07:00
github-actions[bot] 60668224b1 Sync generated provider output 2026-07-28 16:48:18 +00:00
Paul BakausandClaude Fable 5 25934b9f6f The contract survives the compiler, and the roll has no skip condition
Transcript archaeology on the failed gallery batch split the binding
break three ways. One model authored a complete, correct contract that
Astro then erased: the compiler strips a slot's leading comment while
keeping deeper ones, so the contract now belongs to the root layout's
body as its first child, and the first production build gets grepped
for the seed key, because a contract the build erased is a contract
nobody can audit. Another model simply skipped the roll and built the
exact category default the seed exists to refuse; the roll step now
states outright that it has no substitute and no skip condition. The
third failure was the worker watchdog, fixed separately.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 09:47:36 -07:00
github-actions[bot] 042b81cb8d Sync generated provider output 2026-07-28 16:02:34 +00:00
Paul BakausandGitHub 963e13e040 Merge pull request #425 from vinaypokharkar/fix/detect-system-chrome-gpu-window
fix(detect): use system Chrome on Windows to stop GPU crash-loop window (#372)
2026-07-28 09:02:01 -07:00
github-actions[bot] 1cf7d7ab0f Sync generated provider output 2026-07-28 03:28:15 +00:00
Paul BakausandClaude Fable 5 7cd43c0365 The contract carries the exit condition, because the file outlives attention
Two probe runs on two different harnesses built complete pages and
declared done without ever entering the finish sequence: the reference
was read once near turn four and the finish choreography had fallen out
of attention thirty turns later. The one text a model rereads on every
edit is its own artifact, so the direction contract now closes with a
FINISH line naming the exit condition verbatim: unreviewed and
undocumented is unfinished; this build ends with the finish review, the
verdict, and DESIGN.md. A page that looks complete with that line
undischarged is not done, it is abandoned at the finish line.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 20:27:44 -07:00
github-actions[bot] d8f1deb35d Sync generated provider output 2026-07-28 03:06:14 +00:00
Paul BakausandClaude Fable 5 8b6324d1b9 View every image by its workspace-relative path
A sandboxed harness rejected view_image on an absolute path to a mock
the model had itself just produced under .impeccable/mocks/, killing
the run. The relative-path rule existed only for downloaded quality-bar
cards; it now covers every image the flow produces or references, in
the comp round and in the asset producer's comparison step.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 20:05:37 -07:00
Paul BakausandClaude Fable 5 68b1129634 Release bumps: skill 4.0.3, CLI 3.4.0, extension 1.3.0, with synced output
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ext-v1.3.0 cli-v3.4.0 skill-v4.0.3
2026-07-27 19:17:09 -07:00
Paul BakausandClaude Fable 5 ce4dcf9a93 Split breadth from rating in the challenger and staging pools
Rating grades quality, breadth says whether a world can serve an
arbitrary build at all; while they shared one field, the only way to
hold a narrow world back was calling it marginal, which made excellent
but narrow unrecordable and corrupted the ratings as a calibration
signal for the next authoring round. Both axes now exclude
independently, either kind of hold keeps its approval for direct
briefs, an all-niche tier falls back rather than starving, and
stagings honour the same gate with the same fallback. Tests cover the
niche exclusion at strength, the fallback parity with marginal-only
tiers, and the staging gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 19:17:09 -07:00
github-actions[bot] d3c7b05a3e Sync generated provider output 2026-07-28 01:56:22 +00:00
Paul BakausandClaude Fable 5 690e24129a CLAUDE.md: the rule engine is a facade now; drop the dead line numbers
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:55:52 -07:00
Paul BakausandClaude Fable 5 33a1c5fcae Ban kickers outright: one eyebrow above a heading is one too many
The detector's repeated-section-kickers rule waited for three tracked
labels before calling the pattern; generated pages earn the finding on
the first one. Retire that id and replace it with kicker-above-heading,
which flags any tracked-caps or small-caps label block sitting directly
above an h1-h4 or heading-role element, at full warning severity.

The candidate gate absorbs the false-positive shapes the repetition
count used to paper over: editorial category-and-date meta lines,
breadcrumbs with separators, legal and chapter numbering, application
panel context labels, nav landmarks before page titles, and stat
callouts with the label below the number. Hero-scale h1 eyebrows stay
with hero-eyebrow-chip so one element gets one finding, and the static
cascade now carries font-variant so small-caps kickers register.

The craft floor entry moves from caution to ban in the same breath.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:55:52 -07:00
github-actions[bot] 806a48aef2 Sync generated provider output 2026-07-28 01:52:13 +00:00
Paul BakausandClaude Fable 5 6a7d75b6fe Bound the finish by verdict, not by count, and teach the matrix medium and type
The hard stop landed one step early: one review, one batched fix, one
recapture, then done, with nobody ever judging whether the fixes reached
the quality the findings named. A recapture measures positions; the
model then presented mechanical confirmation as artistic success over a
page whose display face, material, and hero legibility had all drifted
from the approved comp. The finish now ends on a verdict: the recaptured
screenshots go back to the same reviewer, which scores every material
fix resolved, partial, or unresolved and names at most three regressions
the batch introduced, no new hunt. Partial and unresolved fixes earn
exactly one more round; two rounds is the ceiling, the second verdict
ends the work whatever it says, and the final verdict table goes to the
user as it stands, open items included.

Three blindnesses from the same run close alongside. The matrix gains
two mandatory rows: TYPE, where a display face of a different character
is contradicted however the layout matches, and MATERIAL, where flat CSS
standing in for painted, textured, or dimensional artwork is contradicted
regardless of placement. And the Truth check now requires every produced
asset visibly present in the screenshots, because a paper texture at
0.16 opacity is a compliance token, not a shipped material.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:51:32 -07:00
github-actions[bot] 270f177d1d Sync generated provider output 2026-07-28 01:19:54 +00:00
Paul BakausandClaude Fable 5 09a33bc58b One sketch, one agent: retire the batch producer and its supervision
The batch producer was the clumsy piece: one subagent owning eight
jobs needed heartbeat rules, reclaim windows, and a page full of
fallbacks to survive its own opacity. The unit of work is now a single
card. With parallel subagents, the set fans out one agent per card, up
to four in flight, landing everything in roughly the time of one; a
single-sketch agent has no planning phase and no batch to stall, so a
failure costs one slot and its remedies fit one sentence: regenerate an
empty slot when its agent returns, drop it when the user answers first.
Without parallel subagents, the main thread generates in reading order
after serving, and the harness's own generation display carries the
progress. The page-side streaming is unchanged; it never cared who
writes the files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:19:25 -07:00
github-actions[bot] f59c5223a4 Sync generated provider output 2026-07-28 01:13:37 +00:00
Paul BakausandClaude Fable 5 4329f757f5 Only the visible card face is interactive
A hidden backface still hit-tests in Chrome, so after flipping a card
the front's picture-in-picture sat invisibly over the back's chips,
showing its zoom cursor and eating the flip-back click. Pointer events
now follow visibility: the back is inert until the card flips, and the
front goes inert while it is flipped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:13:04 -07:00
github-actions[bot] 69bf1e9523 Sync generated provider output 2026-07-28 01:11:27 +00:00
Paul BakausandClaude Fable 5 ca88ea008b Patience while sketches land, honesty when standing in
Field data: the first image of a real batch took ninety seconds and the
page's 150-second fallback then silently promoted catalog art to full
bleed, unlabeled, which is exactly the this-is-your-design misread the
picture-in-picture treatment exists to prevent. The policy is now
patience while there is progress: a slot shows its inspiration only
after waiting four minutes with nothing landing anywhere on the page
for four minutes, the stand-in is dimmed and labeled 'inspiration ·
sketch pending', and polling continues so the real sketch still swaps
in whenever it arrives. Slots with no inspiration keep the honest
elapsed shimmer instead of folding. The parent's reclaim rule matches:
files landing steadily is health at any pace, and only total silence,
no first file in three minutes, takes the batch back inline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:10:56 -07:00
github-actions[bot] c3fe6d8064 Sync generated provider output 2026-07-28 01:06:14 +00:00
Paul BakausandClaude Fable 5 17bc2701f3 Put the full read on the card's back; the front is for choosing
Field feedback: with every fact stacked under the media the cards ran
past a screen tall. The front now carries only what the choice needs,
sketch, lineage, title, thesis, identity, and the honest risk clamped
to two lines, while first viewport and the case read on the back behind
a Details chip, sharing the face with the board when the world has one.
Risk stays on the front because the counterweights are pointless if the
downside hides behind a flip, and once the sketch lands the first
viewport is a picture anyway. The schema notes now ask for one-sentence
facts, since a long fact should cost the reader a flip, not the page
its scanability.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 18:05:41 -07:00
github-actions[bot] aef8cbac34 Sync generated provider output 2026-07-28 00:49:56 +00:00
Paul BakausandClaude Fable 5 0eb443d29b Bound the hand, greek the copy, and treat waiting as supervision
A codex field run dealt six challengers into an eight-sketch batch
behind an opaque subagent, and the user stared at a page of shimmer
asking whether anything was happening at all. Four fixes from that run.
A hand now holds at most three challengers, the rest banked for
re-rolls, so fairness within the hand stops multiplying into a queue.
Sketches greek everything but the product's real name and one real
headline, because an invented spec, price, or ship date in a sketch is
a claim PRODUCT.md never made, and comps have solved this for a century.
Sketch production follows the user's reading order with the first file
doubling as the producer's heartbeat, and the parent's --wait loop
checks the sketch directory each pass, reclaiming the batch inline when
two minutes pass with nothing landed. And a failed --start now captures
the daemon's stderr to a per-key log and names the sandbox as the usual
suspect, instead of reporting only that failure occurred. The shimmer
counts its elapsed seconds, and gives up at 150 instead of 300.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:49:21 -07:00
github-actions[bot] 149d71a772 Sync generated provider output 2026-07-28 00:15:34 +00:00
Paul BakausandClaude Fable 5 58a2d3dccd Bleed the deck to the viewport, fade the fuller side, let the glance take over
Three field notes from a live review. The deck now escapes the content
column and runs edge to edge, so a cut-off card sits at the screen edge
where it reads as more cards instead of at an invisible container edge
where it reads as a bug; the first card still aligns with the column
via scroll padding. Whichever side hides more content wears a fade, and
a hard edge means the end. The vertical pager grows from a bare chevron
into labeled Back and More pills, because in a column deck it is the
primary way forward. And hovering the inspiration thumb now takes over
the whole media region instead of a timid zoom; the sketch is the
promise, the inspiration is a glance, and the glance must cost nothing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:14:59 -07:00
github-actions[bot] f5827256d0 Sync generated provider output 2026-07-28 00:11:12 +00:00
Paul BakausandClaude Fable 5 e6612ea8ef Page the deck on its long axis, and never let decoration hide the cards
Field-checked in a real browser, which surfaced three defects the DOM
tests could not: the generic .media img display rule defeated [hidden]
and floated an empty block over the shimmer and its sketching note; the
deal animation left every card at opacity zero in an unfocused tab,
because rAF throttling is real and decoration must never gate content;
and the sketch poll's cache-busting query missed the anchored /img
route, so a landed sketch kept shimmering forever.

The grid is now a snap-scrolling deck: one row in a wide viewport, one
column in a tall one, with edge arrows that appear only on overflow and
page one card at a time.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:10:39 -07:00
github-actions[bot] a07e4ed787 Sync generated provider output 2026-07-28 00:03:34 +00:00
Paul BakausandClaude Fable 5 d89ee5f87c Deal every card the same hand: anatomy, sketches, and the standing door
The decision page compared unlike things: the grounded direction was a
wall of text beside curated catalog art, the catalog art read as a
promise of the build, the weighing silently shrank the challenger set,
and the standing exit hid in the footer under the cards it must not
soften. Every card now shares one anatomy (thesis, palette chips,
material tags, first viewport, case, risk), every dealt challenger is
presented with the weighing written on it rather than applied to it,
the catalog image rides picture-in-picture as labeled inspiration with
the lightbox a click away, and canonCard renders the category standard
as one honest, subordinate card.

When image generation exists, each card declares a sketch slot the page
polls: serve first, generate after, through one shared deliberately
unfinished frame, so the comparison stays about direction instead of
rendering luck. The asset producer takes the batch when subagents
exist; the chosen sketch returns in ANSWER to seed at most one comp
probe, and the comp round still renders its full set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 17:03:05 -07:00
github-actions[bot] 5bec5408e5 Sync generated provider output 2026-07-27 22:52:08 +00:00
Paul BakausandClaude Fable 5 f482d9405e Teach the reading-heavy subagents to write before the ceiling lands
Raising the reviewer's turn budget did not change its fate, only its
reading: 43 tool uses instead of 22, still reaped mid-read with nothing
written, because the SDK ends a run at max-turns without warning and the
model never feels the deadline. The definitions now carry the deadline
themselves: reading is an allowance, batch Reads per turn, take the
decisive inputs first, sample instead of walking the tree, and write by
mid-budget, naming what went unread. A review built from what you saw
beats a perfect review that never arrives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 15:51:36 -07:00
github-actions[bot] d7d07cb0d6 Sync generated provider output 2026-07-27 22:14:03 +00:00
Paul BakausandClaude Fable 5 c9213835e7 Review fidelity against the comp itself, not the builder's summary of it
A codex run turned an approved comp into a related second art direction
and the finish reviewer passed it: the review anchored on the direction
contract, a lossy abstraction the builder wrote, and every element that
abstraction dropped passed silently. Four changes close that chain. The
reviewer inventories the comp's salient elements before reading the
contract and classifies each one (match, adaptation, missing,
contradicted, added without approval), with adaptations citing the
answer, brief, accessibility need, or product truth that forced them,
and fidelity failures outranking craft in material_fixes. The visualize
inventory gate records compositional commitments alongside asset media,
since the 150-word contract cannot carry them. The north-star allowance
now says what it permits: translation, never recomposition. And the
finish sequence recaptures the same viewports once after the fix batch,
so what the documenter records is what actually shipped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 15:13:27 -07:00
github-actions[bot] d52077414c Sync generated provider output 2026-07-27 21:50:35 +00:00
Paul BakausandClaude Fable 5 9e4990765f Give the reading-heavy subagents turn budgets that survive their inputs
A finish review reads the artifact, two full-page screenshots, the
approved comp, the quality-bar cards, and the contract before it may
write a word; at max-turns 12 the SDK reaps it mid-read and the parent
receives the opening sentence as the whole review. Observed twice in a
row (spawn and respawn) on the first real subagent run. The documenter
reads at least as much, and the asset producer pays per asset.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 14:50:03 -07:00
dependabot[bot]andGitHub e0144ed585 Bump the bun-minor-and-patch group with 9 updates (#429)
Prepared with AI assistance from OpenAI Codex under maintainer automation instructions.
2026-07-27 10:24:55 -07:00
github-actions[bot] 839dd10079 Sync generated provider output 2026-07-27 17:07:25 +00:00
Paul BakausandGitHub 9b613ef931 Merge pull request #419 from pbakaus/diff-base-detection
Detect the diff base in context-signals instead of assuming main/master
2026-07-27 10:06:48 -07:00
Vinaywho c9c0fc887b Merge remote-tracking branch 'upstream/main' into fix/detect-system-chrome-gpu-window
# Conflicts:
#	scripts/test-suites.mjs
2026-07-27 15:13:14 +05:30
Vinaywho a4b691c5a2 detect: preserve system-Chrome launch error as fallback cause 2026-07-27 15:12:29 +05:30
Paul BakausandClaude Code 01d5d357c5 The develop candidate leads with an advertised develop default rev
Round eight closes the stale-local class completely: the develop
candidate sits before the remote-default entries, so when origin/HEAD
itself points at develop, its name claim let a stale local develop win
over the fresher origin/develop. The candidate now leads with any
remote-advertised develop rev, exactly as the remote-default and
upstream candidates already lead with theirs. main/master were already
covered since their remote-default entries come first in the order.
Failing-first test forces local develop two commits behind.

Prepared with AI assistance (Claude Code), directed by @pbakaus.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-26 19:04:48 -07:00