mirror of
https://github.com/pbakaus/impeccable.git
synced 2026-09-12 14:16:28 +03:00
Generate visual cues for document seed color picks
Document seed mode's palette pick used to ask the user to imagine colors from words. Add a visual-cues pipeline: compose four-hex palettes from the brief (mood, references, PRODUCT.md), generate a hero scene plus a cream artifact sheet per concept, crop the sheet into isolated artifacts, and write cues.json so the user can pick a direction by eye. Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
@@ -63,6 +63,8 @@ Thumbs.db
|
||||
.impeccable/config.local.json
|
||||
# API keys for optional image-generation cues (document seed mode).
|
||||
.impeccable/.env
|
||||
# Generated visual-cue images + cues.json (document seed Step 4).
|
||||
.impeccable/visual-cues/
|
||||
**/.impeccable/hook.pending.json
|
||||
.impeccable/provider-smoke/
|
||||
src/__impeccable_provider_smoke_*.html
|
||||
|
||||
@@ -407,7 +407,7 @@ Interview answers are words; a palette is easier picked by eye. Before writing t
|
||||
- **No native path**: pause and {{ask_instruction}} whether the user wants generated visual cues to pick a palette by eye. *"I can generate a few small palette-and-mood images so you choose a direction visually instead of from descriptions. That needs an image-generation API key, stored as `IMAGE_GEN_API_KEY` in `.impeccable/.env`. Add one, or skip straight to the seed?"* If a key arrives, write it to `.impeccable/.env`, confirm that file is listed in the project's `.gitignore` (add it if missing; a committed key is a leak), and ask which provider it belongs to so you call the right API.
|
||||
- **The user opts out, or no key arrives**: go to Step 5 and seed from the answers alone.
|
||||
|
||||
When generation is available, produce **2-4** cue images from the interview answers and Step 2 observations: each carries one palette direction as swatches on the chosen background, one type mood, one texture or motif. These are direction tests, not mocks; vary the hue anchor or color strategy across them, not minor tweaks. Show them, ask which feels closest and what feels off, and carry the pick into the seed as the confirmed color direction. One round; refinement belongs to implementation, not the seed.
|
||||
When generation is available, **stop and load [visual-cues.md](visual-cues.md)** and follow its pipeline; it owns everything from the one-line user announcement and concept drafting through parallel or serial generation, cropping, and `cues.json`. Do not restate its mechanics here or in chat. When its `cues.json` is written, tell the user the cues are ready and **end your turn**. The pick round is a separate later step; Steps 5-6 run after that pick, or immediately when the user opted out of generation.
|
||||
|
||||
### Step 5: Write seed DESIGN.md
|
||||
|
||||
|
||||
@@ -0,0 +1,201 @@
|
||||
# Visual Cues Pipeline
|
||||
|
||||
Loaded by `{{command_prefix}}impeccable document` seed mode (Step 4) when image generation is available. Input: the five seed interview answers, the asset observations from seed Step 2, and PRODUCT.md. Output: cue images plus `cues.json` under `.impeccable/visual-cues/`, ready for the user to pick from by eye in a later round.
|
||||
|
||||
Tell the user once, before generating: *"Generating visual cues; this can take a minute or two."* Then work without narration. Chat carries no per-image commentary, no prompt dumps, no palette tables; the folder is the deliverable.
|
||||
|
||||
## The two images
|
||||
|
||||
Each cue is **two generations by the same agent**, in sequence:
|
||||
|
||||
```text
|
||||
HERO [slug].png (1500x1500) ARTIFACT SHEET masters/[slug]-artifacts.png
|
||||
+---------------------------+ +-------------+-------------+ 1500x1500
|
||||
| | | [obj A] | [obj B] |
|
||||
| one full-bleed scene, | | centered, | centered, |
|
||||
| the product's world, | | ~2/3 of | clear |
|
||||
| all four artifact | | its cell | margins |
|
||||
| objects visible in it, | +-------------+-------------+ 750
|
||||
| four palette colors | | [obj C] | [obj D] |
|
||||
| with named carriers | | | |
|
||||
| | | flat cream #FDFCF6 across |
|
||||
+---------------------------+ | the whole canvas, no cell |
|
||||
saved as-is, NO crop | borders, soft shadows OK |
|
||||
+-------------+-------------+
|
||||
quadrant-cropped into
|
||||
[slug]-2..5.png, NO matting
|
||||
```
|
||||
|
||||
- The **hero** is the visual cue: one atmospheric full-bleed composition that stages the concept's palette and mood; this is what the user will pick between. No grid, no regions: the whole frame is the scene. The four artifact objects all appear inside it.
|
||||
- The **artifact sheet** is the second generation, with the hero attached as the reference image: the same four objects re-photographed individually, one per quadrant, each isolated and centered on one continuous flat warm-cream background, `#FDFCF6`. The objects inherit the hero's materials and colors; only the setting changes. ("Sheet" is our name for the file; the prompt never uses it.)
|
||||
- **No transparency, ever.** No alpha channels, no chroma keys, no "transparent background" in any prompt (image models paint a checkerboard instead). Crops keep the cream; soft contact shadows are welcome, they ground the objects.
|
||||
- **Isolation is a hard rule on the sheet.** Every object centered on its quadrant's center point, filling about two-thirds of it, with clear cream margin on every side: nothing comes near the canvas edge, another object, or the quadrant midlines. The crop cuts exactly at the midlines, so anything crossing one gets clipped.
|
||||
|
||||
Example, hero = a flower atelier's worktable: sheet quadrants carry the wrapping ribbon, a single stem, a row of loose petals, and the ceramic vase. All four are visible in the hero scene.
|
||||
|
||||
## Step 1: Compose the palettes from the brief
|
||||
|
||||
The palette is designed before any image exists; the image stages it. Do **not** generate first and read colors off the result: that yields moody near-monochromes, not a usable system. And do not pull colors from a generator or seed bank: everything a designer would research is already in hand. Work the way designers work, brief first:
|
||||
|
||||
1. **Reread the brief.** PRODUCT.md (personality, audience, positioning, what the product sells or shows), the interview answers (Q1 color strategy and hue anchor, Q4 the three named references, Q5 the anti-reference), and the seed Step 2 asset observations. This is the moodboard material.
|
||||
2. **Write one mood phrase per palette**, specific enough to compose from. Good: "dawn delivery run, cut stems in cold water, the city still gray". Bad: "modern and clean", "warm and inviting"; a phrase that fits any brand composes nothing.
|
||||
3. **Map the mood to color.** Pick each palette's hue territory from what the mood should make this audience feel and what the named references actually look like: color psychology plus reference study, not a random draw. The Q1 hue anchor leads one or two palettes; the others take territories the brief also supports (an adjacent hue, a complement, a dark register, a warm register). Six palettes on one hue are six versions of one mood; the pick round exists so the user can choose between genuinely different color stories. The anti-reference (Q5) is a hard constraint on all six.
|
||||
|
||||
Then compose each palette as four exact hex values with a **60-30-10 balance**. These are website/app colors, headed for tokens, not scene colors:
|
||||
|
||||
- **neutral** (~60%, the dominant): the surface, what most of a screen will be. An off-white or near-white with a temperature tint, or a near-black when the mood calls for dark. Never pure `#FFFFFF` or `#000000`.
|
||||
- **primary** (~30%): the brand color, the mood's main carrier. Must read clearly against the neutral.
|
||||
- **secondary**: structure and support: an adjacent hue, or the primary shifted in lightness and chroma. Visibly a different swatch, not a darker copy of primary.
|
||||
- **tertiary** (~10%, the accent): the most saturated of the four and used smallest; distinct in hue from primary so it keeps signal value.
|
||||
|
||||
Hard rules:
|
||||
|
||||
- **Every color earns its place.** For each role, state in one line what it does and why it fits this product. A color you can't justify in one line gets replaced, not kept because it looks nice. The lines guide your composition; they don't go in chat or `cues.json`.
|
||||
- **Contrast is non-negotiable.** Primary must read clearly on the neutral; tertiary must pop against both. A palette that fails either is not done.
|
||||
- Within one palette, any two roles must be nameable apart at a glance. A dark green primary next to a dark green neutral is one color, not two.
|
||||
- Across the set, every palette takes a **different direction**: a different mood phrase and a different harmony scheme (analogous deepened, complementary accent, dark-dominant, warm neutral with the anchor demoted to accent, near-monochrome with one vivid accent). At most two palettes may share a hue family; if two would look alike as four swatches side by side, replace one.
|
||||
|
||||
Done when: 6 (or 4 when serial, Step 5b) palettes exist, each with a mood phrase from the brief, four hexes with roles and reasons, and no two palettes interchangeable.
|
||||
|
||||
## Step 2: Draft the concepts
|
||||
|
||||
Attach one one-line cue concept to each palette. **Every concept lives in the product's own world.** Reread PRODUCT.md first: what the product is, who it serves, what it sells or shows. The subject of every hero scene comes from that world, named with the product's own nouns. A concept that could belong to any other product is not done; sharpen it until it could only be this brand. Diversity comes from the rest of the matrix:
|
||||
|
||||
- Each concept leads with a different anchor from what you already hold: its palette's mood phrase, the color strategy (interview Q1), the type direction (Q2), the three named references (Q4), PRODUCT.md's positioning and personality, the seed Step 2 asset observations.
|
||||
- Give each concept a distinct **material world**: botanical, ceramic, paper/print, textile, metal, glass, stone, food. The material world is the supporting cast around the product's subject, never a replacement for it. For a florist, ceramic means the vases and the kiln-room shelf behind the arrangement; paper means the wrapping bench; the flowers stay in frame.
|
||||
- Vary the scene's register too: one concept at work (hands mid-task), one at rest (the finished thing displayed), one in detail (macro), one in place (the room). Six takes on one world beats six unrelated worlds.
|
||||
- Name the four artifact objects now. Three tests, each against a known failure: chosen **from inside the scene**, so each plausibly sits in the hero composition; **compact**, so it sits centered in a square-ish quadrant (no edge-to-edge ladles or full-width garlands); carrying **no writing** (no tags, labels, packaging, printed cards, or stationery), because a branded tag invites the model to render typography, and text on an artifact ruins it.
|
||||
- The anti-reference (Q5) is a shared negative constraint on every concept.
|
||||
- Name each concept with a two-word slug (`amber-dusk`, `coastal-glass`). The slug is the cue id in filenames and `cues.json`.
|
||||
|
||||
Done when: every palette has a concept with a slug, a subject from the product's world, a material world no other concept uses, and four named artifacts.
|
||||
|
||||
## Step 3: Build the hero prompt
|
||||
|
||||
Expand each concept into a self-contained hero prompt using this template. Write it like screenplay direction, not a keyword list: subject doing something, in a place, in a light. Name every palette color twice, as a plain-language color and as its hex, and tie each to a physical carrier in the scene; a hex with no carrier gets ignored. Keep prohibitions to the final guardrail line.
|
||||
|
||||
```text
|
||||
One full-bleed photograph, 1500x1500 pixels: [one atmospheric scene from
|
||||
the product's world: subject and what it is doing, setting, time of day].
|
||||
The scene contains [artifact A], [artifact B], [artifact C], and
|
||||
[artifact D], all plainly visible. Lighting: [direction and quality, e.g.
|
||||
"low afternoon window light raking in from the left"]. Camera: [framing
|
||||
and lens, e.g. "85mm still life at waist level, shallow depth of field"].
|
||||
Mood: [two or three adjectives from PRODUCT.md's personality]. The scene
|
||||
is art-directed to a strict four-color story, every color plainly visible:
|
||||
[color name] ([neutral hex]) as the dominant ground and backdrop, about
|
||||
60% of the frame; [color name] ([primary hex]) carried by [the main
|
||||
subject]; [color name] ([secondary hex]) on [a supporting element];
|
||||
[color name] ([tertiary hex]) as one small vivid accent on [a specific
|
||||
object]. Rich, saturated, editorial color. Photorealistic, real texture.
|
||||
No text, no labels, no numbers, no borders, no watermark.
|
||||
```
|
||||
|
||||
Fill every `[bracketed]` slot from the concept; never leave template language in the prompt. Done when every concept has a complete hero prompt whose four colors each have a named carrier.
|
||||
|
||||
## Step 4: Build the artifact-sheet prompt
|
||||
|
||||
The sheet prompt is the same for every concept except the object list. It runs as the **second** generation, with the concept's hero attached as the reference/input image, so the objects match the scene instead of being reinvented. Describe it as plain product photography; any word that implies an editorial layout ("catalog", "sheet", "spread", "grid") invites the model to design a page with titles and captions instead of photographing objects:
|
||||
|
||||
```text
|
||||
Using the attached photograph as the exact reference for objects, materials,
|
||||
and colors: one square photograph, 1500x1500 pixels, of four objects from
|
||||
that scene, each re-photographed individually from directly overhead in soft
|
||||
even studio light. This is a plain photograph of objects resting on a bare
|
||||
surface. Absolutely no text anywhere in the image: no letters, no words, no
|
||||
numbers, no labels, no captions, no title.
|
||||
|
||||
The surface is one perfectly flat warm-cream background, hex #FDFCF6,
|
||||
continuous edge to edge. Picture the canvas divided into four equal
|
||||
quadrants: [artifact A] top-left, [artifact B] top-right, [artifact C]
|
||||
bottom-left, [artifact D] bottom-right.
|
||||
|
||||
Each object sits exactly centered on its quadrant's center point, filling
|
||||
about two-thirds of its quadrant, with clear cream margin on every side:
|
||||
nothing comes anywhere near the canvas edges or the horizontal and vertical
|
||||
centerlines of the image. A soft gentle contact shadow under each object is
|
||||
welcome.
|
||||
|
||||
No frames, no cell borders, no dividing lines, no watermark, no typography
|
||||
of any kind; the background stays one uninterrupted #FDFCF6 everywhere.
|
||||
```
|
||||
|
||||
The cream is constant across all cues (`#FDFCF6`; the artifacts land on this exact surface downstream), so never restyle it per concept. Done when every concept has its sheet prompt with the four artifacts assigned to quadrants.
|
||||
|
||||
## Step 5a: Generate in parallel (harness has subagents)
|
||||
|
||||
If the harness has any subagent/spawn tool, parallel is **required**: run all 6 concepts as one **wave**, every spawn call emitted in the same tool-call round, one concept per subagent, each with a self-contained task (subagents start without your context). **Never generate the images yourself one at a time when a subagent tool exists**; a serial loop in a subagent-capable harness is a failure, not a fallback. Attach the harness's image-generation skill to each spawn when the harness expects that (Codex: the `imagegen` skill).
|
||||
|
||||
Spawn each subagent with this task, filling the slots:
|
||||
|
||||
```text
|
||||
You generate exactly two images in sequence and report. Use the harness's
|
||||
native image generation tool. Do not fall back to CLIs or APIs; do not edit
|
||||
repo files.
|
||||
|
||||
1. Generate the HERO image, 1500x1500 (or the nearest supported square),
|
||||
with this prompt:
|
||||
|
||||
[the full Step 3 hero prompt for this concept]
|
||||
|
||||
2. Generate the ARTIFACT SHEET, same size, passing the hero image you just
|
||||
generated as the reference/input image (the tool's image-edit or
|
||||
reference-image mode), with this prompt:
|
||||
|
||||
[the full Step 4 sheet prompt for this concept]
|
||||
|
||||
3. Look at the sheet you generated. If any object crosses the canvas edge or
|
||||
the horizontal or vertical centerline, or any text appears anywhere,
|
||||
regenerate the ARTIFACT SHEET once: same reference image, same prompt,
|
||||
plus this line appended: "Make every object smaller, at most half of its
|
||||
quadrant, pulled in tight to its quadrant's center, with even more empty
|
||||
cream between the objects and around the edges." Never retry more than
|
||||
once; keep the second sheet regardless.
|
||||
4. Reply with exactly these three lines and nothing else:
|
||||
|
||||
COMPLETED [slug]
|
||||
HERO [absolute path to the hero PNG]
|
||||
ARTIFACTS [absolute path to the final sheet PNG]
|
||||
|
||||
If either generation fails, reply instead with one line:
|
||||
ERROR [slug] [short reason]
|
||||
```
|
||||
|
||||
Six concepts fit the observed Codex ceiling of 6 concurrent subagents, so one wave normally covers everything. If a spawn is rejected with a thread-limit error, collect the accepted wave, close those agents to release their slots, then run a second wave for the rejects. If every spawn in the first wave ERRORs because subagents lack the image tool, fall back to Step 5b. Close every agent after collecting its report.
|
||||
|
||||
Done when: every concept has either a three-line COMPLETED report or an ERROR line. An ERROR concept is dropped, not retried more than once; five good cues beat a stalled pipeline.
|
||||
|
||||
## Step 5b: Generate in series (no subagents)
|
||||
|
||||
Only when the harness has no subagent tool at all: generate **4** concepts yourself, one pair at a time (hero, then sheet with the hero as reference), with the same Step 3 and Step 4 prompts. After each pair: apply the same look-and-retry rule as the subagent task, and record the same facts a subagent would report (slug, both paths). Same done-condition as 5a, over 4 concepts.
|
||||
|
||||
## Step 6: Crop and compile
|
||||
|
||||
For each completed concept, run one command, carrying the slug and that concept's **planned palette from Step 1** (you composed it; no subagent echo needed):
|
||||
|
||||
```text
|
||||
node {{scripts_path}}/visual-cues.mjs crop [hero.png] [artifacts.png] \
|
||||
--slug [slug] \
|
||||
--palette "primary=#RRGGBB;secondary=#RRGGBB;tertiary=#RRGGBB;neutral=#RRGGBB" \
|
||||
--out .impeccable/visual-cues
|
||||
```
|
||||
|
||||
The script copies the hero untouched to `[slug].png`, keeps the sheet under `masters/[slug]-artifacts.png`, and quadrant-crops the sheet into `[slug]-2.png` through `[slug]-5.png` (cream stays; nothing is matted). For each palette role it searches the hero for the closest rendered pixel (`snapped`, with its hero position), then updates `cues.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"cues": ["amber-dusk", "coastal-glass"],
|
||||
"supporting-artifacts": {
|
||||
"amber-dusk": ["amber-dusk-2", "amber-dusk-3", "amber-dusk-4", "amber-dusk-5"]
|
||||
},
|
||||
"palette": {
|
||||
"amber-dusk": { "primary": { "hex": "#B8422E", "snapped": "#B4402F", "at": [312, 540] } }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Done when: `cues.json` lists one entry per completed concept and every listed slug has its five PNGs on disk (hero plus four artifacts).
|
||||
|
||||
## Step 7: Pause
|
||||
|
||||
Tell the user in one or two lines that the visual cues are ready at `.impeccable/visual-cues/` (name the count), then end your turn. The pick round is a separate later step: do not show or describe the images, do not ask which the user prefers, and do not write DESIGN.md in this turn.
|
||||
@@ -0,0 +1,397 @@
|
||||
#!/usr/bin/env node
|
||||
// visual-cues.mjs — crop + compile for document seed visual cues.
|
||||
// Pipeline doc: skill/reference/visual-cues.md (canonical; this help text is not).
|
||||
//
|
||||
// Each cue is two images: a full-bleed hero scene and an artifact sheet
|
||||
// (four objects on one flat cream canvas, one per quadrant). No alpha, no
|
||||
// chroma key: crops keep the cream.
|
||||
//
|
||||
// node visual-cues.mjs crop <hero.png> <artifacts.png> --slug <two-word-slug>
|
||||
// [--palette "primary=#RRGGBB;secondary=...;tertiary=...;neutral=..."]
|
||||
// [--out <dir>] (default: .impeccable/visual-cues)
|
||||
// Copies the hero untouched to <slug>.png, keeps the sheet under
|
||||
// <out>/masters/<slug>-artifacts.png, quadrant-crops the sheet into
|
||||
// <slug>-2..5.png, finds each planned palette hex's closest pixel in
|
||||
// the hero, and updates <out>/cues.json.
|
||||
//
|
||||
// Dependency-free: PNG decode/encode on node:zlib. Rejects interlaced and
|
||||
// indexed-color PNGs; convert those with sips/ImageMagick/PIL first.
|
||||
|
||||
import { readFileSync, writeFileSync, mkdirSync, copyFileSync, existsSync } from 'node:fs';
|
||||
import { join, resolve } from 'node:path';
|
||||
import { pathToFileURL } from 'node:url';
|
||||
import zlib from 'node:zlib';
|
||||
|
||||
// ---------------------------------------------------------------- PNG codec
|
||||
|
||||
const PNG_SIG = Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]);
|
||||
|
||||
// Every PNG chunk carries a CRC-32 trailer (the spec's fixed polynomial,
|
||||
// 0xedb88320); precompute the 256-entry lookup table once instead of doing
|
||||
// the bit-by-bit division per byte.
|
||||
const CRC_TABLE = (() => {
|
||||
const t = new Int32Array(256);
|
||||
for (let n = 0; n < 256; n++) {
|
||||
let c = n;
|
||||
for (let k = 0; k < 8; k++) c = c & 1 ? 0xedb88320 ^ (c >>> 1) : c >>> 1;
|
||||
t[n] = c;
|
||||
}
|
||||
return t;
|
||||
})();
|
||||
|
||||
function crc32(buf) {
|
||||
let c = 0xffffffff;
|
||||
for (let i = 0; i < buf.length; i++) c = CRC_TABLE[(c ^ buf[i]) & 0xff] ^ (c >>> 8);
|
||||
return (c ^ 0xffffffff) >>> 0;
|
||||
}
|
||||
|
||||
// PNG filter type 4 (Paeth): predicts a byte from its left (a), above (b),
|
||||
// and above-left (c) neighbors, picking whichever of a, b, or a+b-c lands
|
||||
// closest to the actual gradient. Used only by decodePng's unfilter step;
|
||||
// encodePng always writes filter 0, so it never needs the inverse.
|
||||
function paeth(a, b, c) {
|
||||
const p = a + b - c;
|
||||
const pa = Math.abs(p - a);
|
||||
const pb = Math.abs(p - b);
|
||||
const pc = Math.abs(p - c);
|
||||
if (pa <= pb && pa <= pc) return a;
|
||||
if (pb <= pc) return b;
|
||||
return c;
|
||||
}
|
||||
|
||||
export function decodePng(buf) {
|
||||
if (!buf.subarray(0, 8).equals(PNG_SIG)) throw new Error('not a PNG file');
|
||||
// Walk the chunk stream: each chunk is [4-byte length][4-byte type][data][4-byte crc].
|
||||
// IHDR carries the header fields; IDAT is the (possibly multi-chunk)
|
||||
// compressed pixel data, concatenated below before inflating; other
|
||||
// chunk types (tEXt, iCCP, etc.) are skipped since nothing here needs them.
|
||||
let pos = 8;
|
||||
let ihdr = null;
|
||||
const idat = [];
|
||||
while (pos + 8 <= buf.length) {
|
||||
const len = buf.readUInt32BE(pos);
|
||||
const type = buf.toString('ascii', pos + 4, pos + 8);
|
||||
const data = buf.subarray(pos + 8, pos + 8 + len);
|
||||
if (type === 'IHDR') {
|
||||
ihdr = {
|
||||
width: data.readUInt32BE(0),
|
||||
height: data.readUInt32BE(4),
|
||||
bitDepth: data[8],
|
||||
colorType: data[9],
|
||||
interlace: data[12],
|
||||
};
|
||||
} else if (type === 'IDAT') {
|
||||
idat.push(data);
|
||||
} else if (type === 'IEND') {
|
||||
break;
|
||||
}
|
||||
pos += 12 + len; // length + type + data + crc
|
||||
}
|
||||
if (!ihdr) throw new Error('PNG has no IHDR chunk');
|
||||
const { width, height, bitDepth, colorType, interlace } = ihdr;
|
||||
if (interlace) throw new Error('interlaced PNG not supported; re-save without interlacing (sips, ImageMagick, or PIL)');
|
||||
if (colorType === 3) throw new Error('indexed-color PNG not supported; convert to RGB/RGBA first (sips, ImageMagick, or PIL)');
|
||||
if (bitDepth !== 8 && bitDepth !== 16) throw new Error(`unsupported bit depth ${bitDepth}; convert to 8-bit first`);
|
||||
const channels = { 0: 1, 2: 3, 4: 2, 6: 4 }[colorType];
|
||||
if (!channels) throw new Error(`unsupported color type ${colorType}`);
|
||||
|
||||
const sampleBytes = bitDepth / 8;
|
||||
const bpp = channels * sampleBytes; // bytes per pixel
|
||||
const stride = width * bpp; // bytes per scanline, excluding the filter-type byte
|
||||
const raw = zlib.inflateSync(Buffer.concat(idat));
|
||||
|
||||
// Each scanline in the inflated stream is prefixed with a 1-byte filter
|
||||
// type (0-4) that says how it was delta-encoded against the row above
|
||||
// and/or the pixel to the left; undo that in place, row by row, since
|
||||
// filter 2-4 need the already-unfiltered previous row to reconstruct.
|
||||
const px = Buffer.alloc(height * stride);
|
||||
let rp = 0;
|
||||
for (let y = 0; y < height; y++) {
|
||||
const filter = raw[rp++];
|
||||
const row = px.subarray(y * stride, (y + 1) * stride);
|
||||
raw.copy(row, 0, rp, rp + stride);
|
||||
rp += stride;
|
||||
const prev = y > 0 ? px.subarray((y - 1) * stride, y * stride) : null;
|
||||
if (filter === 0) continue; // None: bytes are already the real pixel values
|
||||
if (filter === 1) {
|
||||
// Sub: each byte was stored as (value - left).
|
||||
for (let i = bpp; i < stride; i++) row[i] = (row[i] + row[i - bpp]) & 0xff;
|
||||
} else if (filter === 2) {
|
||||
// Up: each byte was stored as (value - above).
|
||||
if (prev) for (let i = 0; i < stride; i++) row[i] = (row[i] + prev[i]) & 0xff;
|
||||
} else if (filter === 3) {
|
||||
// Average: each byte was stored as (value - floor((left + above) / 2)).
|
||||
for (let i = 0; i < stride; i++) {
|
||||
const left = i >= bpp ? row[i - bpp] : 0;
|
||||
const up = prev ? prev[i] : 0;
|
||||
row[i] = (row[i] + ((left + up) >> 1)) & 0xff;
|
||||
}
|
||||
} else if (filter === 4) {
|
||||
// Paeth: each byte was stored as (value - paeth(left, above, above-left)).
|
||||
for (let i = 0; i < stride; i++) {
|
||||
const a = i >= bpp ? row[i - bpp] : 0;
|
||||
const b = prev ? prev[i] : 0;
|
||||
const c = prev && i >= bpp ? prev[i - bpp] : 0;
|
||||
row[i] = (row[i] + paeth(a, b, c)) & 0xff;
|
||||
}
|
||||
} else {
|
||||
throw new Error(`unknown PNG filter ${filter} at row ${y}`);
|
||||
}
|
||||
}
|
||||
|
||||
// Normalize every supported color type (grayscale, RGB, grayscale+alpha,
|
||||
// RGBA) down to one consistent RGBA8 buffer, so everything past this
|
||||
// point (crop, palette search, re-encode) only ever deals with one shape.
|
||||
// 16-bit samples keep only the high byte; visual cues never need more
|
||||
// than 8 bits of precision per channel.
|
||||
const rgba = Buffer.alloc(width * height * 4);
|
||||
const at = (base, ch) => px[base + ch * sampleBytes];
|
||||
for (let i = 0; i < width * height; i++) {
|
||||
const base = i * bpp;
|
||||
let r, g, b, a;
|
||||
if (colorType === 0) {
|
||||
r = g = b = at(base, 0);
|
||||
a = 255;
|
||||
} else if (colorType === 2) {
|
||||
r = at(base, 0); g = at(base, 1); b = at(base, 2);
|
||||
a = 255;
|
||||
} else if (colorType === 4) {
|
||||
r = g = b = at(base, 0);
|
||||
a = at(base, 1);
|
||||
} else {
|
||||
r = at(base, 0); g = at(base, 1); b = at(base, 2); a = at(base, 3);
|
||||
}
|
||||
const o = i * 4;
|
||||
rgba[o] = r; rgba[o + 1] = g; rgba[o + 2] = b; rgba[o + 3] = a;
|
||||
}
|
||||
return { width, height, rgba, hasAlpha: colorType === 4 || colorType === 6 };
|
||||
}
|
||||
|
||||
// Wraps one chunk's payload with its length header, type tag, and CRC
|
||||
// trailer, matching the layout decodePng's chunk walk expects.
|
||||
function pngChunk(type, data) {
|
||||
const out = Buffer.alloc(12 + data.length);
|
||||
out.writeUInt32BE(data.length, 0);
|
||||
out.write(type, 4, 'ascii');
|
||||
data.copy(out, 8);
|
||||
out.writeUInt32BE(crc32(out.subarray(4, 8 + data.length)), 8 + data.length);
|
||||
return out;
|
||||
}
|
||||
|
||||
// Always writes 8-bit RGBA with filter type 0 (None) on every scanline: the
|
||||
// crops here are small and this script has no bandwidth concerns, so the
|
||||
// simplicity of never predicting/unpredicting bytes outweighs the larger
|
||||
// file size a real filter choice would save.
|
||||
export function encodePng(rgba, width, height) {
|
||||
const ihdr = Buffer.alloc(13);
|
||||
ihdr.writeUInt32BE(width, 0);
|
||||
ihdr.writeUInt32BE(height, 4);
|
||||
ihdr[8] = 8; // bit depth
|
||||
ihdr[9] = 6; // color type 6 = RGBA
|
||||
const stride = width * 4;
|
||||
// One extra byte per row for the filter-type prefix (always 0 here).
|
||||
const raw = Buffer.alloc((stride + 1) * height);
|
||||
for (let y = 0; y < height; y++) {
|
||||
raw[y * (stride + 1)] = 0; // filter: None
|
||||
rgba.copy(raw, y * (stride + 1) + 1, y * stride, (y + 1) * stride);
|
||||
}
|
||||
const idat = zlib.deflateSync(raw, { level: 9 });
|
||||
return Buffer.concat([PNG_SIG, pngChunk('IHDR', ihdr), pngChunk('IDAT', idat), pngChunk('IEND', Buffer.alloc(0))]);
|
||||
}
|
||||
|
||||
// ------------------------------------------------------------ quadrant math
|
||||
|
||||
// The artifact sheet is one 2x2 grid on a flat cream canvas, one object per
|
||||
// quadrant, in reading order: q2 top-left, q3 top-right, q4 bottom-left,
|
||||
// q5 bottom-right. Proportional, so any square-ish sheet cuts the same way.
|
||||
export function quadrants(width, height) {
|
||||
const mx = Math.round(width / 2);
|
||||
const my = Math.round(height / 2);
|
||||
return {
|
||||
q2: { x: 0, y: 0, w: mx, h: my },
|
||||
q3: { x: mx, y: 0, w: width - mx, h: my },
|
||||
q4: { x: 0, y: my, w: mx, h: height - my },
|
||||
q5: { x: mx, y: my, w: width - mx, h: height - my },
|
||||
};
|
||||
}
|
||||
|
||||
// Copies one rectangle r = {x, y, w, h} out of img.rgba, row by row (rows
|
||||
// aren't contiguous across the crop boundary in the source buffer).
|
||||
function cropRegion(img, r) {
|
||||
const out = Buffer.alloc(r.w * r.h * 4);
|
||||
for (let y = 0; y < r.h; y++) {
|
||||
const src = ((r.y + y) * img.width + r.x) * 4;
|
||||
img.rgba.copy(out, y * r.w * 4, src, src + r.w * 4);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// ----------------------------------------------------------------- palette
|
||||
|
||||
// role=#RRGGBB per entry; a legacy trailing @x,y is accepted and ignored
|
||||
// (the search below beats model-reported coordinates every time).
|
||||
const PALETTE_ENTRY = /^([a-z][a-z-]*)=(#[0-9a-fA-F]{6})(?:@\d+,\d+)?$/;
|
||||
|
||||
function parsePalette(str) {
|
||||
const out = {};
|
||||
for (const part of str.split(';')) {
|
||||
const m = part.trim().match(PALETTE_ENTRY);
|
||||
if (!m) throw new Error(`bad palette entry "${part.trim()}" (expected role=#RRGGBB)`);
|
||||
out[m[1]] = { hex: m[2].toUpperCase() };
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// The parent designed the palette, so the planned hex is known; what needs
|
||||
// measuring is where and how faithfully the hero staged it. Search the whole
|
||||
// hero for the pixel closest to each planned hex. hex stays the planned
|
||||
// value; snapped is the closest rendered pixel; at is its hero position.
|
||||
function snapPalette(img, palette) {
|
||||
const out = {};
|
||||
// Sample on a grid instead of every pixel: ~150 samples per axis is dense
|
||||
// enough to find a representative patch of any staged color, and scanning
|
||||
// a 1500x1500 hero at full resolution for every role adds up otherwise.
|
||||
const step = Math.max(1, Math.floor(Math.min(img.width, img.height) / 150));
|
||||
for (const [role, entry] of Object.entries(palette)) {
|
||||
const pr = parseInt(entry.hex.slice(1, 3), 16);
|
||||
const pg = parseInt(entry.hex.slice(3, 5), 16);
|
||||
const pb = parseInt(entry.hex.slice(5, 7), 16);
|
||||
let best = Infinity;
|
||||
let bx = 0;
|
||||
let by = 0;
|
||||
// Squared Euclidean distance in RGB space; skipping the sqrt is fine
|
||||
// since only the relative ordering of distances matters here.
|
||||
for (let y = 0; y < img.height; y += step) {
|
||||
for (let x = 0; x < img.width; x += step) {
|
||||
const o = (y * img.width + x) * 4;
|
||||
const dr = img.rgba[o] - pr;
|
||||
const dg = img.rgba[o + 1] - pg;
|
||||
const db = img.rgba[o + 2] - pb;
|
||||
const d = dr * dr + dg * dg + db * db;
|
||||
if (d < best) { best = d; bx = x; by = y; }
|
||||
}
|
||||
}
|
||||
const o = (by * img.width + bx) * 4;
|
||||
const snapped = `#${[img.rgba[o], img.rgba[o + 1], img.rgba[o + 2]]
|
||||
.map((v) => v.toString(16).padStart(2, '0'))
|
||||
.join('')
|
||||
.toUpperCase()}`;
|
||||
out[role] = { hex: entry.hex, snapped, at: [bx, by] };
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------- cues.json
|
||||
|
||||
// Reads the existing cues.json (if any) and merges this cue in, so cropping
|
||||
// the six concepts one after another accumulates into one shared manifest
|
||||
// instead of each crop overwriting the last.
|
||||
function updateCuesJson(outDir, slug, artifactIds, palette) {
|
||||
const path = join(outDir, 'cues.json');
|
||||
let data = {};
|
||||
if (existsSync(path)) data = JSON.parse(readFileSync(path, 'utf8'));
|
||||
data.cues = data.cues || [];
|
||||
data['supporting-artifacts'] = data['supporting-artifacts'] || {};
|
||||
if (!data.cues.includes(slug)) data.cues.push(slug);
|
||||
data['supporting-artifacts'][slug] = artifactIds;
|
||||
if (palette) {
|
||||
data.palette = data.palette || {};
|
||||
data.palette[slug] = palette;
|
||||
}
|
||||
writeFileSync(path, JSON.stringify(data, null, 2) + '\n');
|
||||
return data;
|
||||
}
|
||||
|
||||
// -------------------------------------------------------------------- CLI
|
||||
|
||||
// Minimal flag parser: positional args collect into `_`, everything after
|
||||
// a `--name` becomes args.name. Good enough for this script's small,
|
||||
// fixed set of options; no need for a dependency here.
|
||||
function parseArgs(argv) {
|
||||
const args = { _: [] };
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
if (argv[i].startsWith('--')) {
|
||||
args[argv[i].slice(2)] = argv[i + 1];
|
||||
i++;
|
||||
} else {
|
||||
args._.push(argv[i]);
|
||||
}
|
||||
}
|
||||
return args;
|
||||
}
|
||||
|
||||
// Errors surface as JSON on stderr (matching the success shape on stdout)
|
||||
// so the calling agent can parse either outcome the same way.
|
||||
function fail(msg) {
|
||||
console.error(JSON.stringify({ ok: false, error: msg }));
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
function cmdCrop(args) {
|
||||
const [heroFile, sheetFile] = args._;
|
||||
const slug = args.slug;
|
||||
if (!heroFile || !sheetFile || !slug) {
|
||||
fail('usage: visual-cues.mjs crop <hero.png> <artifacts.png> --slug <slug> [--palette "..."] [--out <dir>]');
|
||||
}
|
||||
if (!/^[a-z0-9]+(-[a-z0-9]+)+$/.test(slug)) fail(`slug "${slug}" must be lowercase words joined by hyphens (e.g. amber-dusk)`);
|
||||
const outDir = resolve(args.out || '.impeccable/visual-cues');
|
||||
const hero = decodePng(readFileSync(resolve(heroFile)));
|
||||
const sheet = decodePng(readFileSync(resolve(sheetFile)));
|
||||
|
||||
mkdirSync(join(outDir, 'masters'), { recursive: true });
|
||||
const heroPath = join(outDir, `${slug}.png`);
|
||||
copyFileSync(resolve(heroFile), heroPath); // the hero ships untouched, no crop
|
||||
const keptSheet = join(outDir, 'masters', `${slug}-artifacts.png`);
|
||||
copyFileSync(resolve(sheetFile), keptSheet); // uncropped sheet, kept for reference
|
||||
|
||||
// q2..q5 in reading order (top-left, top-right, bottom-left, bottom-right)
|
||||
// become <slug>-2.png..<slug>-5.png, matching the numbering documented in
|
||||
// reference/visual-cues.md and expected by cues.json readers.
|
||||
const qs = quadrants(sheet.width, sheet.height);
|
||||
const files = [heroPath];
|
||||
const artifactIds = [];
|
||||
const order = ['q2', 'q3', 'q4', 'q5'];
|
||||
for (let i = 0; i < order.length; i++) {
|
||||
const r = qs[order[i]];
|
||||
const id = `${slug}-${i + 2}`;
|
||||
artifactIds.push(id);
|
||||
const outPath = join(outDir, `${id}.png`);
|
||||
writeFileSync(outPath, encodePng(cropRegion(sheet, r), r.w, r.h));
|
||||
files.push(outPath);
|
||||
}
|
||||
|
||||
// --palette is optional: the agent may crop before it has finished
|
||||
// designing the palette, and can re-run crop later once it has hexes.
|
||||
let palette = null;
|
||||
if (args.palette) palette = snapPalette(hero, parsePalette(args.palette));
|
||||
|
||||
updateCuesJson(outDir, slug, artifactIds, palette);
|
||||
|
||||
console.log(JSON.stringify({
|
||||
ok: true,
|
||||
slug,
|
||||
hero: heroPath,
|
||||
artifacts: keptSheet,
|
||||
files,
|
||||
palette,
|
||||
cuesJson: join(outDir, 'cues.json'),
|
||||
}, null, 2));
|
||||
}
|
||||
|
||||
function main() {
|
||||
const [cmd, ...rest] = process.argv.slice(2);
|
||||
const args = parseArgs(rest);
|
||||
try {
|
||||
if (cmd === 'crop') cmdCrop(args);
|
||||
else fail('usage: visual-cues.mjs crop <hero.png> <artifacts.png> --slug <slug> [options] (see reference/visual-cues.md)');
|
||||
} catch (err) {
|
||||
fail(err.message);
|
||||
}
|
||||
}
|
||||
|
||||
// Only auto-run when invoked directly (`node visual-cues.mjs ...`), not
|
||||
// when another module imports its exports (decodePng, encodePng, etc.),
|
||||
// e.g. from a test file.
|
||||
if (process.argv[1] && import.meta.url === pathToFileURL(resolve(process.argv[1])).href) {
|
||||
main();
|
||||
}
|
||||
Reference in New Issue
Block a user