diff --git a/skills/.curated/imagegen/SKILL.md b/skills/.curated/imagegen/SKILL.md index 2e3d8bd..8819588 100644 --- a/skills/.curated/imagegen/SKILL.md +++ b/skills/.curated/imagegen/SKILL.md @@ -1,81 +1,144 @@ --- name: "imagegen" -description: "Use when the user asks to generate or edit images via the OpenAI Image API (for example: generate image, edit/inpaint/mask, background removal or replacement, transparent background, product shots, concept art, covers, or batch variants); run the bundled CLI (`scripts/image_gen.py`) and require `OPENAI_API_KEY` for live calls." +description: "Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output should be a bitmap asset rather than repo-native code or vector. Do not use when the task is better handled by editing existing SVG/vector/code-native assets, extending an established icon or logo system, or building the visual directly in HTML/CSS/canvas." --- - # Image Generation Skill -Generates or edits images for the current project (e.g., website assets, game assets, UI mockups, product mockups, wireframes, logo design, photorealistic images, infographics). Defaults to `gpt-image-1.5` and the OpenAI Image API, and prefers the bundled CLI for deterministic, reproducible runs. +Generates or edits images for the current project (for example website assets, game assets, UI mockups, product mockups, wireframes, logo design, photorealistic images, or infographics). + +## Top-level modes and rules + +This skill has exactly two top-level modes: + +- **Default built-in tool mode (preferred):** built-in `image_gen` tool for normal image generation and editing. Does not require `OPENAI_API_KEY`. +- **Fallback CLI mode (explicit-only):** `scripts/image_gen.py` CLI. Use only when the user explicitly asks for the CLI path. Requires `OPENAI_API_KEY`. + +Within the explicit CLI fallback only, the CLI exposes three subcommands: + +- `generate` +- `edit` +- `generate-batch` + +Rules: +- Use the built-in `image_gen` tool by default for all normal image generation and editing requests. +- Never switch to CLI fallback automatically. +- If the built-in tool fails or is unavailable, tell the user the CLI fallback exists and that it requires `OPENAI_API_KEY`. Proceed only if the user explicitly asks for that fallback. +- If the user explicitly asks for CLI mode, use the bundled `scripts/image_gen.py` workflow. Do not create one-off SDK runners. +- Never modify `scripts/image_gen.py`. If something is missing, ask the user before doing anything else. + +Built-in save-path policy: +- In built-in tool mode, Codex saves generated images under `$CODEX_HOME/*` by default. +- Do not describe or rely on OS temp as the default built-in destination. +- Do not describe or rely on a destination-path argument (if any) on the built-in `image_gen` tool. If a specific location is needed, generate first and then move or copy the selected output from `$CODEX_HOME/generated_images/...`. +- Save-path precedence in built-in mode: + 1. If the user names a destination, move or copy the selected output there. + 2. If the image is meant for the current project, move or copy the final selected image into the workspace before finishing. + 3. If the image is only for preview or brainstorming, render it inline; the underlying file can remain at the default `$CODEX_HOME/*` path. +- Never leave a project-referenced asset only at the default `$CODEX_HOME/*` path. +- Do not overwrite an existing asset unless the user explicitly asked for replacement; otherwise create a sibling versioned filename such as `hero-v2.png` or `item-icon-edited.png`. + +Shared prompt guidance for both modes lives in `references/prompting.md` and `references/sample-prompts.md`. + +Fallback-only docs/resources for CLI mode: +- `references/cli.md` +- `references/image-api.md` +- `references/codex-network.md` +- `scripts/image_gen.py` ## When to use - Generate a new image (concept art, product shot, cover, website hero) -- Edit an existing image (inpainting, masked edits, lighting or weather transformations, background replacement, object removal, compositing, transparent background) -- Batch runs (many prompts, or many variants across prompts) +- Generate a new image using one or more reference images for style, composition, or mood +- Edit an existing image (inpainting, lighting or weather transformations, background replacement, object removal, compositing, transparent background) +- Produce many assets or variants for one task -## Decision tree (generate vs edit vs batch) -- If the user provides an input image (or says “edit/retouch/inpaint/mask/translate/localize/change only X”) → **edit** -- Else if the user needs many different prompts/assets → **generate-batch** -- Else → **generate** +## When not to use +- Extending or matching an existing SVG/vector icon set, logo system, or illustration library inside the repo +- Creating simple shapes, diagrams, wireframes, or icons that are better produced directly in SVG, HTML/CSS, or canvas +- Making a small project-local asset edit when the source file already exists in an editable native format +- Any task where the user clearly wants deterministic code-native output instead of a generated bitmap + +## Decision tree + +Think about two separate questions: + +1. **Intent:** is this a new image or an edit of an existing image? +2. **Execution strategy:** is this one asset or many assets/variants? + +Intent: +- If the user wants to modify an existing image while preserving parts of it, treat the request as **edit**. +- If the user provides images only as references for style, composition, mood, or subject guidance, treat the request as **generate**. +- If the user provides no images, treat the request as **generate**. + +Built-in edit semantics: +- Built-in edit mode is for images already visible in the conversation context, such as attached images or images generated earlier in the thread. +- If the user wants to edit a local image file with the built-in tool, first load it with built-in `view_image` tool so the image is visible in the conversation context, then proceed with the built-in edit flow. +- Do not promise arbitrary filesystem-path editing through the built-in tool. +- If a local file still needs direct file-path control, masks, or other explicit CLI-only parameters, use the explicit CLI fallback only when the user asks for it. +- For edits, preserve invariants aggressively and save non-destructively by default. + +Execution strategy: +- In the built-in default path, produce many assets or variants by issuing one `image_gen` call per requested asset or variant. +- In the explicit CLI fallback path, use the CLI `generate-batch` subcommand only when the user explicitly chose CLI mode and needs many prompts/assets. + +Assume the user wants a new image unless they clearly ask to change an existing one. ## Workflow -1. Decide intent: generate vs edit vs batch (see decision tree above). -2. Collect inputs up front: prompt(s), exact text (verbatim), constraints/avoid list, and any input image(s)/mask(s). For multi-image edits, label each input by index and role; for edits, list invariants explicitly. -3. If batch: write a temporary JSONL under tmp/ (one job per line), run once, then delete the JSONL. -4. Augment prompt into a short labeled spec (structure + constraints) without inventing new creative requirements. -5. Run the bundled CLI (`scripts/image_gen.py`) with sensible defaults (see references/cli.md). -6. For complex edits/generations, inspect outputs (open/view images) and validate: subject, style, composition, text accuracy, and invariants/avoid items. -7. Iterate: make a single targeted change (prompt or mask), re-run, re-check. -8. Save/return final outputs and note the final prompt + flags used. - -## Temp and output conventions -- Use `tmp/imagegen/` for intermediate files (for example JSONL batches); delete when done. -- Write final artifacts under `output/imagegen/` when working in this repo. -- Use `--out` or `--out-dir` to control output paths; keep filenames stable and descriptive. - -## Dependencies (install if missing) -Prefer `uv` for dependency management. - -Python packages: -``` -uv pip install openai pillow -``` -If `uv` is unavailable: -``` -python3 -m pip install openai pillow -``` - -## Environment -- `OPENAI_API_KEY` must be set for live API calls. - -If the key is missing, give the user these steps: -1. Create an API key in the OpenAI platform UI: https://platform.openai.com/api-keys -2. Set `OPENAI_API_KEY` as an environment variable in their system. -3. Offer to guide them through setting the environment variable for their OS/shell if needed. -- Never ask the user to paste the full key in chat. Ask them to set it locally and confirm when ready. - -If installation isn't possible in this environment, tell the user which dependency is missing and how to install it locally. - -## Defaults & rules -- Use `gpt-image-1.5` unless the user explicitly asks for `gpt-image-1-mini` or explicitly prefers a cheaper/faster model. -- Assume the user wants a new image unless they explicitly ask for an edit. -- Require `OPENAI_API_KEY` before any live API call. -- Use the OpenAI Python SDK (`openai` package) for all API calls; do not use raw HTTP. -- If the user requests edits, use `client.images.edit(...)` and include input images (and mask if provided). -- Prefer the bundled CLI (`scripts/image_gen.py`) over writing new one-off scripts. -- Never modify `scripts/image_gen.py`. If something is missing, ask the user before doing anything else. -- If the result isn’t clearly relevant or doesn’t satisfy constraints, iterate with small targeted prompt changes; only ask a question if a missing detail blocks success. +1. Decide the top-level mode: built-in by default, fallback CLI only if explicitly requested. +2. Decide the intent: `generate` or `edit`. +3. Decide whether the output is preview-only or meant to be consumed by the current project. +4. Decide the execution strategy: single asset vs repeated built-in calls vs CLI `generate-batch`. +5. Collect inputs up front: prompt(s), exact text (verbatim), constraints/avoid list, and any input images. +6. For every input image, label its role explicitly: + - reference image + - edit target + - supporting insert/style/compositing input +7. If the edit target is only on the local filesystem and you are staying on the built-in path, inspect it with `view_image` first so the image is available in conversation context. +8. If the user asked for a photo, illustration, sprite, product image, banner, or other explicitly raster-style asset, use `image_gen` rather than substituting SVG/HTML/CSS placeholders. If the request is for an icon, logo, or UI graphic that should match existing repo-native SVG/vector/code assets, prefer editing those directly instead. +9. Augment the prompt based on specificity: + - If the user's prompt is already specific and detailed, normalize it into a clear spec without adding creative requirements. + - If the user's prompt is generic, add tasteful augmentation only when it materially improves output quality. +10. Use the built-in `image_gen` tool by default. +11. If the user explicitly chooses the CLI fallback, then and only then use the fallback-only docs for quality, `input_fidelity`, masks, output format, output paths, and network setup. +12. Inspect outputs and validate: subject, style, composition, text accuracy, and invariants/avoid items. +13. Iterate with a single targeted change, then re-check. +14. For preview-only work, render the image inline; the underlying file may remain at the default `$CODEX_HOME/generated_images/...` path. +15. For project-bound work, move or copy the selected artifact into the workspace and update any consuming code or references. Never leave a project-referenced asset only at the default `$CODEX_HOME/generated_images/...` path. +16. For batches, persist only the selected finals in the workspace unless the user explicitly asked to keep discarded variants. +17. Always report the final saved path for any workspace-bound asset, plus the final prompt and whether the built-in tool or fallback CLI mode was used. ## Prompt augmentation -Reformat user prompts into a structured, production-oriented spec. Only make implicit details explicit; do not invent new requirements. + +Reformat user prompts into a structured, production-oriented spec. Make the user's goal clearer and more actionable, but do not blindly add detail. + +Treat this as prompt-shaping guidance, not a closed schema. Use only the lines that help, and add a short extra labeled line when it materially improves clarity. + +### Specificity policy + +Use the user's prompt specificity to decide how much augmentation is appropriate: + +- If the prompt is already specific and detailed, preserve that specificity and only normalize/structure it. +- If the prompt is generic, you may add tasteful augmentation when it will materially improve the result. + +Allowed augmentations: +- composition or framing hints +- polish level or intended-use hints +- practical layout guidance +- reasonable scene concreteness that supports the stated request + +Not allowed augmentations: +- extra characters or objects that are not implied by the request +- brand names, slogans, palettes, or narrative beats that are not implied +- arbitrary side-specific placement unless the surrounding layout supports it ## Use-case taxonomy (exact slugs) + Classify each request into one of these buckets and keep the slug consistent across prompts and references. Generate: - photorealistic-natural — candid/editorial lifestyle scenes with real texture and natural lighting. - product-mockup — product/packaging shots, catalog imagery, merch concepts. -- ui-mockup — app/web interface mockups that look shippable. +- ui-mockup — app/web interface mockups and wireframes; specify the desired fidelity. - infographic-diagram — diagrams/infographics with structured layout and text. - logo-brand — logo/mark exploration, vector-friendly. - illustration-story — comics, children’s book art, narrative scenes. @@ -85,90 +148,132 @@ Generate: Edit: - text-localization — translate/replace in-image text, preserve layout. - identity-preserve — try-on, person-in-scene; lock face/body/pose. -- precise-object-edit — remove/replace a specific element (incl. interior swaps). +- precise-object-edit — remove/replace a specific element (including interior swaps). - lighting-weather — time-of-day/season/atmosphere changes only. - background-extraction — transparent background / clean cutout. - style-transfer — apply reference style while changing subject/scene. - compositing — multi-image insert/merge with matched lighting/perspective. - sketch-to-render — drawing/line art to photoreal render. -Quick clarification (augmentation vs invention): -- If the user says “a hero image for a landing page”, you may add *layout/composition constraints* that are implied by that use (e.g., “generous negative space on the right for headline text”). -- Do not introduce new creative elements the user didn’t ask for (e.g., adding a mascot, changing the subject, inventing brand names/logos). +## Shared prompt schema -Template (include only relevant lines): -``` +Use the following labeled spec as shared prompt scaffolding for both top-level modes: + +```text Use case: Asset type: Primary request: -Scene/background: +Input images: (optional) +Scene/backdrop: Subject:
Style/medium: Composition/framing: Lighting/mood: Color palette: Materials/textures: -Quality: -Input fidelity (edits): Text (verbatim): "" Constraints: Avoid: ``` +Notes: +- `Asset type` and `Input images` are prompt scaffolding, not dedicated CLI flags. +- `Scene/backdrop` refers to the visual setting. It is not the same as the fallback CLI `background` parameter, which controls output transparency behavior. +- Fallback-only execution notes such as `Quality:`, `Input fidelity:`, masks, output format, and output paths belong in the explicit CLI path only. Do not treat them as built-in `image_gen` tool arguments. + Augmentation rules: -- Keep it short; add only details the user already implied or provided elsewhere. -- Always classify the request into a taxonomy slug above and tailor constraints/composition/quality to that bucket. Use the slug to find the matching example in `references/sample-prompts.md`. -- If the user gives a broad request (e.g., "Generate images for this website"), use judgment to propose tasteful, context-appropriate assets and map each to a taxonomy slug. -- For edits, explicitly list invariants ("change only X; keep Y unchanged"). +- Keep it short. +- Add only the details needed to improve the prompt materially. +- For edits, explicitly list invariants (`change only X; keep Y unchanged`). - If any critical detail is missing and blocks success, ask a question; otherwise proceed. ## Examples ### Generation example (hero image) -``` -Use case: stylized-concept +```text +Use case: product-mockup Asset type: landing page hero Primary request: a minimal hero image of a ceramic coffee mug Style/medium: clean product photography -Composition/framing: centered product, generous negative space on the right +Composition/framing: wide composition with usable negative space for page copy if needed Lighting/mood: soft studio lighting Constraints: no logos, no text, no watermark ``` ### Edit example (invariants) -``` +```text Use case: precise-object-edit Asset type: product photo background replacement -Primary request: replace the background with a warm sunset gradient +Primary request: replace only the background with a warm sunset gradient Constraints: change only the background; keep the product and its edges unchanged; no text; no watermark ``` -## Prompting best practices (short list) -- Structure prompt as scene -> subject -> details -> constraints. +## Prompting best practices +- Structure prompt as scene/backdrop -> subject -> details -> constraints. - Include intended use (ad, UI mock, infographic) to set the mode and polish level. - Use camera/composition language for photorealism. +- Only use SVG/vector stand-ins when the user explicitly asked for vector output or a non-image placeholder. - Quote exact text and specify typography + placement. - For tricky words, spell them letter-by-letter and require verbatim rendering. -- For multi-image inputs, reference images by index and describe how to combine them. +- For multi-image inputs, reference images by index and describe how they should be used. - For edits, repeat invariants every iteration to reduce drift. - Iterate with single-change follow-ups. -- For latency-sensitive runs, start with quality=low; use quality=high for text-heavy or detail-critical outputs. -- For strict edits (identity/layout lock), consider input_fidelity=high. -- If results feel “tacky”, add a brief “Avoid:” line (stock-photo vibe; cheesy lens flare; oversaturated neon; harsh bloom; oversharpening; clutter) and specify restraint (“editorial”, “premium”, “subtle”). +- If the prompt is generic, add only the extra detail that will materially help. +- If the prompt is already detailed, normalize it instead of expanding it. +- For explicit CLI fallback only, see `references/cli.md` and `references/image-api.md` for `quality`, `input_fidelity`, masks, output format, and output-path guidance. -More principles: `references/prompting.md`. Copy/paste specs: `references/sample-prompts.md`. +More principles shared by both modes: `references/prompting.md`. +Copy/paste specs shared by both modes: `references/sample-prompts.md`. ## Guidance by asset type Asset-type templates (website assets, game assets, wireframes, logo) are consolidated in `references/sample-prompts.md`. -## CLI + environment notes +## Fallback CLI mode only + +### Temp and output conventions +These conventions apply only to the explicit CLI fallback. They do not describe built-in `image_gen` output behavior. +- Use `tmp/imagegen/` for intermediate files (for example JSONL batches); delete them when done. +- Write final artifacts under `output/imagegen/`. +- Use `--out` or `--out-dir` to control output paths; keep filenames stable and descriptive. + +### Dependencies +Prefer `uv` for dependency management in this repo. + +Required Python package: +```bash +uv pip install openai +``` + +Optional for downscaling only: +```bash +uv pip install pillow +``` + +Portability note: +- If you are using the installed skill outside this repo, install dependencies into that environment with its package manager. +- In uv-managed environments, `uv pip install ...` remains the preferred path. + +### Environment +- `OPENAI_API_KEY` must be set for live API calls. +- Do not ask the user for `OPENAI_API_KEY` when using the built-in `image_gen` tool. +- Never ask the user to paste the full key in chat. Ask them to set it locally and confirm when ready. + +If the key is missing, give the user these steps: +1. Create an API key in the OpenAI platform UI: https://platform.openai.com/api-keys +2. Set `OPENAI_API_KEY` as an environment variable in their system. +3. Offer to guide them through setting the environment variable for their OS/shell if needed. + +If installation is not possible in this environment, tell the user which dependency is missing and how to install it into their active environment. + +### Script-mode notes - CLI commands + examples: `references/cli.md` - API parameter quick reference: `references/image-api.md` -- If network approvals / sandbox settings are getting in the way: `references/codex-network.md` +- Network approvals / sandbox settings for CLI mode: `references/codex-network.md` ## Reference map -- **`references/cli.md`**: how to *run* image generation/edits/batches via `scripts/image_gen.py` (commands, flags, recipes). -- **`references/image-api.md`**: what knobs exist at the API level (parameters, sizes, quality, background, edit-only fields). -- **`references/prompting.md`**: prompting principles (structure, constraints/invariants, iteration patterns). -- **`references/sample-prompts.md`**: copy/paste prompt recipes (generate + edit workflows; examples only). -- **`references/codex-network.md`**: environment/sandbox/network-approval troubleshooting. +- `references/prompting.md`: shared prompting principles for both modes. +- `references/sample-prompts.md`: shared copy/paste prompt recipes for both modes. +- `references/cli.md`: fallback-only CLI usage via `scripts/image_gen.py`. +- `references/image-api.md`: fallback-only API/CLI parameter reference. +- `references/codex-network.md`: fallback-only network/sandbox troubleshooting for CLI mode. +- `scripts/image_gen.py`: fallback-only CLI implementation. Do not load or use it unless the user explicitly chooses CLI mode. diff --git a/skills/.curated/imagegen/agents/openai.yaml b/skills/.curated/imagegen/agents/openai.yaml index 74750c2..c9cfddb 100644 --- a/skills/.curated/imagegen/agents/openai.yaml +++ b/skills/.curated/imagegen/agents/openai.yaml @@ -1,6 +1,6 @@ interface: display_name: "Image Gen" - short_description: "Generate and edit images using OpenAI" + short_description: "Generate or edit images for websites, games, and more" icon_small: "./assets/imagegen-small.svg" icon_large: "./assets/imagegen.png" - default_prompt: "Generate or edit images for this task and return the final prompt plus selected outputs." + default_prompt: "Generate or edit the visual assets for this task with the built-in `image_gen` tool by default. First confirm that the task actually calls for a raster image; if the project already has SVG/vector/code-native assets and the user wants to extend or match those, do not use this skill. If the task includes reference images, treat them as references unless the user clearly wants an existing image modified. For multi-asset requests, loop built-in calls rather than treating batch as a separate top-level mode. Only use the fallback CLI if the user explicitly asks for it, and keep CLI-only controls such as `generate-batch`, `quality`, `input_fidelity`, masks, and output paths on that fallback path." diff --git a/skills/.curated/imagegen/references/cli.md b/skills/.curated/imagegen/references/cli.md index f5ab9e2..8cb0663 100644 --- a/skills/.curated/imagegen/references/cli.md +++ b/skills/.curated/imagegen/references/cli.md @@ -1,11 +1,13 @@ # CLI reference (`scripts/image_gen.py`) -This file contains the “command catalog” for the bundled image generation CLI. Keep `SKILL.md` as overview-first; put verbose CLI details here. +This file is for the fallback CLI mode only. Read it only after the user explicitly asks to use `scripts/image_gen.py` instead of the built-in `image_gen` tool. + +`generate-batch` is a CLI subcommand in this fallback path. It is not a top-level mode of the skill. ## What this CLI does -- `generate`: generate new images from a prompt -- `edit`: edit an existing image (optionally with a mask) — inpainting / background replacement / “change only X” -- `generate-batch`: run many jobs from a JSONL file (one job per line) +- `generate`: generate a new image from a prompt +- `edit`: edit one or more existing images +- `generate-batch`: run many generation jobs from a JSONL file Real API calls require **network access** + `OPENAI_API_KEY`. `--dry-run` does not. @@ -17,116 +19,142 @@ export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}" export IMAGE_GEN="$CODEX_HOME/skills/imagegen/scripts/image_gen.py" ``` +Install dependencies into that environment with its package manager. In uv-managed environments, `uv pip install ...` remains the preferred path. + +## Quick start + Dry-run (no API call; no network required; does not require the `openai` package): -``` -python "$IMAGE_GEN" generate --prompt "Test" --dry-run -``` - -Generate (requires `OPENAI_API_KEY` + network): - -``` -uv run --with openai python "$IMAGE_GEN" generate --prompt "A cozy alpine cabin at dawn" --size 1024x1024 -``` - -No `uv` installed? Use your active Python env: - -``` -python "$IMAGE_GEN" generate --prompt "A cozy alpine cabin at dawn" --size 1024x1024 -``` - -## Guardrails (important) -- Use `python "$IMAGE_GEN" ...` (or equivalent full path) for generations/edits/batch work. -- Do **not** create one-off runners (e.g. `gen_images.py`) unless the user explicitly asks for a custom wrapper. -- **Never modify** `scripts/image_gen.py`. If something is missing, ask the user before doing anything else. - -## Defaults (unless overridden by flags) -- Model: `gpt-image-1.5` -- Size: `1024x1024` -- Quality: `auto` -- Output format: `png` -- Background: unspecified (API default). If you set `--background transparent`, also set `--output-format png` or `webp`. - -## Quality + input fidelity -- `--quality` works for `generate`, `edit`, and `generate-batch`: `low|medium|high|auto`. -- `--input-fidelity` is **edit-only**: `low|high` (use `high` for strict edits like identity or layout lock). - -Example: -``` -python "$IMAGE_GEN" edit --image input.png --prompt "Change only the background" --quality high --input-fidelity high -``` - -## Masks (edits) -- Use a **PNG** mask; an alpha channel is strongly recommended. -- The mask should match the input image dimensions. -- In the edit prompt, repeat invariants (e.g., “change only the background; keep the subject unchanged”) to reduce drift. - -## Optional deps -Prefer `uv run --with ...` for an out-of-the-box run without changing the current project env; otherwise install into your active env: - -``` -uv pip install openai -``` - -## Common recipes - -Generate + also write a downscaled copy for fast web loading: - -``` -uv run --with openai --with pillow python "$IMAGE_GEN" generate \ - --prompt "A cozy alpine cabin at dawn" \ - --size 1024x1024 \ - --downscale-max-dim 1024 +```bash +python "$IMAGE_GEN" generate \ + --prompt "Test" \ + --out output/imagegen/test.png \ + --dry-run ``` Notes: -- Downscaling writes an extra file next to the original (default suffix `-web`, e.g. `output-web.png`). -- Downscaling requires Pillow (use `uv run --with pillow ...` or install it into your env). +- One-off dry-runs print the API payload and the computed output path(s). +- Repo-local finals should live under `output/imagegen/`. + +Generate (requires `OPENAI_API_KEY` + network): + +```bash +python "$IMAGE_GEN" generate \ + --prompt "A cozy alpine cabin at dawn" \ + --size 1024x1024 \ + --out output/imagegen/alpine-cabin.png +``` + +Edit: + +```bash +python "$IMAGE_GEN" edit \ + --image input.png \ + --prompt "Replace only the background with a warm sunset" \ + --out output/imagegen/sunset-edit.png +``` + +## Guardrails +- Use the bundled CLI directly (`python "$IMAGE_GEN" ...`) after activating the correct environment. +- Do **not** create one-off runners (for example `gen_images.py`) unless the user explicitly asks for a custom wrapper. +- **Never modify** `scripts/image_gen.py`. If something is missing, ask the user before doing anything else. + +## Defaults +- Model: `gpt-image-1.5` +- Supported model family for this CLI: GPT Image models (`gpt-image-*`) +- Size: `1024x1024` +- Quality: `auto` +- Output format: `png` +- Default one-off output path: `output/imagegen/output.png` +- Background: unspecified unless `--background` is set + +## Quality, input fidelity, and masks (CLI fallback only) +These are explicit CLI controls. They are not built-in `image_gen` tool arguments. + +- `--quality` works for `generate`, `edit`, and `generate-batch`: `low|medium|high|auto` +- `--input-fidelity` is **edit-only** and validated as `low|high` +- `--mask` is **edit-only** + +Example: + +```bash +python "$IMAGE_GEN" edit \ + --image input.png \ + --prompt "Change only the background" \ + --quality high \ + --input-fidelity high \ + --out output/imagegen/background-edit.png +``` + +Mask notes: +- For multi-image edits, pass repeated `--image` flags. Their order is meaningful, so describe each image by index and role in the prompt. +- The CLI accepts a single `--mask`. +- Use a PNG mask when possible; the script treats mask handling as best-effort and does not perform full preflight validation beyond file checks/warnings. +- In the edit prompt, repeat invariants (`change only the background; keep the subject unchanged`) to reduce drift. + +## Output handling +- Use `tmp/imagegen/` for temporary JSONL inputs or scratch files. +- Use `output/imagegen/` for final outputs. +- Reruns fail if a target file already exists unless you pass `--force`. +- `--out-dir` changes one-off naming to `image_1.`, `image_2.`, and so on. +- Downscaled copies use the default suffix `-web` unless you override it. + +## Common recipes Generate with augmentation fields: -``` +```bash python "$IMAGE_GEN" generate \ --prompt "A minimal hero image of a ceramic coffee mug" \ - --use-case "landing page hero" \ + --use-case "product-mockup" \ --style "clean product photography" \ - --composition "centered product, generous negative space" \ - --constraints "no logos, no text" + --composition "wide product shot with usable negative space for page copy" \ + --constraints "no logos, no text" \ + --out output/imagegen/mug-hero.png +``` + +Generate + also write a downscaled copy for fast web loading: + +```bash +python "$IMAGE_GEN" generate \ + --prompt "A cozy alpine cabin at dawn" \ + --size 1024x1024 \ + --downscale-max-dim 1024 \ + --out output/imagegen/alpine-cabin.png ``` Generate multiple prompts concurrently (async batch): -``` -mkdir -p tmp/imagegen +```bash +mkdir -p tmp/imagegen output/imagegen/batch cat > tmp/imagegen/prompts.jsonl << 'EOF' -{"prompt":"Cavernous hangar interior with a compact shuttle parked center-left, open bay door","use_case":"game concept art environment","composition":"wide-angle, low-angle, cinematic framing","lighting":"volumetric light rays through drifting fog","constraints":"no logos or trademarks; no watermark","size":"1536x1024"} -{"prompt":"Gray wolf in profile in a snowy forest, crisp fur texture","use_case":"wildlife photography print","composition":"100mm, eye-level, shallow depth of field","constraints":"no logos or trademarks; no watermark","size":"1024x1024"} +{"prompt":"Cavernous hangar interior with a compact shuttle parked near the center","use_case":"stylized-concept","composition":"wide-angle, low-angle","lighting":"volumetric light rays through drifting fog","constraints":"no logos or trademarks; no watermark","size":"1536x1024"} +{"prompt":"Gray wolf in profile in a snowy forest","use_case":"photorealistic-natural","composition":"eye-level","constraints":"no logos or trademarks; no watermark","size":"1024x1024"} EOF -python "$IMAGE_GEN" generate-batch --input tmp/imagegen/prompts.jsonl --out-dir out --concurrency 5 +python "$IMAGE_GEN" generate-batch \ + --input tmp/imagegen/prompts.jsonl \ + --out-dir output/imagegen/batch \ + --concurrency 5 -# Cleanup (recommended) rm -f tmp/imagegen/prompts.jsonl ``` Notes: -- Use `--concurrency` to control parallelism (default `5`). Higher concurrency can hit rate limits; the CLI retries on transient errors. -- Per-job overrides are supported in JSONL (e.g., `size`, `quality`, `background`, `output_format`, `n`, and prompt-augmentation fields). +- `generate-batch` requires `--out-dir`. +- generate-batch requires --out-dir. +- Use `--concurrency` to control parallelism (default `5`). +- Per-job overrides are supported in JSONL (for example `size`, `quality`, `background`, `output_format`, `output_compression`, `moderation`, `n`, `model`, `out`, and prompt-augmentation fields). - `--n` generates multiple variants for a single prompt; `generate-batch` is for many different prompts. -- Treat the JSONL file as temporary: write it under `tmp/` and delete it after the run (don’t commit it). - -Edit: - -``` -python "$IMAGE_GEN" edit --image input.png --mask mask.png --prompt "Replace the background with a warm sunset" -``` +- In batch mode, per-job `out` is treated as a filename under `--out-dir`. ## CLI notes - Supported sizes: `1024x1024`, `1536x1024`, `1024x1536`, or `auto`. - Transparent backgrounds require `output_format` to be `png` or `webp`. -- Default output is `output.png`; multiple images become `output-1.png`, `output-2.png`, etc. -- Use `--no-augment` to skip prompt augmentation. +- `--prompt-file`, `--output-compression`, `--moderation`, `--max-attempts`, `--fail-fast`, `--force`, and `--no-augment` are supported. +- This CLI is intended for GPT Image models. Do not assume older non-GPT image-model behavior applies here. ## See also -- API parameter quick reference: `references/image-api.md` -- Prompt examples: `references/sample-prompts.md` +- API parameter quick reference for fallback CLI mode: `references/image-api.md` +- Prompt examples shared across both top-level modes: `references/sample-prompts.md` +- Network/sandbox notes for fallback CLI mode: `references/codex-network.md` diff --git a/skills/.curated/imagegen/references/codex-network.md b/skills/.curated/imagegen/references/codex-network.md index 9e8ec8d..249d628 100644 --- a/skills/.curated/imagegen/references/codex-network.md +++ b/skills/.curated/imagegen/references/codex-network.md @@ -1,28 +1,33 @@ # Codex network approvals / sandbox notes +This file is for the fallback CLI mode only. Read it only after the user explicitly asks to use `scripts/image_gen.py`. + This guidance is intentionally isolated from `SKILL.md` because it can vary by environment and may become stale. Prefer the defaults in your environment when in doubt. -## Why am I asked to approve every image generation call? -Image generation uses the OpenAI Image API, so the CLI needs outbound network access. In many Codex setups, network access is disabled by default (especially under stricter sandbox modes), and/or the approval policy may require confirmation before networked commands run. +## Why am I asked to approve image generation calls? +The fallback CLI uses the OpenAI Image API, so it needs outbound network access. In many Codex setups, network access is disabled by default and/or the approval policy requires confirmation before networked commands run. -## How do I reduce repeated approval prompts (network)? -If you trust the repo and want fewer prompts, enable network access for the relevant sandbox mode and relax the approval policy. +## Important note about approvals vs network +- `--ask-for-approval never` suppresses approval prompts. +- It does **not** by itself enable network access. +- In `workspace-write`, network access still depends on your Codex configuration (for example `[sandbox_workspace_write] network_access = true`). + +## How do I reduce repeated approval prompts? +If you trust the repo and want fewer prompts, use a configuration or profile that both: +- enables network for the sandbox mode you plan to use +- sets an approval policy that matches your risk tolerance Example `~/.codex/config.toml` pattern: -``` -approval_policy = "never" +```toml +approval_policy = "on-request" sandbox_mode = "workspace-write" [sandbox_workspace_write] network_access = true ``` -Or for a single session: - -``` -codex --sandbox workspace-write --ask-for-approval never -``` +If you want quieter automation after network is enabled, you can choose a stricter approval policy, but do that intentionally and with care. ## Safety note -Use caution: enabling network and disabling approvals reduces friction but increases risk if you run untrusted code or work in an untrusted repository. +Enabling network and reducing approvals lowers friction, but increases risk if you run untrusted code or work in an untrusted repository. diff --git a/skills/.curated/imagegen/references/image-api.md b/skills/.curated/imagegen/references/image-api.md index ba87782..a2750c1 100644 --- a/skills/.curated/imagegen/references/image-api.md +++ b/skills/.curated/imagegen/references/image-api.md @@ -1,36 +1,49 @@ # Image API quick reference +This file is for the fallback CLI mode only. Use it only after the user explicitly asks to use `scripts/image_gen.py` instead of the built-in `image_gen` tool. + +These parameters describe the Image API and bundled CLI fallback surface. Do not assume they are normal arguments on the built-in `image_gen` tool. + +## Scope +- This fallback CLI is intended for GPT Image models (`gpt-image-1.5`, `gpt-image-1`, and `gpt-image-1-mini`). +- The built-in `image_gen` tool and the fallback CLI do not expose the same controls. + ## Endpoints - Generate: `POST /v1/images/generations` (`client.images.generate(...)`) - Edit: `POST /v1/images/edits` (`client.images.edit(...)`) -## Models -- Default: `gpt-image-1.5` -- Alternatives: `gpt-image-1-mini` (for faster, lower-cost generation) - -## Core parameters (generate + edit) +## Core parameters for GPT Image models - `prompt`: text prompt - `model`: image model - `n`: number of images (1-10) - `size`: `1024x1024`, `1536x1024`, `1024x1536`, or `auto` - `quality`: `low`, `medium`, `high`, or `auto` -- `background`: `transparent`, `opaque`, or `auto` (transparent requires `png`/`webp`) +- `background`: output transparency behavior (`transparent`, `opaque`, or `auto`) for generated output; this is not the same thing as the prompt's visual scene/backdrop - `output_format`: `png` (default), `jpeg`, `webp` - `output_compression`: 0-100 (jpeg/webp only) - `moderation`: `auto` (default) or `low` ## Edit-specific parameters -- `image`: one or more input images (first image is primary) -- `mask`: optional mask image (same size, alpha channel required) -- `input_fidelity`: `low` (default) or `high` (support varies by model) - set it to `high` if the user needs a very specific edit and you can't achieve it with the default `low` fidelity. +- `image`: one or more input images. For GPT Image models, you can provide up to 16 images. +- `mask`: optional mask image +- `input_fidelity`: `low` (default) or `high` + +Model-specific note for `input_fidelity`: +- `gpt-image-1` and `gpt-image-1-mini` preserve all input images, but the first image gets richer textures and finer details. +- `gpt-image-1.5` preserves the first 5 input images with higher fidelity. ## Output - `data[]` list with `b64_json` per image +- The bundled `scripts/image_gen.py` CLI decodes `b64_json` and writes output files for you. -## Limits & notes +## Limits and notes - Input images and masks must be under 50MB. -- Use edits endpoint when the user requests changes to an existing image. +- Use the edits endpoint when the user requests changes to an existing image. - Masking is prompt-guided; exact shapes are not guaranteed. - Large sizes and high quality increase latency and cost. -- For fast iteration or latency-sensitive runs, start with `quality=low`; raise to `high` for text-heavy or detail-critical outputs. -- Use `input_fidelity=high` for strict edits (identity preservation, layout lock, or precise compositing). +- High `input_fidelity` can materially increase input token usage. +- If a request fails because a specific option is unsupported by the selected GPT Image model, retry manually without that option. + +## Important boundary +- `quality`, `input_fidelity`, explicit masks, `background`, `output_format`, and related parameters are fallback-only execution controls. +- Do not assume they are built-in `image_gen` tool arguments. diff --git a/skills/.curated/imagegen/references/prompting.md b/skills/.curated/imagegen/references/prompting.md index 3968fbf..8b6c684 100644 --- a/skills/.curated/imagegen/references/prompting.md +++ b/skills/.curated/imagegen/references/prompting.md @@ -1,81 +1,98 @@ -# Prompting best practices (gpt-image-1.5) +# Prompting best practices + +These prompting principles are shared by both top-level modes of the skill: +- built-in `image_gen` tool (default) +- explicit `scripts/image_gen.py` CLI fallback + +This file is about prompt structure, specificity, and iteration. Fallback-only execution controls such as `quality`, `input_fidelity`, masks, output format, and output paths live in the fallback docs. ## Contents - [Structure](#structure) -- [Specificity](#specificity) -- [Avoiding “tacky” outputs](#avoiding-tacky-outputs) -- [Composition & layout](#composition--layout) -- [Constraints & invariants](#constraints--invariants) +- [Specificity policy](#specificity-policy) +- [Allowed and disallowed augmentation](#allowed-and-disallowed-augmentation) +- [Composition and layout](#composition-and-layout) +- [Constraints and invariants](#constraints-and-invariants) - [Text in images](#text-in-images) -- [Multi-image inputs](#multi-image-inputs) +- [Input images and references](#input-images-and-references) - [Iterate deliberately](#iterate-deliberately) -- [Quality vs latency](#quality-vs-latency) +- [Fallback-only execution controls](#fallback-only-execution-controls) - [Use-case tips](#use-case-tips) - [Where to find copy/paste recipes](#where-to-find-copypaste-recipes) ## Structure -- Use a consistent order: scene/background -> subject -> key details -> constraints -> output intent. -- Include intended use (ad, UI mock, infographic) to set the mode and polish level. -- For complex requests, use short labeled lines instead of a long paragraph. +- Use a consistent order: scene/backdrop -> subject -> key details -> constraints -> output intent. +- Include intended use (ad, UI mock, infographic) to set the level of polish. +- For complex requests, use short labeled lines instead of one long paragraph. -## Specificity -- Name materials, textures, and visual medium (photo, watercolor, 3D render). -- For photorealism, include camera/composition language (lens, framing, lighting). -- Add targeted quality cues only when needed (film grain, textured brushstrokes, macro detail); avoid generic "8K" style prompts. +## Specificity policy +- If the user prompt is already specific and detailed, normalize it into a clean spec without adding creative requirements. +- If the prompt is generic, you may add tasteful detail when it materially improves the output. +- Treat examples in `sample-prompts.md` as fully-authored recipes, not as the default amount of augmentation to add to every request. -## Avoiding “tacky” outputs -- Don’t use vibe-only buzzwords (“epic”, “cinematic”, “trending”, “8k”, “award-winning”, “unreal engine”, “artstation”) unless the user explicitly wants that look. -- Specify restraint: “minimal”, “editorial”, “premium”, “subtle”, “natural color grading”, “soft contrast”, “no harsh bloom”, “no oversharpening”. -- For 3D/illustration, name the finish you want: “matte”, “paper grain”, “ink texture”, “flat color with soft shadow”; avoid “glossy plastic” unless requested. -- Add a short negative line when needed (especially for marketing art): “Avoid: stock-photo vibe; cheesy lens flare; oversaturated neon; excessive bokeh; fake-looking smiles; clutter”. +## Allowed and disallowed augmentation -## Composition & layout -- Specify framing and viewpoint (close-up, wide, top-down) and placement ("logo top-right"). -- Call out negative space if you need room for UI or overlays. +Allowed augmentation for generic prompts: +- composition and framing cues +- intended-use or polish-level hints +- practical layout guidance +- reasonable scene concreteness that supports the request -## Constraints & invariants -- State what must not change ("keep background unchanged"). -- For edits, say "change only X; keep Y unchanged" and repeat invariants on every iteration to reduce drift. +Do not add: +- extra characters, props, or objects that are not implied +- brand palettes, slogans, or story beats that are not implied +- arbitrary side-specific placement unless the surrounding layout supports it + +## Composition and layout +- Specify framing and viewpoint (close-up, wide, top-down) and placement only when it materially helps. +- Call out negative space if the asset clearly needs room for UI or copy. +- Avoid making left/right layout decisions unless the user or surrounding layout supports them. + +## Constraints and invariants +- State what must not change (`keep background unchanged`). +- For edits, say `change only X; keep Y unchanged` and repeat invariants on every iteration to reduce drift. ## Text in images - Put literal text in quotes or ALL CAPS and specify typography (font style, size, color, placement). - Spell uncommon words letter-by-letter if accuracy matters. - For in-image copy, require verbatim rendering and no extra characters. -## Multi-image inputs -- Reference inputs by index and role ("Image 1: product, Image 2: style"). -- Describe how to combine them ("apply Image 2's style to Image 1"). -- For compositing, specify what moves where and what must remain unchanged. +## Input images and references +- Do not assume that every provided image is an edit target. +- Label each image by index and role (`Image 1: edit target`, `Image 2: style reference`). +- If the user provides images for style, composition, or mood guidance and does not ask to modify them, treat the request as generation with references. +- If the user asks to preserve an existing image while changing specific parts, treat the request as an edit. +- For compositing, describe how the images interact (`place the subject from Image 2 into Image 1`). ## Iterate deliberately - Start with a clean base prompt, then make small single-change edits. - Re-specify critical constraints when you iterate. +- Prefer one targeted follow-up at a time over rewriting the whole prompt. -## Quality vs latency -- For latency-sensitive runs, start at `quality=low` and only raise it if needed. -- Use `quality=high` for text-heavy or detail-critical images. -- For strict edits (identity preservation, layout lock), consider `input_fidelity=high`. +## Fallback-only execution controls +- `quality`, `input_fidelity`, explicit masks, output format, and output paths are fallback-only execution controls. +- Do not assume they are built-in `image_gen` tool arguments. +- If the user explicitly chooses CLI fallback, see `references/cli.md` and `references/image-api.md` for those controls. ## Use-case tips Generate: -- photorealistic-natural: Prompt as if a real photo is captured in the moment; use photography language (lens, lighting, framing); call for real texture (pores, wrinkles, fabric wear, imperfections); avoid studio polish or staging; use `quality=high` when detail matters. +- photorealistic-natural: Prompt as if a real photo is captured in the moment; use photography language (lens, lighting, framing); call for real texture; avoid over-stylized polish unless requested. - product-mockup: Describe the product/packaging and materials; ensure clean silhouette and label clarity; if in-image text is needed, require verbatim rendering and specify typography. -- ui-mockup: Describe a real product; focus on layout, hierarchy, and common UI elements; avoid concept-art language so it looks shippable. -- infographic-diagram: Define the audience and layout flow; label parts explicitly; require verbatim text; use `quality=high`. -- logo-brand: Keep it simple and scalable; ask for a strong silhouette and balanced negative space; avoid gradients and fine detail. -- illustration-story: Define panels or scene beats; keep each action concrete; for continuity, restate character traits and outfit each time. -- stylized-concept: Specify style cues, material finish, and rendering approach (3D, painterly, clay); add a short "Avoid" line to prevent tacky effects. +- ui-mockup: Describe the target fidelity first (shippable mockup or low-fi wireframe), then focus on layout, hierarchy, and practical UI elements; avoid concept-art language. +- infographic-diagram: Define the audience and layout flow; label parts explicitly; require verbatim text. +- logo-brand: Keep it simple and scalable; ask for a strong silhouette and balanced negative space; avoid decorative flourishes unless requested. +- illustration-story: Define panels or scene beats; keep each action concrete. +- stylized-concept: Specify style cues, material finish, and rendering approach (3D, painterly, clay) without inventing new story elements. - historical-scene: State the location/date and required period accuracy; constrain clothing, props, and environment to match the era. Edit: - text-localization: Change only the text; preserve layout, typography, spacing, and hierarchy; no extra words or reflow unless needed. -- identity-preserve: Lock identity (face, body, pose, hair, expression); change only the specified elements; match lighting and shadows; use `input_fidelity=high` if likeness drifts. +- identity-preserve: Lock identity (face, body, pose, hair, expression); change only the specified elements; match lighting and shadows. - precise-object-edit: Specify exactly what to remove/replace; preserve surrounding texture and lighting; keep everything else unchanged. - lighting-weather: Change only environmental conditions (light, shadows, atmosphere, precipitation); keep geometry, framing, and subject identity. -- background-extraction: Request transparent background; crisp silhouette; no halos; preserve label text exactly; optionally add a subtle contact shadow. -- style-transfer: Specify style cues to preserve (palette, texture, brushwork) and what must change; add "no extra elements" to prevent drift. -- compositing: Reference inputs by index; specify what moves where; match lighting, perspective, and scale; keep background and framing unchanged. -- sketch-to-render: Preserve layout, proportions, and perspective; add plausible materials, lighting, and environment; "do not add new elements or text." +- background-extraction: Request a clean cutout; crisp silhouette; no halos; preserve label text exactly; no restyling. +- style-transfer: Specify style cues to preserve (palette, texture, brushwork) and what must change; add `no extra elements` to prevent drift. +- compositing: Reference inputs by index; specify what moves where; match lighting, perspective, and scale; keep the base framing unchanged. +- sketch-to-render: Preserve layout, proportions, and perspective; choose materials and lighting that support the supplied sketch without adding new elements. ## Where to find copy/paste recipes -For copy/paste prompt specs (examples only), see `references/sample-prompts.md`. This file focuses on principles, structure, and iteration patterns. +For copy/paste prompt specs (examples only), see `references/sample-prompts.md`. This file focuses on principles, specificity, and iteration patterns. diff --git a/skills/.curated/imagegen/references/sample-prompts.md b/skills/.curated/imagegen/references/sample-prompts.md index f299666..e4b2b01 100644 --- a/skills/.curated/imagegen/references/sample-prompts.md +++ b/skills/.curated/imagegen/references/sample-prompts.md @@ -1,8 +1,21 @@ # Sample prompts (copy/paste) -Use these as starting points (recipes only). Keep user-provided requirements; do not invent new creative elements. +These prompt recipes are shared across both top-level modes of the skill: +- built-in `image_gen` tool (default) +- explicit `scripts/image_gen.py` CLI fallback -For prompting principles (structure, invariants, iteration), see `references/prompting.md`. +Use these as starting points. They are intentionally complete prompt recipes, not the default amount of augmentation to add to every user request. + +When adapting a user's prompt: +- keep user-provided requirements +- only add detail according to the specificity policy in `SKILL.md` +- do not treat every example below as permission to invent extra story elements + +The labeled lines are prompt scaffolding, not a closed schema. `Asset type` and `Input images` are prompt-only scaffolding; the CLI does not expose them as dedicated flags. + +Execution details such as explicit CLI flags, `quality`, `input_fidelity`, masks, output formats, and local output paths depend on mode. Use the built-in tool by default; only apply CLI-specific controls after the user explicitly opts into fallback mode. + +For prompting principles (structure, specificity, invariants, iteration), see `references/prompting.md`. ## Generate @@ -10,39 +23,36 @@ For prompting principles (structure, invariants, iteration), see `references/pro ``` Use case: photorealistic-natural Primary request: candid photo of an elderly sailor on a small fishing boat adjusting a net -Scene/background: coastal water with soft haze -Subject: weathered skin with wrinkles and sun texture; a calm dog on deck nearby +Scene/backdrop: coastal water with soft haze +Subject: weathered skin with wrinkles and sun texture Style/medium: photorealistic candid photo -Composition/framing: medium close-up, eye-level, 50mm lens +Composition/framing: medium close-up, eye-level Lighting/mood: soft coastal daylight, shallow depth of field, subtle film grain Materials/textures: real skin texture, worn fabric, salt-worn wood Constraints: natural color balance; no heavy retouching; no glamorization; no watermark Avoid: studio polish; staged look -Quality: high ``` ### product-mockup ``` Use case: product-mockup Primary request: premium product photo of a matte black shampoo bottle with a minimal label -Scene/background: clean studio gradient from light gray to white +Scene/backdrop: clean studio gradient from light gray to white Subject: single bottle centered with subtle reflection Style/medium: premium product photography Composition/framing: centered, slight three-quarter angle, generous padding Lighting/mood: softbox lighting, clean highlights, controlled shadows Materials/textures: matte plastic, crisp label printing Constraints: no logos or trademarks; no watermark -Quality: high ``` ### ui-mockup ``` Use case: ui-mockup -Primary request: mobile app UI for a local farmers market with vendors and specials -Scene/background: clean white background with subtle natural accents -Subject: header, vendor list with small photos, "Today's specials" section, location and hours +Primary request: mobile app home screen for a local farmers market with vendors and daily specials +Asset type: mobile app screen Style/medium: realistic product UI, not concept art -Composition/framing: iPhone frame, balanced spacing and hierarchy +Composition/framing: clean vertical mobile layout with clear hierarchy Constraints: practical layout, clear typography, no logos or trademarks, no watermark ``` @@ -50,13 +60,12 @@ Constraints: practical layout, clear typography, no logos or trademarks, no wate ``` Use case: infographic-diagram Primary request: detailed infographic of an automatic coffee machine flow -Scene/background: clean, light neutral background +Scene/backdrop: clean, light neutral background Subject: bean hopper -> grinder -> brew group -> boiler -> water tank -> drip tray Style/medium: clean vector-like infographic with clear callouts and arrows Composition/framing: vertical poster layout, top-to-bottom flow Text (verbatim): "Bean Hopper", "Grinder", "Brew Group", "Boiler", "Water Tank", "Drip Tray" Constraints: clear labels, strong contrast, no logos or trademarks, no watermark -Quality: high ``` ### logo-brand @@ -64,7 +73,7 @@ Quality: high Use case: logo-brand Primary request: original logo for "Field & Flour", a local bakery Style/medium: vector logo mark; flat colors; minimal -Composition/framing: single centered logo on plain background with padding +Composition/framing: single centered logo on a plain background with generous padding Constraints: strong silhouette, balanced negative space; original design only; no gradients unless essential; no trademarks; no watermark ``` @@ -72,7 +81,7 @@ Constraints: strong silhouette, balanced negative space; original design only; n ``` Use case: illustration-story Primary request: 4-panel comic about a pet left alone at home -Scene/background: cozy living room across panels +Scene/backdrop: cozy living room across panels Subject: pet reacting to the owner leaving, then relaxing, then returning to a composed pose Style/medium: comic illustration with clear panels Composition/framing: 4 equal-sized vertical panels, readable actions per panel @@ -83,10 +92,10 @@ Constraints: no text; no logos or trademarks; no watermark ``` Use case: stylized-concept Primary request: cavernous hangar interior with tall support beams and drifting fog -Scene/background: industrial hangar interior, deep scale, light haze -Subject: compact shuttle, parked center-left, bay door open +Scene/backdrop: industrial hangar interior, deep scale, light haze +Subject: compact shuttle parked near the center Style/medium: cinematic concept art, industrial realism -Composition/framing: wide-angle, low-angle, cinematic framing +Composition/framing: wide-angle, low-angle Lighting/mood: volumetric light rays cutting through fog Constraints: no logos or trademarks; no watermark ``` @@ -95,8 +104,8 @@ Constraints: no logos or trademarks; no watermark ``` Use case: historical-scene Primary request: outdoor crowd scene in Bethel, New York on August 16, 1969 -Scene/background: open field, temporary stages, period-accurate tents and signage -Subject: crowd in period-accurate clothing, authentic staging and environment +Scene/backdrop: open field with period-appropriate staging +Subject: crowd in period-accurate clothing, authentic environment Style/medium: photorealistic photo Composition/framing: wide shot, eye-level Constraints: period-accurate details; no modern objects; no logos or trademarks; no watermark @@ -109,24 +118,24 @@ Constraints: period-accurate details; no modern objects; no logos or trademarks; Use case: Asset type: Primary request: -Scene/background: +Scene/backdrop: Subject:
Style/medium: -Composition/framing: +Composition/framing: Lighting/mood: Color palette: -Constraints: +Constraints: ``` ### Website assets example: minimal hero background ``` Use case: stylized-concept Asset type: landing page hero background -Primary request: minimal abstract background with a soft gradient and subtle texture (calm, modern) -Style/medium: matte illustration / soft-rendered abstract background (not glossy 3D) -Composition/framing: wide composition; large negative space on the right for headline +Primary request: minimal abstract background with a soft gradient and subtle texture +Style/medium: matte illustration / soft-rendered abstract background +Composition/framing: wide composition with usable negative space for page copy Lighting/mood: gentle studio glow -Color palette: cool neutrals with a restrained blue accent +Color palette: restrained neutral palette Constraints: no text; no logos; no watermark ``` @@ -134,11 +143,11 @@ Constraints: no text; no logos; no watermark ``` Use case: stylized-concept Asset type: feature section illustration -Primary request: simple abstract shapes suggesting connection and flow (tasteful, minimal) -Scene/background: subtle light-gray backdrop with faint texture +Primary request: simple abstract shapes suggesting connection and flow +Scene/backdrop: subtle light-gray backdrop with faint texture Style/medium: flat illustration; soft shadows; restrained contrast Composition/framing: centered cluster; open margins for UI -Color palette: muted teal and slate, low contrast accents +Color palette: muted neutral palette Constraints: no text; no logos; no watermark ``` @@ -147,9 +156,9 @@ Constraints: no text; no logos; no watermark Use case: photorealistic-natural Asset type: blog header image Primary request: overhead desk scene with notebook, pen, and coffee cup -Scene/background: warm wooden tabletop +Scene/backdrop: warm wooden tabletop Style/medium: photorealistic photo -Composition/framing: wide crop; subject placed left; right side left empty +Composition/framing: wide crop with clean room for page copy Lighting/mood: soft morning light Constraints: no text; no logos; no watermark ``` @@ -159,7 +168,7 @@ Constraints: no text; no logos; no watermark Use case: stylized-concept Asset type: Primary request: -Scene/background: (if applicable) +Scene/backdrop: (if applicable) Subject:
Style/medium: ; Composition/framing: ; ; @@ -172,11 +181,10 @@ Constraints: no logos or trademarks; no watermark Use case: stylized-concept Asset type: game environment concept art Primary request: cavernous hangar interior with tall support beams and drifting fog -Scene/background: industrial hangar interior, deep scale, light haze -Subject: compact shuttle, parked center-left, bay door open -Foreground: painted floor markings; cables; tool carts along edges +Scene/backdrop: industrial hangar interior, deep scale, light haze +Subject: compact shuttle parked near the center Style/medium: cinematic concept art, industrial realism -Composition/framing: wide-angle, low-angle, cinematic framing +Composition/framing: wide-angle, low-angle Lighting/mood: volumetric light rays cutting through fog Constraints: no logos or trademarks; no watermark ``` @@ -186,12 +194,9 @@ Constraints: no logos or trademarks; no watermark Use case: stylized-concept Asset type: game character concept Primary request: desert scout character with layered travel gear -Silhouette: long coat with hood, wide boots, satchel -Outfit/gear: dusty canvas, leather straps, brass buckles -Face/hair: windworn face, short cropped hair +Subject: long coat, satchel, practical travel clothing Style/medium: character render; stylized realism -Pose: neutral hero pose -Background: simple neutral backdrop +Composition/framing: neutral hero pose on a simple backdrop Constraints: no logos or trademarks; no watermark ``` @@ -202,9 +207,7 @@ Asset type: game UI icon Primary request: round shield icon with a subtle rune pattern Style/medium: painted game UI icon Composition/framing: centered icon; generous padding; clear silhouette -Background: transparent -Lighting/mood: subtle highlights; crisp edges -Constraints: no text; no logos or trademarks; no watermark +Constraints: no text; no background scene elements; no logos or trademarks; no watermark ``` ### Game assets example: tileable texture @@ -213,8 +216,7 @@ Use case: stylized-concept Asset type: tileable game texture Primary request: worn sandstone blocks Style/medium: seamless tileable texture; PBR-ish look -Scale: medium tiling -Lighting: neutral / flat lighting +Scene/backdrop: neutral lighting reference only Constraints: seamless edges; no obvious focal elements; no text; no logos or trademarks; no watermark ``` @@ -223,10 +225,9 @@ Constraints: seamless edges; no obvious focal elements; no text; no logos or tra Use case: ui-mockup Asset type: website wireframe Primary request: -Fidelity: low-fi grayscale wireframe; hand-drawn feel; simple boxes -Layout: -Annotations: -Resolution/orientation: +Style/medium: low-fi grayscale wireframe +Composition/framing: +Subject: Constraints: no color; no logos; no real photos; no watermark ``` @@ -235,11 +236,10 @@ Constraints: no color; no logos; no real photos; no watermark Use case: ui-mockup Asset type: website wireframe Primary request: SaaS homepage layout with clear hierarchy -Fidelity: low-fi grayscale wireframe; hand-drawn feel; simple boxes -Layout: top nav; hero with headline and CTA; three feature cards; testimonial strip; pricing preview; footer -Annotations: label each block ("Nav", "Hero", "CTA", "Feature", "Testimonial", "Pricing", "Footer") -Resolution/orientation: landscape (wide) for desktop -Constraints: no color; no logos; no real photos; no watermark +Style/medium: low-fi grayscale wireframe +Subject: top nav; hero with headline and CTA; three feature cards; testimonial strip; pricing preview; footer +Composition/framing: landscape desktop layout +Constraints: label major blocks; no color; no logos; no real photos; no watermark ``` ### Wireframe example: pricing page @@ -247,23 +247,21 @@ Constraints: no color; no logos; no real photos; no watermark Use case: ui-mockup Asset type: website wireframe Primary request: pricing page layout with comparison table -Fidelity: low-fi grayscale wireframe; sketchy lines; simple boxes -Layout: header; plan toggle; 3 pricing cards; comparison table; FAQ accordion; footer -Annotations: label key areas ("Toggle", "Plan Card", "Table", "FAQ") -Resolution/orientation: landscape for desktop or portrait for tablet -Constraints: no color; no logos; no real photos; no watermark +Style/medium: low-fi grayscale wireframe +Subject: header; plan toggle; 3 pricing cards; comparison table; FAQ accordion; footer +Composition/framing: desktop or tablet layout +Constraints: label key areas; no color; no logos; no real photos; no watermark ``` ### Wireframe example: mobile onboarding flow ``` Use case: ui-mockup -Asset type: website wireframe +Asset type: mobile onboarding wireframe Primary request: three-screen mobile onboarding flow -Fidelity: low-fi grayscale wireframe; hand-drawn feel; simple boxes -Layout: screen 1 (logo placeholder, headline, illustration placeholder, CTA); screen 2 (feature bullets); screen 3 (form fields + CTA) -Annotations: label each block and screen number -Resolution/orientation: portrait (tall) for mobile -Constraints: no color; no logos; no real photos; no watermark +Style/medium: low-fi grayscale wireframe +Subject: screen 1 headline and CTA; screen 2 feature bullets; screen 3 form fields and CTA +Composition/framing: portrait mobile layout +Constraints: label screens and blocks; no color; no logos; no real photos; no watermark ``` ### Logo template @@ -286,7 +284,7 @@ Primary request: geometric leaf symbol suggesting sustainability and growth Style/medium: vector logo mark; flat colors; minimal Composition/framing: centered mark; clear silhouette Color palette: deep green and off-white -Constraints: no text; no gradients; no mockups; no 3D; no watermark +Constraints: no text unless requested; no gradients; no mockups; no 3D; no watermark ``` ### Logo example: monogram mark @@ -308,7 +306,6 @@ Primary request: clean wordmark for a modern studio Style/medium: vector wordmark; flat colors; minimal Text (verbatim): "Studio North" Composition/framing: centered text; even letter spacing -Color palette: charcoal on white Constraints: no gradients; no mockups; no 3D; no watermark ``` @@ -318,24 +315,23 @@ Constraints: no gradients; no mockups; no 3D; no watermark ``` Use case: text-localization Input images: Image 1: original infographic -Primary request: translate all in-image text to Spanish +Primary request: replace "Bean Hopper", "Grinder", "Brew Group", "Boiler", "Water Tank", and "Drip Tray" with "Tolva", "Molino", "Grupo de infusión", "Caldera", "Depósito de agua", and "Bandeja de goteo" Constraints: change only the text; preserve layout, typography, spacing, and hierarchy; no extra words; do not alter logos or imagery ``` ### identity-preserve ``` Use case: identity-preserve -Input images: Image 1: person photo; Image 2..N: clothing items +Input images: Image 1: person photo; Image 2..N: clothing references Primary request: replace only the clothing with the provided garments -Constraints: preserve face, body shape, pose, hair, expression, and identity; match lighting and shadows; keep background unchanged; no accessories or text -Input fidelity (edits): high +Constraints: preserve face, body shape, pose, hair, expression, and identity; match lighting and shadows; keep the background unchanged; no accessories or text ``` ### precise-object-edit ``` Use case: precise-object-edit Input images: Image 1: room photo -Primary request: replace ONLY the white chairs with wooden chairs +Primary request: replace only the white chairs with wooden chairs Constraints: preserve camera angle, room lighting, floor shadows, and surrounding objects; keep all other aspects unchanged ``` @@ -345,24 +341,22 @@ Use case: lighting-weather Input images: Image 1: original photo Primary request: make it look like a winter evening with gentle snowfall Constraints: preserve subject identity, geometry, camera angle, and composition; change only lighting, atmosphere, and weather -Quality: high ``` ### background-extraction ``` Use case: background-extraction Input images: Image 1: product photo -Primary request: extract the product on a transparent background -Output: transparent background (RGBA PNG) -Constraints: crisp silhouette, no halos/fringing; preserve label text exactly; no restyling +Primary request: isolate the product on a clean transparent background +Constraints: crisp silhouette; no halos or fringing; preserve label text exactly; no restyling ``` ### style-transfer ``` Use case: style-transfer Input images: Image 1: style reference -Primary request: apply Image 1's visual style to a man riding a motorcycle on a white background -Constraints: preserve palette, texture, and brushwork; no extra elements; plain white background +Primary request: apply Image 1's visual style to a man riding a motorcycle on a plain white backdrop +Constraints: preserve palette, texture, and brushwork; no extra elements ``` ### compositing @@ -370,8 +364,7 @@ Constraints: preserve palette, texture, and brushwork; no extra elements; plain Use case: compositing Input images: Image 1: base scene; Image 2: subject to insert Primary request: place the subject from Image 2 next to the person in Image 1 -Constraints: match lighting, perspective, and scale; keep background and framing unchanged; no extra elements -Input fidelity (edits): high +Constraints: match lighting, perspective, and scale; keep the base framing unchanged; no extra elements ``` ### sketch-to-render @@ -380,5 +373,4 @@ Use case: sketch-to-render Input images: Image 1: drawing Primary request: turn the drawing into a photorealistic image Constraints: preserve layout, proportions, and perspective; choose realistic materials and lighting; do not add new elements or text -Quality: high ``` diff --git a/skills/.curated/imagegen/scripts/image_gen.py b/skills/.curated/imagegen/scripts/image_gen.py index a16f202..57cab20 100644 --- a/skills/.curated/imagegen/scripts/image_gen.py +++ b/skills/.curated/imagegen/scripts/image_gen.py @@ -1,5 +1,7 @@ #!/usr/bin/env python3 -"""Generate or edit images with the OpenAI Image API. +"""Fallback CLI for explicit image generation or editing with GPT Image models. + +Used only when the user explicitly opts into CLI fallback mode. Defaults to gpt-image-1.5 and a structured prompt augmentation workflow. """ @@ -25,10 +27,13 @@ DEFAULT_QUALITY = "auto" DEFAULT_OUTPUT_FORMAT = "png" DEFAULT_CONCURRENCY = 5 DEFAULT_DOWNSCALE_SUFFIX = "-web" +DEFAULT_OUTPUT_PATH = "output/imagegen/output.png" +GPT_IMAGE_MODEL_PREFIX = "gpt-image-" ALLOWED_SIZES = {"1024x1024", "1536x1024", "1024x1536", "auto"} ALLOWED_QUALITIES = {"low", "medium", "high", "auto"} ALLOWED_BACKGROUNDS = {"transparent", "opaque", "auto", None} +ALLOWED_INPUT_FIDELITIES = {"low", "high", None} MAX_IMAGE_BYTES = 50 * 1024 * 1024 MAX_BATCH_JOBS = 500 @@ -43,6 +48,17 @@ def _warn(message: str) -> None: print(f"Warning: {message}", file=sys.stderr) +def _dependency_hint(package: str, *, upgrade: bool = False) -> str: + command = f"uv pip install {'-U ' if upgrade else ''}{package}" + return ( + "Activate the repo-selected environment first, then install it with " + f"`{command}`. If this repo uses a local virtualenv, start with " + "`source .venv/bin/activate`; otherwise use this repo's configured shared fallback " + "environment. If your project declares dependencies, prefer that project's normal " + "`uv sync` flow." + ) + + def _ensure_api_key(dry_run: bool) -> None: if os.getenv("OPENAI_API_KEY"): print("OPENAI_API_KEY is set.", file=sys.stderr) @@ -105,12 +121,25 @@ def _validate_background(background: Optional[str]) -> None: _die("background must be one of transparent, opaque, or auto.") +def _validate_input_fidelity(input_fidelity: Optional[str]) -> None: + if input_fidelity not in ALLOWED_INPUT_FIDELITIES: + _die("input-fidelity must be one of low or high.") + + +def _validate_model(model: str) -> None: + if not model.startswith(GPT_IMAGE_MODEL_PREFIX): + _die( + "model must be a GPT Image model (for example gpt-image-1.5, gpt-image-1, or gpt-image-1-mini)." + ) + + def _validate_transparency(background: Optional[str], output_format: str) -> None: if background == "transparent" and output_format not in {"png", "webp"}: _die("transparent background requires output-format png or webp.") def _validate_generate_payload(payload: Dict[str, Any]) -> None: + _validate_model(str(payload.get("model", DEFAULT_MODEL))) n = int(payload.get("n", 1)) if n < 1 or n > 10: _die("n must be between 1 and 10") @@ -238,9 +267,7 @@ def _downscale_image_bytes(image_bytes: bytes, *, max_dim: int, output_format: s try: from PIL import Image except Exception: - _die( - "Downscaling requires Pillow. Install with `uv pip install pillow` (then re-run)." - ) + _die(f"Downscaling requires Pillow. {_dependency_hint('pillow')}") if max_dim < 1: _die("--downscale-max-dim must be >= 1") @@ -306,8 +333,8 @@ def _decode_write_and_downscale( def _create_client(): try: from openai import OpenAI - except ImportError as exc: - _die("openai SDK not installed. Install with `uv pip install openai`.") + except ImportError: + _die(f"openai SDK not installed in the active environment. {_dependency_hint('openai')}") return OpenAI() @@ -318,9 +345,12 @@ def _create_async_client(): try: import openai as _openai # noqa: F401 except ImportError: - _die("openai SDK not installed. Install with `uv pip install openai`.") + _die( + f"openai SDK not installed in the active environment. {_dependency_hint('openai')}" + ) _die( - "AsyncOpenAI not available in this openai SDK version. Upgrade with `uv pip install -U openai`." + "AsyncOpenAI not available in this openai SDK version. " + f"{_dependency_hint('openai', upgrade=True)}" ) return AsyncOpenAI() @@ -506,8 +536,7 @@ async def _run_generate_batch(args: argparse.Namespace) -> int: _validate_generate_payload(job_payload) effective_output_format = _normalize_output_format(job_payload.get("output_format")) _validate_transparency(job_payload.get("background"), effective_output_format) - if "output_format" in job_payload: - job_payload["output_format"] = effective_output_format + job_payload["output_format"] = effective_output_format n = int(job_payload.get("n", 1)) outputs = _job_output_paths( @@ -557,8 +586,7 @@ async def _run_generate_batch(args: argparse.Namespace) -> int: _validate_generate_payload(payload) effective_output_format = _normalize_output_format(payload.get("output_format")) _validate_transparency(payload.get("background"), effective_output_format) - if "output_format" in payload: - payload["output_format"] = effective_output_format + payload["output_format"] = effective_output_format outputs = _job_output_paths( out_dir=out_dir, output_format=effective_output_format, @@ -634,12 +662,21 @@ def _generate(args: argparse.Namespace) -> None: output_format = _normalize_output_format(args.output_format) _validate_transparency(args.background, output_format) - if "output_format" in payload: - payload["output_format"] = output_format + payload["output_format"] = output_format output_paths = _build_output_paths(args.out, output_format, args.n, args.out_dir) + downscaled = None + if args.downscale_max_dim is not None: + downscaled = [str(_derive_downscale_path(p, args.downscale_suffix)) for p in output_paths] if args.dry_run: - _print_request({"endpoint": "/v1/images/generations", **payload}) + _print_request( + { + "endpoint": "/v1/images/generations", + "outputs": [str(p) for p in output_paths], + "outputs_downscaled": downscaled, + **payload, + } + ) return print( @@ -693,16 +730,26 @@ def _edit(args: argparse.Namespace) -> None: output_format = _normalize_output_format(args.output_format) _validate_transparency(args.background, output_format) - if "output_format" in payload: - payload["output_format"] = output_format + payload["output_format"] = output_format + _validate_input_fidelity(args.input_fidelity) output_paths = _build_output_paths(args.out, output_format, args.n, args.out_dir) + downscaled = None + if args.downscale_max_dim is not None: + downscaled = [str(_derive_downscale_path(p, args.downscale_suffix)) for p in output_paths] if args.dry_run: payload_preview = dict(payload) payload_preview["image"] = [str(p) for p in image_paths] if mask_path: payload_preview["mask"] = str(mask_path) - _print_request({"endpoint": "/v1/images/edits", **payload_preview}) + _print_request( + { + "endpoint": "/v1/images/edits", + "outputs": [str(p) for p in output_paths], + "outputs_downscaled": downscaled, + **payload_preview, + } + ) return print( @@ -797,7 +844,7 @@ def _add_shared_args(parser: argparse.ArgumentParser) -> None: parser.add_argument("--output-format") parser.add_argument("--output-compression", type=int) parser.add_argument("--moderation") - parser.add_argument("--out", default="output.png") + parser.add_argument("--out", default=DEFAULT_OUTPUT_PATH) parser.add_argument("--out-dir") parser.add_argument("--force", action="store_true") parser.add_argument("--dry-run", action="store_true") @@ -824,7 +871,9 @@ def _add_shared_args(parser: argparse.ArgumentParser) -> None: def main() -> int: - parser = argparse.ArgumentParser(description="Generate or edit images via the Image API") + parser = argparse.ArgumentParser( + description="Fallback CLI for explicit image generation or editing via GPT Image models" + ) subparsers = parser.add_subparsers(dest="command", required=True) gen_parser = subparsers.add_parser("generate", help="Create a new image") @@ -866,6 +915,7 @@ def main() -> int: _validate_size(args.size) _validate_quality(args.quality) _validate_background(args.background) + _validate_model(args.model) _ensure_api_key(args.dry_run) args.func(args)