Commit Graph
809 Commits
Author SHA1 Message Date
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> 2c78dfcca3 feat(semantic-spacetime): add format-compliant M1 skill core
Add the semantic-spacetime skill skeleton: a thin SKILL.md router with
triggers/anti-triggers, Load By Need, Quick Start, Related Skills, gotchas,
and exit conditions; a human-facing README; MIT license; deep
provenance-marked theory references (foundations, glossary, bibliography);
the versioned sst-model-v1 template pair; a 6-case schema-v1 eval manifest;
the root README catalog entry; regenerated marketplace and llms artifacts;
and the skill-triggers row.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-12 22:33:47 -04:00
Magnus HedemarkandGitHub c438315ae4 Merge pull request #313 from magnus919/docs/issue-311-skill-vetting
docs(agent-skills): add third-party skill vetting guidance and deterministic-script rule
2026-08-12 00:50:59 -04:00
Magnus Hedemark cc18d8602c Merge remote-tracking branch 'origin/main' into docs/issue-311-skill-vetting 2026-08-12 00:48:28 -04:00
Magnus Hedemark 8e5bc023a0 docs(agent-skills): add third-party skill vetting guidance and deterministic-script rule
Adds references/vetting-third-party-skills.md with a dependency-style
vetting checklist (provenance, SKILL.md body, scripts, references),
safe first-run practice, and reporting guidance, citing the Snyk
ToxicSkills audit as the primary source for ecosystem risk statistics.

SKILL.md gains the match-prescriptiveness-to-fragility decision rule,
the run-vs-reference intent rule for bundled scripts, and an
Adopting Third-Party Skills section. README and eval manifest updated.

Closes #311
2026-08-12 00:44:58 -04:00
Magnus HedemarkandGitHub 0413a05034 Merge pull request #312 from magnus919/feat/promise-theory-skill
feat(skill): promise theory — expert methodology for hybrid human+AI agent coordination
2026-08-12 00:20:31 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> 31a2fd16bb feat(skill): promise-theory quick start documents --json and --dry-run
Quick Start now documents the --json and --dry-run flags with their
exact semantics (single JSON object on stdout; read-only no-write
guard) and points to `python3 scripts/promise-contract.py --help` for
the full flag list, so a no-prior-knowledge user can drive the CLI
end-to-end (VAL-USE-013). No other content changes.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-12 00:12:47 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> ffd05b2e1d feat(skill): promise-theory catalog integration
Add the promise-theory skill to the repository catalog: root README entry,
regenerated llms.txt and marketplace/plugin packaging, and a
references/skill-triggers.md row.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-11 23:41:52 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> a30274fa68 feat(skill): promise-theory evals manifest + trigger probes
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-11 23:39:26 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> a2f27c036e feat(skill): promise-theory script + unit tests
Add promise-contract.py, a stdlib-only Python 3.10+ CLI that lints
promise-manifest v1 contracts (restricted-YAML or JSON) against the pinned
schema and renders a promise-graph summary. lint exits 0 on valid + full
coverage, 1 on lint errors/coverage gaps (accumulated, no fail-fast), and 2
on usage/IO errors; --json preserves the {valid, errors, warnings, coverage,
bindings} shape even on parse errors; --dry-run is a no-op guard. Robustness
handles empty/whitespace files, non-UTF-8 bytes, CRLF/BOM, JSON type errors,
and deep nesting without Python tracebacks.

Add tests/test_promise_contract.py covering valid contracts (YAML + JSON),
coverage gaps, schema violations, malformed input, --json, --dry-run, render,
and the robustness cases (empty, dup ids, bindings, enums, expires, encoding,
usage errors). 36 tests pass via unittest discover.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-11 23:32:21 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> 4a780c4757 feat(skill): promise-theory templates + human README
Add the three fillable templates (promise manifest YAML, agent contract,
promise review) and the human-facing README with the five required
sections. The manifest template is a lint-clean, fully covered example of
the pinned v1 schema with per-field comments; all intra-template id
references (accepts, expectations.about) resolve cross-agent. The contract
template carries the seven mandated sections with schema-aligned severity
and type vocabulary; the review template carries the five retrospective
sections with the three diagnosis categories.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-11 23:22:38 -04:00
Magnus Hedemark e856ddb52f feat(skill): promise-theory references — diagnosis & debugging + glossary
Complete the seven-reference set for promise-theory. diagnosis-and-debugging.md
maps the Cemri et al. multi-agent failure taxonomy onto promise-theory breach
categories (specification issues ↔ broken promise bodies; inter-agent conflicts
↔ failed acceptance/incompatible co-languages; task verification problems ↔
missing assessment), adds withdrawal failure as a fourth promise-theoretic
class, documents a four-step diagnostic procedure (walk the promise graph →
check bindings → check evaluation loop → check withdrawal semantics) with a
worked example, and states the theory's limitations and open problems (no
coordination-quality benchmark, guarantees don't compose across handoffs, LLM
promises lack causal teeth, stochasticity, ambiguity. glossary.md defines all
27 architecture §4.2 terms as heading-/bold-led entries with citations plus a
related-terms section. Both files stay under 60k chars, resolve all backtick
*.md references and markdown links, and carry consistent EXTRAPOLATION /
[UNVERIFIED] provenance markers.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
EOF
)
2026-08-11 23:19:09 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> 93e0d965ca feat(skill): promise-theory references — coordination, patterns, trust & verification
Adds the three coordination references per architecture 4.2/6/7:
- agent-coordination.md: the core thesis with the fixed 11-concept mapping
  table (concept -> concrete agent-coordination practice), the hybrid
  human+agent boundary (humans as acceptors/evaluators, calibrated
  subordination, causal vs moral responsibility, HITL escalation,
  three-languages problem, swarms vs teams), the multi-agent lineage, and
  the agent-council routing statement; cites Burgess arXiv:2604.10505 and
  states the scarcity of direct literature.
- patterns.md: all seven canonical patterns with worked examples, the M12
  ladder, the Ye & Tan contract tuple and lifecycle with degradation
  semantics, the named ESCALATE-2 trigger, and the workflow-architect
  routing via the bundles path.
- trust-and-verification.md: the two-component trust model, belief/evidence,
  P_succ, verification rates as attention budgets incl. Dunbar budgets,
  gameable assessment, semantic-promise measurement guidance, "confine,
  don't convince", the versioned promise ledger, and routing to
  agent-evals-and-observability and artifact-pyramids.
All files under the 60k-char reference cap with [UNVERIFIED]/EXTRAPOLATION
provenance markers per architecture 4.2.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-11 23:12:37 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> d01d7a1578 feat(skill): promise-theory references foundations + applications-infrastructure
Adds references/foundations.md (academic core with Burgess/Bergstra
citations: promise definition and notation S ─b→ R, proposals, scope,
impositions, obligation-as-derived, polarity/bindings, assessment/belief/
evidence, trust as discounting, exact/empty promises, deception, matrices/
graphs, valence/bundles, roles, discovery, Downstream Principle, evaluation
loops, history, adjacent frameworks, critiques; honest formal-status section)
and references/applications-infrastructure.md (CFEngine case study incl. the
promise-keeping-was-never-stored-as-data lesson, IaC comparison table,
distributed-systems connection, adoption history, LLM-reasoning-layer
argument). Both files < 60k chars with provenance markers per architecture
§4.2.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-11 23:06:23 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> 236d678a9d feat(skill): promise-theory scaffold — SKILL.md router + MIT LICENSE
Add the promise-theory skill's thin router (SKILL.md) and LICENSE per
architecture §4.1: frontmatter (name, trigger-oriented description with a
negative boundary, license MIT), a 8-line core-model summary with all six
mandated elements, all six use triggers, all five anti-triggers, a
Load-By-Need routing table covering the seven planned references, a Quick
Start (draft from template, lint with scripts/promise-contract.py),
cross-references to the six sibling skills, and all five gotchas. Grounded
in the mission research reports; sibling links resolve; references/*.md
links land with later reference features.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-11 22:56:28 -04:00
Magnus HedemarkandGitHub 4b9d347662 Merge pull request #310 from magnus919/feat/travel-guide-section-footers
feat(skill): travel-guide section-end footers — field notes, next-up, ghost mark
2026-08-10 20:09:05 -04:00
Magnus Hedemark 3cd705f7e6 feat(skill): travel-guide section-end footers — field notes, next-up, ghost mark
Fills the white space between sections with a bottom-of-page footer per
section: a content-derived field note (first anchor failure mode, first day
alternative, practical recheck item, or first skip reason) when one exists, a
next-section line with the following section's number, and a faint ghost
route mark. Footers hug the page bottom via flex column + margin-top auto;
multi-page sections carry the footer at the end of the section. Sheets fill
the print page so the footer lands at the bottom instead of floating.

Field notes repeat model content in one line and never invent new plans;
sections with nothing worth saying render the next-up line only. QA gate,
editorial reference, SKILL.md, and README updated; test suite extended to
cover footer presence, next-section wiring, and field-note content.

AI assistance: implementation and tests drafted by Jasper (Hermes Agent),
design reviewed and approved by Magnus Hedemark.
2026-08-10 20:01:17 -04:00
Magnus HedemarkandGitHub b87d18e15f feat(skills): expose audio/video recording URLs in fireflies transcripts get (#308) (#309) 2026-08-10 17:37:13 -04:00
Magnus HedemarkandGitHub 15b517906a Merge pull request #307 from magnus919/docs/supabase-evals-attribution
docs(skill): attribute supabase/evals harness reference (Apache-2.0)
2026-08-09 18:17:22 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> 2114147c21 docs(skill): attribute supabase/evals harness reference (Apache-2.0)
The reference's concepts, runtime descriptions, and commands are derived
from the supabase/evals README, which is Apache-2.0. Add an attribution
section to references/agent-evals.md with the license link and list the
harness repository in references/source-index.md, per the repository's
attribution convention.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-09 18:14:15 -04:00
Magnus HedemarkandGitHub 7c61659d1b Merge branch 'main' into dependabot/pip/loguru-gte-0.7.3 2026-08-09 18:12:19 -04:00
Magnus HedemarkandGitHub 7b5845d52d Merge pull request #306 from magnus919/feat/supabase-evals-harness
feat(skill): incorporate supabase/evals harness into supabase skill
2026-08-09 18:09:16 -04:00
Magnus Hedemarkandfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> 4a5f18e435 feat(skill): incorporate supabase/evals harness into supabase skill
Add references/agent-evals.md documenting the official supabase/evals
harness: eval/experiment concepts, the tools and local-stack runtimes,
run and result-viewing commands, and a mapping of harness scenarios to
the skill's operating references. Route to it from the supabase
"Choose the path" table and from postgres, agent-evals-and-observability,
backend-engineering, and data-engineering. Add two eval cases covering
the new reference and keep the generated catalog artifacts current.

Closes #271

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-09 18:05:33 -04:00
Magnus HedemarkandGitHub 6cdc1b1c32 Merge pull request #305 from magnus919/fix/meshcore-packet-capture-v2.2.0
docs(meshcore-packet-capture): refresh skill for upstream v2.2.0
2026-08-09 13:55:21 -04:00
Magnus Hedemark 9f68c2aef3 test(meshcore-packet-capture): add eval manifest for v2.2.0 surface
The repo's eval-coverage ratchet fails modified skills without a schema-valid
eval manifest (coverage 62% >= 50% fail-on-modify threshold). Add five cases
covering the v2.2.0 additions: --neighbors-now/--neighbors-exit CLI, per-broker
neighbors opt-in + IATA requirement, payload-decoding scope limits, the
--user-service install/uninstall flow, and config precedence.

Part of #304
2026-08-09 13:51:21 -04:00
Magnus HedemarkandMagnus Hedemark b4e7c582a6 docs(meshcore-packet-capture): refresh skill for upstream v2.2.0
Bring SKILL.md and references up to date with agessaman/meshcore-packet-capture
v2.2.0 (26 commits past the v2.0.0 source index):

- CLI boundary: document --neighbors-now / --neighbors-exit
- Config: payload decoding (decode_payloads, include_decoded, hashtag
  channels, channel keys), neighbors publishing (interval, discover window,
  scope timeouts, max), log rotation, ble_pin, per-broker owner/email
- MQTT: neighbors and decoded topics, per-broker include_decoded/neighbors
- Deployment: --user-service install/uninstall flow, meshcore ==2.3.8 pin
- Source index: refresh commit/version, cover payload_decode.py and
  neighbors.py

Closes #304

Co-authored-by: Magnus Hedemark <magnus@users.noreply.github.com>
2026-08-09 13:39:39 -04:00
Magnus HedemarkandGitHub 2823d82585 Merge pull request #303 from magnus919/feat/travel-guide-visual-system
feat(skill): travel-guide visual system — journey line, day strip, meters, photo grade
2026-08-08 15:27:42 -04:00
Magnus Hedemark 3f7c46e368 feat(skill): travel-guide visual system — journey line, day strip, meters, photo grade
Adds the visual deltas verified against the mock: a route journey line on the
cover for multi-stop trips, a color-coded trip-at-a-glance day strip after the
brief, pace and budget meters, ghost section numbers, a unified warm photo
grade on anchor images, and a lede drop cap. Day cards gain an optional kind
field (arrive, city, excursion, coast) that drives the strip colors; anchor
cards and the validator accept optional image fields. Photo-sourcing guidance
for free-license images added to the research reference, QA gate updated, and
the test suite extended to cover the new renderer and validator behavior.

AI assistance: implementation and tests drafted by Jasper (Hermes Agent),
design reviewed and approved by Magnus Hedemark.
2026-08-08 15:20:11 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
6fe5aed86e feat(skill): rename writing skill to writers-helper (#302)
Rename the skill directory to writers-helper and update the name field,
eval manifest skill_name, skill README title and example paths, root
README catalog entry, and the skill-triggers index. Regenerate llms.txt,
.claude-plugin/marketplace.json, .codex-plugin/plugin.json, and
.agents/plugins/marketplace.json from their generators. Content is
unchanged.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-08 13:06:46 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
ecdfc658c1 feat(skill): add writing skill for the full writing lifecycle (#301)
Add writing/, a comprehensive writer's personal skill distilled from a
44-book writing-craft and publishing library. Ships 10 expert references
(planning and research, craft and structure, prose and style, drafting,
editing and revision, blocks and prompts, habits and lifestyle,
publishing and career, genres and formats, pitfalls and solutions),
13 fill-in templates (premise canvas through book proposal, query
letter, and submission log), and 5 Python helper scripts (position-aware
prompt generator, session planner, manuscript stats analyzer, habit
journal, submission tracker). All content is original paraphrase and
synthesis; no copyrighted source material is reproduced.

Add the catalog entry at its sorted position and regenerate llms.txt,
.claude-plugin/marketplace.json, .codex-plugin/plugin.json, and
.agents/plugins/marketplace.json from their generators.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-08 12:53:55 -04:00
Magnus HedemarkandGitHub 29a367e90b fix: SkillOpt optimization of travel-guide skill (3 epochs)
SkillOpt optimization of the travel-guide skill (3 autonomous greenfield epochs, 9/9 proposals accepted): page breaks ahead of all section headers, filesystem hygiene with dedicated working folders, sanitizer defaults redacting profile preferences/constraints, audience and private-by-default in the trip contract, process scaling for narrow questions, renderer-hang guidance, private-artifact image sourcing, shareable edition-phrasing review, and a narrow-question eval case.

Authored by Jasper (AI agent on behalf of Magnus Hedemark).
2026-08-08 03:01:24 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
9a42be585e feat(skill): add genius-life creativity practice skill (#299)
* feat(skill): add genius-life core skill files

Add genius-life/SKILL.md and the six references/ files (creative-process,
practice-mode, development-mode, practices-catalog, evidence-basis,
scope-and-safety) as original synthesis from the mission research library,
with copyright-compliant paraphrase, named-fellow attribution, and honest
framing. Templates, README, evals, and catalog integration ship in later
features.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): add genius-life templates, README, and evals

Add six fillable worksheets (talent audit, session plan, project
worksheet, conditions audit, incubation log, risk and failure review),
a human-facing README with the repository's required sections, and a
12-case eval manifest covering both modes and the required boundary
behaviors.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): add genius-life to catalog and regenerate artifacts

Add the genius-life README catalog entry at its sorted position and
regenerate llms.txt, .claude-plugin/marketplace.json, .codex-plugin/plugin.json,
and .agents/plugins/marketplace.json from their generators.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-08 01:31:20 -04:00
Magnus HedemarkandGitHub c44685800e feat: add personalized travel-guide skill
Add the personalized travel dossier skill, its portable evaluation cases, rendering and privacy helpers, tests, and synchronized discovery catalogs.

Authored by Jasper (AI agent on behalf of Magnus Hedemark).
2026-08-07 23:34:19 -04:00
Magnus HedemarkandGitHub 43fd4b9518 fix(anydoc): clarify rendered-layout inspection route (#297) 2026-08-07 16:13:42 -04:00
Magnus HedemarkandGitHub 935299d271 Optimize anydoc skill guidance (#296) 2026-08-06 23:02:50 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
f37dc73829 feat(skill): add anydoc — office documents to GitHub-Flavored Markdown (#295)
* feat(skill): add anydoc core content and references

Add the anydoc skill content tree: SKILL.md (progressive-disclosure index
with frontmatter per ALLOWED_FIELDS), human-facing README, the five reference
files (formats, cli-reference, errors, workflows, sources), 24 committed
fixtures (valid + error cases), and a fixture-grounded eval manifest with 8
cases. Every documented behavior, exit code, and error message was verified
against the real pinned CLI (npx -y @firecrawl/anydoc@0.1.6); verbatim --help
and error transcripts are reproduced character-for-character.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): add anydoc wrapper script and unit tests

Implements scripts/anydoc, a stdlib-only Python wrapper around the pinned
@firecrawl/anydoc@0.1.6 CLI: convert/batch/info subcommands, global
--json/--dry-run, input and output pre-validation, friendly hints for the
no-OCR/encrypted/malformed/unsupported error classes, Node >= 20 and npx
availability checks, deterministic batch output naming with documented
duplicate/collision behavior, and exit codes 0/1/2. Adds offline unittest
suite (46 tests, real-CLI tests skip when npx is unavailable) and keeps the
wrapper contract documented in cli-reference.md and errors.md.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): ratchet anydoc evals to 14 grounded cases

Verify the pre-authored 8-case manifest and extend it with six
high-signal cases (PDF lower-fidelity pipeline, legacy .ppt table
flattening, ODP same-serializer, RTF, EPUB, CSV header promotion),
each grounded in real pinned-CLI runs against the committed fixtures.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* feat(skill): integrate anydoc into repo catalog and artifacts

Add the sorted anydoc catalog entry to README.md (between agent-skills
and api-design-and-evolution), regenerate the tracked catalog artifacts
(.claude-plugin/marketplace.json, .codex-plugin/plugin.json,
.agents/plugins/marketplace.json, llms.txt) with the ruby generators,
and add a routing note to documents/SKILL.md pointing office-document
to-markdown conversion at the anydoc skill.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* fix(skill): polish anydoc wrapper timeout, JSON shape, and docs

- run_cli raises CliTimeoutError on the 120s timeout; convert/batch with
  --json now emit one parseable JSON error envelope (error_class "timeout")
  on stdout before exiting, so --json always yields exactly one JSON doc
- batch JSON failure entries (pre-validation and CLI) now carry error_class
  ("io" for missing/dir inputs, mapped classes for CLI failures), so all
  batch failure entries share the same shape
- build_cli_command places -o/-f before the -- separator for dash-leading
  filenames, so `convert -f csv -- -weird` converts instead of misparsing
  ("unexpected second input"); absolute-path inputs unchanged
- workflows.md vault-ingestion recipe globs notes/* instead of docs/* and
  warns to run from a temp/vault dir, never touching repo-root docs/
- unit tests: +6 (timeout envelope x4, batch error_class shape,
  dash-leading filename); suite grows 46 -> 52

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-06 20:11:19 -04:00
Magnus HedemarkandGitHub cff17c5974 fix: SkillOpt 3-epoch optimization of forward-deployed-engineering bundle (#294)
* fix: SkillOpt Epoch 1 — forward-deployed-engineering optimization

Inline the nine-stage contract table into SKILL.md (required question, minimum
output, stop condition per stage) with template links and the entry-evidence
rule; dedup the stage table out of references/lifecycle-and-artifacts.md into
a pointer. Name agent-evals-and-observability and production-readiness inline
in the applied-AI release gate (loading protocol step 5).

Validated: 2/2 held-out edits accepted (non-regression, all-pass baseline),
repo validators green (validate-skills, validate-skill-quality,
validate-bundles, validate-evals).

* fix: SkillOpt Epoch 2 — forward-deployed-engineering optimization

Add a 'Where to enter the lifecycle' table (existing state -> entry stage,
with the neckbeard route for bounded changes) and the entry-evidence rule for
mid-stream joins. Replace the flat 'When not to use' list with a proactive
Scenario | Reach for | Why routing table covering the six boundary routes.

Validated: 2/2 held-out edits accepted (non-regression, all-pass baseline)
plus a regression probe on epistemic labels; repo validators green.

* fix: SkillOpt Epoch 3 — forward-deployed-engineering optimization

Add references/worked-example-engagement.md, a fully synthetic depth
calibration artifact showing the charter, evidence-labeled ledger, stage
handoff, evaluation and release decision, adoption scorecard, outcome
measurement record, and productization record for one engagement. Add a File
map row, enumerate the templates row (surfacing engagement-status), and add a
depth-calibration pointer in the Lifecycle section.

Validated: 2/2 held-out edits accepted (non-regression, all-pass baseline);
repo validators green; sanitization scan clean (no private identifiers).
2026-08-06 03:20:45 -04:00
Magnus HedemarkandGitHub a4db8e7d43 Add forward-deployed-engineering bundle (#292)
* feat: add forward deployed engineering bundle

* fix: close FDE bundle review findings
2026-08-06 00:37:47 -04:00
Magnus HedemarkandGitHub eadb82e069 fix(skills): reconcile fireflies CLI with live GraphQL schema (#290)
Fixes #289

- transcripts list: drop removed TranscriptsQueryScope type (scope is a
  String in the live schema), require [String!] for organizers and
  participants, add title/organizer-email/participant-email filters
- bites create: use the live transcript_Id argument name and the
  BitePrivacy enum (public, team, participants)
- add ergonomic commands for documented gaps found in the audit:
  askfred get, meetings update-channel, meetings share --expiry-days,
  live add-to (addToLiveMeeting), live soundbite (createLiveSoundbite),
  audio create-upload/confirm-upload (two-phase upload), users set-role
- add eval manifest (5 cases) to satisfy the modified-skill eval ratchet
- update SKILL.md, cli-reference, api-reference, source-index, workflows
  to match the audited surface and record the 2026-08-05 schema audit
2026-08-05 22:29:36 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
b8a5092c26 feat(linear): project mutations and richer issue verbs in CLI (#288)
## What this adds

Implements the request in #287 and the Tier 1 audit gaps for the `linear` skill's `scripts/linear` CLI, reconciled against the live Linear GraphQL schema.

### New verbs
- `linear project update` — name, description, status, start/target dates, priority, with the same `--dry-run`/`--confirm` gate as issue mutations, and a local 255-character description guard matching Linear's `projectUpdate` limit (Linear rejects longer descriptions with a generic error).
- `linear issue archive` / `linear issue unarchive` — both gated, returning `IssueArchivePayload.entity`.
- `linear state list --team ENG` — first-class workflow-state discovery (previously states were only visible in the `issue move` failure path).

### Richer issue verbs
- `issue create` now accepts `--project`, `--parent`, `--assignee`, `--label` (repeatable), `--state`, `--due`.
- `issue update` now accepts `--assignee`, `--label` (add), `--remove-label`, `--due`, `--project`.

### Resolution rules (all require exactly one match, mirroring `resolve_team`)
- Project: UUID or exact name
- Parent: issue identifier or UUID
- Assignee: exact name, display name, or email (via `users`)
- Label: exact name within the issue's team (via `team.labels`)
- Workflow state: exact name within the issue's team (existing `team.states` resolver, now reusable for `--state` on create)
- Project status: exact name or type (via `projectStatuses`)

### Docs, tests, evals
- SKILL.md command map, state-change gate, and error/recovery sections; README; `domain-and-workflows.md` (project semantics + 255-char limit), `graphql-contract.md` (resolution queries), `integration-boundaries.md` (intentional exclusions list), `sources.md` (2026-08-05 schema re-verification note).
- 15 new offline tests (45 total) covering resolution, gates, dry-run intent, payload shapes, and field guards.
- Added a sixth eval case (`safe-project-and-issue-mutations`).

## Validation
- `python3 -m unittest linear/tests/test_linear.py` — 45/45 pass
- `python3 scripts/validate-evals.py`, `ruby scripts/validate-skills.rb`, `python3 scripts/check-artifacts.py`, `python3 scripts/eval-coverage.py --modified-from origin/main`, skill-quality validator, marketplace/codex/llms freshness, jscpd — all green locally

Closes #287

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-05 16:43:36 -04:00
dependabot[bot]andGitHub f03c48de94 chore(deps-dev): update ruff requirement from >=0.16.0 to >=0.16.1
Updates the requirements on [ruff](https://github.com/astral-sh/ruff) to permit the latest version.
- [Release notes](https://github.com/astral-sh/ruff/releases)
- [Changelog](https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md)
- [Commits](https://github.com/astral-sh/ruff/compare/0.16.0...0.16.1)

---
updated-dependencies:
- dependency-name: ruff
  dependency-version: 0.16.1
  dependency-type: direct:development
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-08-05 19:25:28 +00:00
dependabot[bot]andGitHub 5f789ccf3e chore(deps-dev): update loguru requirement from >=0.7 to >=0.7.3
Updates the requirements on [loguru](https://github.com/Delgan/loguru) to permit the latest version.
- [Release notes](https://github.com/Delgan/loguru/releases)
- [Changelog](https://github.com/Delgan/loguru/blob/master/CHANGELOG.rst)
- [Commits](https://github.com/Delgan/loguru/compare/0.7.0...0.7.3)

---
updated-dependencies:
- dependency-name: loguru
  dependency-version: 0.7.3
  dependency-type: direct:development
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-08-05 19:25:05 +00:00
Magnus HedemarkandGitHub 5d101007ef fix(skill): SkillOpt 3-epoch optimization — dsm5 navigation, decisions, and answer patterns (#281)
Greenfield post-publication SkillOpt run (epoch 0 = 3eb7bd4) on dsm5.

Epoch 1 (prominence/navigation, 4 edits):
- Promote scripts/lookup.py into workflow step 3 for unknown-condition routing
- Anchor split-chapter index-first pattern at the routing table
- Cross-reference Completion criteria from workflow step 9
- Add read-scope guard to step 4 (index + specific part, no whole chapters)

Epoch 2 (decision intelligence, 3 edits):
- Answer-first rule: provisional answer with marked unknowns, then high-yield follow-ups
- Third-party rule: asker may not be subject; never diagnose third parties
- Code-verification guard: cite codes/specifiers/prevalence only from read files

Epoch 3 (pattern expansion, 2 edits):
- Add Answer shape section: six-part structure calibrating output depth
- Add existing-diagnosis handling to Audience adaptation (explain, don't re-derive)

Validation: 21 held-out runs (baseline+candidate), zero regressions; final
installed-artifact regression passed. Description unchanged; catalogs unaffected.
2026-08-05 13:00:09 -04:00
Magnus HedemarkandGitHub e99c6d2994 fix(raleigh): skip token-gated imagery folders in discovery and canary (#280) 2026-08-05 11:16:33 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
3eb7bd4096 feat(validation): enforce 60K-char cap on skill reference files (#279)
* feat(validation): enforce 60K-char cap on skill reference files

Implements issue #277:

- validate-references.rb: new oversized_reference_errors check — every
  references/*.md must be <= 60,000 characters; error reports path, size,
  and the split-and-reindex remediation; wired into validate-skills.rb
- test-validate-skills.rb: 5 fixture tests (under-limit passes, over-limit
  fails with path+size, exactly-at-limit passes, remediation message,
  non-.md ignored); the suite now runs in validate.yml after the format
  check (it was previously untested in CI)
- Docs: agent-skills/SKILL.md, agent-skills/references/best-practices.md,
  and the AGENTS.md Format Compliance table document the cap and the
  split-and-reindex procedure
- Compliance: split remote-systems-administration/references/ansible.md
  and programming-principles/references/refactoring-guru.full.md into an
  index + focused parts (content moved verbatim); SKILL.md routing,
  README, and source-index references updated; pre-existing stale
  refactoring-guru-smells.md reference repointed to the index
- Fix pre-existing quality-gate violations in the programming-principles
  and remote-systems-administration descriptions (imperative verb +
  negative boundary) so this PR's CI quality step passes; regenerated
  llms.txt and marketplace artifacts

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* test(evals): add eval manifests to modified skills for ratchet

The eval-coverage ratchet requires schema-valid eval manifests for any
skill modified once coverage is past 50%. This PR modifies
programming-principles and remote-systems-administration (splitting
their oversized references), so add evals/evals.json to both:

- programming-principles: 6 output-quality cases (task-to-book mapping,
  principled code review, refactor-vs-rewrite, no-op detection, rule
  distillation, principle conflicts)
- remote-systems-administration: 6 output-quality cases (discovery
  before change, smallest control plane, rollback planning, platform
  identification, verification evidence, escalation on missing
  authority)

Coverage: 87/145 (60.0%) schema-valid; ratchet clean.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-04 22:39:14 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
c312b36166 fix(skill): split oversized dsm5 references into index + parts (#278)
Agent harnesses truncate file reads around ~60k characters, so the
largest dsm5 reference files (up to 132k chars) were being cut off
mid-file (reported: "The neurodevelopmental file was truncated").

- Split 15 reference files over 50k chars into a small index (original
  filename preserved, so all existing links keep resolving) plus part
  files of <= ~40k chars each, organized by disorder group
- Updated SKILL.md routing rows to point at indexes and read the part
  for the condition; added large-file handling guidance
- Updated dsm5/README.md What You Get table; documented the size
  convention in 00-overview-and-method.md (Maintaining this library)
- Verified: no reference file exceeds 50k chars (66 files), all 466
  relative links resolve, validators pass, lookup.py lists all parts

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-04 22:10:15 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
d346970bf8 feat(skill): add dsm5 — evidence-based companion to the DSM-5-TR (#276)
Adds the dsm5 skill: an evidence-based conversational expert grounded in
the DSM-5-TR (American Psychiatric Association, 2022) for clinicians,
practitioners, patients, and family members.

- SKILL.md: safety-first conversation workflow (triage -> clarify ->
  route -> compare criteria -> differentials -> calibrated conclusion),
  reference routing table, crisis protocol, audience adaptation
- references/: 28 files — foundation (00-02), all 22 DSM-5-TR diagnostic
  classes (10-31), Part III measures/culture/AMPD/conditions-for-further-
  study (32-33), and cross-cutting differentials (40). Criteria are
  paraphrased with exact counts, durations, specifiers, and ICD-10-CM
  codes, plus per-disorder clinician and patient/family conversation
  guides
- scripts/lookup.py: stdlib keyword search across the reference library
  (--json/--list/--max/-q)
- evals/evals.json: 9 output-quality cases (schema v1)
- README.md: human-facing overview, install notes, and APA attribution
- Catalog entries and generated artifacts (llms.txt, marketplace
  plugins) regenerated; all repo validators pass

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-04 21:30:11 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
6f67a34ef1 feat(skill): cross-pollinate the new tool wave into catalog routing (#275)
* feat(skill): cross-pollinate the new tool wave into catalog routing

Wire the recent tool skill wave into the two-layer routing graph so the
new tool skills are reachable from the methodology skills that own their
domains, and vice versa:

- methodology -> tool down-routes: platform-engineering -> kubernetes,
  terraform, telemetry, postgres, grafana; site-reliability-engineering ->
  telemetry, grafana; data-engineering and backend-engineering -> postgres;
  frontend-engineering -> mobile-development; verification-methodology ->
  playwright, documents; technical-documentation -> documents
- neckbeard: add mobile-development and documents routing rows plus
  change-surface coverage entries, and cross-link the lightweight
  test-hardening path to qa-methodology's bounded mutation-review material
- collaboration layer: chief-of-staff-methodology -> slack/notion/email,
  go-to-market -> crm, conditional-customer-success -> crm; fix the dead
  seo-content-optimization reference in go-to-market (now seo-audit)
- references/skill-triggers.md: add trigger rows for the 14 new skills
- bring go-to-market's description up to the quality validator's
  imperative-verb + negative-boundary requirement and regenerate catalogs

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

* fix(skill): add eval manifest for technical-documentation

The eval-coverage ratchet fails on modified skills without a schema-valid
manifest once coverage passes 50%. technical-documentation was modified by
the routing cross-pollination change and lacked one; add six output-quality
cases covering README authorship, API reference generation, CLI help design,
agent-facing docs, documentation-site IA, and troubleshooting sections.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>

---------

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-04 12:46:06 -04:00
Magnus HedemarkandGitHub f9db3dbe4b fix(calculator): honest burn-multiple and runway labels, surface model assumptions (#274)
* fix(calculator): honest burn-multiple and runway labels, surface model assumptions

- Burn Multiple now reports Graham's metric (net burn / net new ARR);
  the net burn / MRR ratio is reported separately as Burn to Revenue.
  The qualifier (efficient/healthy/warning/critical) is derived from the
  real burn multiple, so DEAD verdicts no longer print 'efficient'.
- ALIVE verdicts no longer print a misleading 'Runway: 120 months'
  (projection cap); output now shows 'Projected cash-out' with 'none
  within the 10-year projection' when the company never runs out.
- Model assumptions (fixed/variable burn split, variable burn ratio,
  growth decay, projection cap, safety buffer) are now surfaced in the
  human report and in JSON model_assumptions.
- SKILL.md: fix dead paulgraham.com/default.html source URL to aord.html;
  update output-field docs and examples to real model output.
- Add regression tests (tests/integration/test_default_alive.py).

Fixes #272
Fixes #273

* docs(calculator): add When Not to Use boundary (validator requirement)
2026-08-04 11:30:05 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
0ff373467c fix: correct vllm models-check test name and stripe cancel boolean (#270)
- vllm: rename test_empty_models_is_a_failure to
  test_models_check_parses_from_stub and fix its misleading docstring;
  it asserts positive-path parsing of the stub's served model list, not an
  empty-models failure.
- stripe: pass cancel_at_period_end as the boolean True instead of the
  string 'true', and normalize booleans to lowercase true/false during
  form encoding so the wire payload stays Stripe-compatible.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-03 20:37:46 -04:00
Magnus HedemarkGitHubfactory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
3256a87bcb feat(skill): add collaboration & business-app tool layer (Slack, Notion, email, CRM, payments) (#269)
Adds five top-level operational tool skills, one per named tool:

- slack: messages, channels, threads, search, files, and webhook signature
  verification (HMAC-SHA256) via a bounded, stdlib-only slack-cli.
- notion: pages, database queries, search, and guarded page updates via
  notion-cli.
- email: transactional email via Twilio SendGrid (send, deliverability
  bounces/spam reports, Signed Event Webhook verification with a
  self-contained ECDSA P-256 verifier) via email-cli.
- crm: HubSpot CRM records, contact search, and deal pipeline views with
  guarded stage updates via crm-cli.
- stripe: read-only-first balance, payment, and subscription queries with
  a guarded period-end subscription cancellation via stripe-cli.

Each skill ships an executable script (--json output, --limit bounded reads,
--dry-run/--yes mutation gate), a human README with the five required
sections, a schema-v1 evals/evals.json with six output-quality cases, a dated
source index + operations reference, and a deterministic unittest suite run
by check-artifacts. All five are indexed in the top-level README and the
generated catalogs were regenerated. Eval coverage rises from 78/139 to
83/144.

Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-03 20:19:26 -04:00