Add a new skill connecting demand, performance, reliability, and spend
decisions. Covers capacity models, unit economics, budget/quota controls,
load/soak test evidence, and SLO-cost tradeoffs with structured templates.
Includes:
- SKILL.md with connected-dimensions framework, working method, four
labeled scenarios (growth, peak, degraded, cost-constrained), and
routing table to six adjacent skills
- README.md with five required human-facing sections
- references/discovery-brief.md comparing ownership boundaries across
financial-modeling, platform-engineering, SRE, product-analytics,
production-readiness, product-roadmapping, and resilience-and-recovery
- Five fillable templates: capacity-model, unit-economics-record,
budget-quota-decision, load-soak-test-plan, slo-cost-tradeoff-record
- evals/evals.json with five output-quality cases: growth-forecast,
peak-event, slo-cost-conflict, quota-decision, misleading-unit-cost
- Regenerated marketplace, Codex, and llms.txt catalogs (117 skills)
- Updated root README catalog section and skill-triggers index
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 18:41:03 -04:00
Magnus HedemarkGitHubusername <username>factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(production-readiness): add production-readiness skill
Add a cross-domain production-readiness skill that assembles production
evidence into a risk-scaled launch decision. Includes:
- SKILL.md: three risk classes (Low/Standard/High) with proportional
evidence requirements, 11-category evidence checklist with named source
or explicit gap for every category, four launch-decision outcomes
(go/no-go/defer/exception) with accountable owners, exception routing
to explicit human approval, and a route-to table for 12 specialist skills.
- README.md: human-facing with all five required sections.
- references/discovery-brief.md: bounded survey of existing production
and engineering skills with concrete ownership boundaries against
release-engineering and site-reliability-engineering.
- references/readiness-record.md: fillable readiness record template.
- evals/evals.json: five output-quality cases covering low-risk docs,
user-facing launch, migration-dependent release, missing owner evidence
(blocked), and exception requiring human approval.
Closes#196
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* chore(production-readiness): update catalog files for production-readiness
Update root README catalog, skill-triggers index, and three generated
marketplace catalog files to include the new production-readiness skill.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
---------
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 18:14:22 -04:00
Magnus HedemarkGitHubusername <username>factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(product-lifecycle-learning): add product lifecycle learning skill (#194)
Introduce a new skill to close the launch-to-learning loop for product features
and capabilities. Covers:
- Post-launch outcome review with explicit epistemic categories
(expected/observed/uncertain/inferred)
- Assumption ledger updates with confidence shifts
- Multi-dimensional feature health assessment
- Six lifecycle decisions: continue/improve/harvest/pivot/pause/retire
- Full retirement lifecycle: deprecation communication, migration paths,
customer treatment during sunset, and internal cleanup
- Durable retained learning records that feed back into roadmap, analytics,
adoption, experimentation, and specifications
Ships 4 references (discovery brief, epistemic discipline, retirement lifecycle,
feedback destinations), 6 templates (outcome review, assumption ledger update,
feature health record, retirement decision, sunset plan, retained learning
record), and 7 eval cases including adversarial coverage.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* fix(product-lifecycle-learning): regenerate marketplace with corrected description
The Claude marketplace JSON contained the original description starting with
"Close" which was replaced with "Compare" to satisfy the imperative-verb
quality check. Regenerate to match the corrected SKILL.md frontmatter.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* fix(product-lifecycle-learning): regenerate llms.txt with corrected description
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
---------
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 17:33:10 -04:00
Magnus HedemarkGitHubusername <username>factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Add a conditional skill for products with recurring human relationships.
Covers success plans, health evidence, renewal/expansion signals, QBRs,
handoffs, escalation, and closed-loop Voice of Customer. The skill is
conditional: it declines and routes away when the product has no accounts,
renewals, QBRs, or customer-success team.
Includes:
- SKILL.md with conditional frontmatter, trigger sections, four product-
model adaptations (B2B subscription, transactional, public-service,
internal product), core artifacts, privacy and human-judgment boundaries,
and routing to product-analytics-and-measurement, product-adoption,
product-experimentation, go-to-market, and product-lifecycle-learning.
- README.md with all five required human-facing sections.
- references/discovery-brief.md surveying existing content and defining
ownership boundaries and routing.
- references/privacy-and-human-judgment.md with consent framework,
surveillance-risk guidance, decision-support rules, and data
classification tiers.
- templates/applicability-decision.md, templates/success-plan.md,
templates/health-risk-record.md, and templates/escalation-and-
feedback-closure.md.
- evals/evals.json with 5 output-quality cases covering B2B subscription,
internal-tool decline (negative trigger), public-service routing,
renewal-risk with mixed signals, and conflicting health evidence.
- Updated root README.md catalog section, references/skill-triggers.md,
and regenerated marketplace/Codex/llms.txt catalogs.
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 17:30:51 -04:00
Magnus HedemarkGitHubusername <username>factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Add a reusable implementation-planning skill for turning approved
requirements or specifications into executable, dependency-aware delivery
plans. Covers work breakdown into vertical slices, dependency mapping with
critical-path analysis, ownership assignment, sequencing and parallelism,
staged rollout strategy with rollback paths, and verification traceability
against the original requirement.
Includes:
- SKILL.md with valid frontmatter, entry gate for prerequisite approval,
progressive-disclosure file map, and handoff table to specialist skills
- README.md with all five required human-facing sections
- references/discovery-brief.md comparing existing planning material and
defining ownership boundaries
- templates/ for implementation plan, dependency record, and risk/decision/
verification sections
- evals/evals.json with six output-quality cases covering ambiguous
requirements, cross-repository dependencies, data migration, risky
rollout, unapproved prerequisite rejection, and multi-team ownership
conflict
- Catalog and routing updates (README, skill-triggers, generated catalogs)
Closes#186
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 17:29:00 -04:00
Magnus HedemarkGitHubusername <username>factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(product-analytics-and-measurement): add product analytics and measurement skill
Add a new skill that turns intended product outcomes into observable,
governed evidence: metric trees with leading/lagging indicators and
countermetrics, event/tracking plans with identity/session/data-quality/
ownership considerations, instrumentation QA across client/server/pipeline/
end-to-end layers, dashboard contracts, privacy-aware measurement, and
decision cadence for outcome reviews.
Includes:
- SKILL.md with Loading Guide, When to Use/Not to Use, related-skill routing
- README.md with all 5 required human-facing sections
- references/discovery-brief.md mapping ownership boundaries vs existing skills
- references/metric-tree.md with measurability gates and contextual examples
- templates/tracking-plan.md (event taxonomy, identity resolution, privacy)
- templates/instrumentation-qa-checklist.md (4-layer QA)
- templates/outcome-review.md (decision cadence template)
- evals/evals.json with 6 cases covering new feature, internal product,
public service, conflicting metrics, unmeasurable North Star rejection,
and privacy-boundary measurement
- Updated root README catalog, skill-triggers.md, and regenerated catalogs
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(product-analytics-and-measurement): add product analytics and measurement skill
Add a new skill that turns intended product outcomes into observable,
governed evidence: metric trees with leading/lagging indicators and
countermetrics, event/tracking plans with identity/session/data-quality/
ownership considerations, instrumentation QA across client/server/pipeline/
end-to-end layers, dashboard contracts, privacy-aware measurement, and
decision cadence for outcome reviews.
Includes:
- SKILL.md with Loading Guide, When to Use/Not to Use, related-skill routing
- README.md with all 5 required human-facing sections
- references/discovery-brief.md mapping ownership boundaries vs existing skills
- references/metric-tree.md with measurability gates and contextual examples
- templates/tracking-plan.md (event taxonomy, identity resolution, privacy)
- templates/instrumentation-qa-checklist.md (4-layer QA)
- templates/outcome-review.md (decision cadence template)
- evals/evals.json with 6 cases covering new feature, internal product,
public service, conflicting metrics, unmeasurable North Star rejection,
and privacy-boundary measurement
- Updated root README catalog, skill-triggers.md, and regenerated catalogs
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
---------
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
2026-08-02 16:58:08 -04:00
Magnus HedemarkGitHubusername <username>factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* fix(product-strategy): correct stale RICE reference
Fix the misspelled prose reference at product-strategy/references/product-strategy.md:69
to point at the canonical rice-framework.md owned by product-methodology. Add a
repository check that scans references/*.md for stale prose backtick references
to nonexistent files (the bug class the SKILL.md link-resolution pass cannot
see), wired into validate-skills.rb, with a regression test suite proving the
stale reference is caught when reintroduced. product-strategy and
product-methodology remain grandfathered; no evals manifests are added.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* fix(validate-references): case-sensitive resolution for stale-reference scan
The references scan resolved backtick tokens case-insensitively on hosts with
case-insensitive filesystems (default macOS APFS), so a token such as
`EVIDENCE-LEDGER.md` matched an existing lowercase `evidence-ledger.md` and
escaped detection locally while failing CI's Linux runners. Resolve candidates
against exact directory entries so results match CI on every host, and treat
neckbeard delivery-packet field names (EVIDENCE-LEDGER, DELIVERY-SPEC, REVIEW,
V2-SPEC) as doc-type names rather than file references. Adds regression tests
for case-mismatched and exact-case references.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
---------
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add delivery packet reference
Author references/delivery-packet.md: the durable cross-phase handoff for
change-request runs. Defines the nine field groups (a-i), the artifact
ownership map (writer/reviewer/gate/path per artifact), field-group write
ownership per phase, lifecycle states with allowed transitions and terminal
semantics, blocked-state semantics, resumability rules with a changed-head
procedure (material and non-material branches plus SHA-update recording) and a
concrete resume-after-context-boundary example, exact-head binding for every
verdict, baseline-vs-post-change evidence with boundary labels, skip
transparency, and a portability statement. Link the packet from the SKILL.md
file map for progressive-disclosure discovery.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add delivery packet template and extend contract templates
Create templates/delivery-packet.md mirroring the nine field groups from
references/delivery-packet.md with fillable sections (placeholder + fill
instruction or example per section; resumability, gates, and lifecycle
sections demonstrate SHA, verdict, and state entry). Update
templates/change-contract.md and templates/evidence-ledger.md to add
change-request provenance, skip-reason, gate-verdict, PR/CI status, and
release-status fields with identical names and semantics across all three
templates.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): expand routing table into conditional applicability matrix
Transform routing-table.md from a simple stage-routing lookup into a
conditional applicability matrix. Retain all 15 pre-existing rows and
the two narrower specialists (product-methodology, c4-diagramming).
Add new rows: backend-engineering, frontend-engineering, cli-builder,
data-engineering, data-architect, agent-evals-and-observability,
platform-engineering, security-audit-methodology, and a conditional
opensource-contributions row (public/OSS repos only).
Every row carries a concrete applicability signal (file-system/artifact
observable) and a concrete skip rule (non-tautological complement).
Add explicit change-surface coverage table mapping all ten mandated
surfaces to >=1 row. Add multi-row composition rule (one lead per
stage, per-stage leads recorded in packet) and no-specialist-needed
fallback. Public catalog names only; link to specialists, do not
duplicate their methodology.
Fulfills: VAL-ROUTING-001, VAL-ROUTING-002, VAL-ROUTING-003,
VAL-ROUTING-004, VAL-ROUTING-005, VAL-ROUTING-015, VAL-ROUTING-016,
VAL-ROUTING-017.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add five gates and four delivery paths to stages.md
Define five explicit change-request gates with verdict semantics
(pass/conditional/blocked), phase boundaries, and verdict owners:
- Gate 1: architecture/design delta approval before planning
- Gate 3: spec + task-plan completeness before planning exit
- Gate 2: QA-owned verification plan before implementation
- Gate 4: independent review with per-dimension reviewer mapping
- Gate 5: boundary verification with material/non-material definitions
Add four delivery paths (lightweight, full, refactor, high-risk) with
mandatory vs conditional phase matrices and skip criteria. Stages.md
declared single source of truth for path matrices (VAL-ROUTING-023).
Fulfills: VAL-ROUTING-006 through VAL-ROUTING-014, VAL-ROUTING-018
through VAL-ROUTING-022.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add canonical change-request journey reference
Author references/journey.md with the 9-phase platform-neutral journey:
intake/provenance, discovery/reproduction, architecture/design delta,
specification/decomposition, pre-implementation test planning, domain
implementation, independent review/boundary verification, readiness/CI
loops with exact-final-head re-verification, and authorized post-merge
release/closeout. Each phase specifies owner, input, output, gate, and
escalation condition with a GitHub/enterprise platform mapping column.
Includes four delivery paths (lightweight, full, refactor, high-risk)
referencing stages.md as single source of truth, conditional specialist
routing with recorded skip reasons, no-change-needed termination, the
materiality rule referencing stages.md canonical definition, PR readiness
vs release authority separation, and contiguous phase input/output chain
through named delivery-packet field groups.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add change lifecycle reference with GitHub and enterprise modes
Author references/lifecycle.md with two documented modes sharing the same
nine-phase sequence, five gates, and delivery-packet field group (i):
- GitHub reference mode: issue snapshot before planning; four-class authority
(contributor/maintainer/merge/release); duplicate/PR/branch/maintainer-direction
checks; fork-vs-branch determination; issue linkage + closing-keyword discovery;
readiness/submission/merge separation; CI triage and review monitoring; material
post-submission re-verification through gates 4 and 5; exact-final-head binding
updated per review round; terminal states (merged/closed/blocked/released) with
evidence; release readiness vs release activity separation; post-release
verification evidence; external cancellation path; conditional delegation to
opensource-contributions for public/OSS repos only; no repo-specific hardcoding.
- Enterprise mode: source-of-truth snapshot (ticket/email/verbal); ticket-tracker
dedup/existing-CR check; explicit named-approver approval gate feeding merge;
enterprise CI; change-governance boundaries (CAB, change freeze); release
authority separation; packet portability mapping.
Also update routing-table.md with platform-neutrality framing note for
bundle-wide consistency (VAL-FORMAT-009).
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add merge gate and release gate to risk-authority-gates
Define the merge gate (exact-final-head SHA + CI passing + approved
review status as mandatory preconditions) and the release gate (explicit
authorization beyond merge authority: human grant or documented
pre-delegated permission). Separate pre-merge release readiness
assessment from release activity execution. All existing content
(authority classes, mutation gate, hard stops, stop-and-escalate rules,
recording requirement) is retained unchanged.
Satisfies VAL-LIFECYCLE-018, VAL-LIFECYCLE-019, VAL-CROSS-020,
VAL-CROSS-021.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): wire journey, lifecycle, and delivery packet into umbrella SKILL.md and README
Extend the SKILL.md description trigger to cover change-request / issue-to-PR /
issue-to-release work while retaining all existing fix/build/refactor/review/
verify/release triggers. Add a negative boundary: the journey is not loaded for
plain fixes, refactors, or reviews without an issue/ticket trajectory.
Add a "Change-request work (conditional)" body section pointing to journey.md
with explicit exclusion rules. Update the file map with rows for journey.md,
lifecycle.md, delivery-packet.md, and templates/delivery-packet.md, each with
both a positive scope (change-request work) and a negative scope (not for simple
fix/refactor/review without an issue trajectory).
Update README.md What You Get table with rows for the new references, template,
and evals. Update the Triggers section with change-request and issue-to-PR
delivery triggers additively.
Regenerate llms.txt and .claude-plugin/marketplace.json to reflect the updated
description.
Fulfills: VAL-JOURNEY-010, VAL-JOURNEY-013, VAL-FORMAT-001..005, VAL-FORMAT-010,
VAL-FORMAT-017, VAL-FORMAT-018, VAL-FORMAT-021, VAL-FORMAT-022, VAL-CROSS-010,
VAL-CROSS-018.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add schema-v1 evals manifest with nine trajectory cases
Create bundles/neckbeard/evals/evals.json (schema v1, skill_name: neckbeard)
with nine output-quality cases covering all required trajectory scenarios:
1. Straightforward bug fix with reproduction and regression test
2. Ambiguous feature requiring product discovery and scope gate
3. Multi-surface change (backend/frontend/API/data routing)
4. Schema/migration change with rollback and release-readiness evidence
5. Refactor with characterization tests and architecture review
6. Docs-only reduced path with comprehensive skip recording
7. Existing PR/duplicate work detection and deferral
8. Review round changing final head requiring re-verification
9. Release-authority-blocked terminal state
Cases cover both successful trajectories (merged) and bounded
escalation/blockage (blocked, closed). Assertions verify routing
decisions, artifact production, gate behavior, skip reasons,
exact-head binding, and terminal lifecycle state — not response
length. Each assertion is judge-decidable from run-produced
artifacts (delivery packet fields, gate verdicts, bound SHAs).
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(neckbeard): add trajectory evaluation fixtures and extend harness
Add integrated multi-phase trajectory fixtures under eval/fixtures/trajectories/
and extend run_eval.py to validate them without breaking the 10 existing
single-task fixtures.
- full-change-request: nine-phase intake→release fixture traversing all 9
journey phases to a released terminal state with all 5 gates recorded and
final verdict bound to a head SHA.
- reduced-docs-only: lightweight path fixture ending closed with recorded
skips for phases 2-5, gates 1-3, and 8 routing-table specialists.
- run_eval.py: classify fixtures by kind discriminator, validate trajectory
sub-schema (phase/gate labels against journey.md canonical names, path
membership, terminal state, head SHA binding for full path), report counts
for both fixture kinds.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* docs(neckbeard): update evaluation docs for trajectory fixtures
- evaluation.md: add worked example (validate-only command + expected
stdout), trajectory fixtures section, trajectory scoring guidance
(outcome dims not phase-count/shape), strengthened claims scoping
- baseline-protocol.md: add trajectory comparison section (context-
equivalence, shape-neutrality, skip transparency, terminal state
equivalence) and claims scoping section
- task-schema.md: add trajectory fixture schema table (kind, path,
phases, gates, terminal_state, skipped_phases, skipped_gates,
final_head_sha, routing fields) with full-path constraints and
trajectory layout
Fulfills VAL-EVALS-009, VAL-EVALS-012, VAL-EVALS-015, VAL-EVALS-016.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* fix(neckbeard): use rsplit for phase-label parsing and document routing fields as metadata-only
Make _validate_phase_labels consistent with _validate_gate_labels by
splitting on the last ": " (rsplit) instead of the first occurrence
(find), so colons inside phase names do not break skip-reason parsing.
Document routing_selected/routing_skipped as metadata-only in
task-schema.md: the runner does not cross-validate these fields against
the routing table, because coupling the harness to the routing table
markdown format would add fragile parsing without improving fixture
correctness. Reviewers verify routing entries during trajectory scoring.
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
---------
Co-authored-by: username <username>
Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
* feat(raleigh): add Raleigh-Wake ECC active incident feed adapter
Add an isolated, capability-limited adapter for the RWECC public
incident feed (incidents.rwecc.com/getdata). The adapter:
- Calls only the fixed public host and exact read-only endpoint
- Supports agency, incident type, and bounded result filters
- Preserves source identifiers without implying a historical archive
- Labels every response as a filtered active feed, not all 911 calls
- Fails/warns distinctly on empty, stale, malformed, and unavailable
responses; empty feed is not reported as proof of no incidents
- Deduplicates records and guards against schema drift
- Can be disabled independently via RALEIGH_DISABLE_INCIDENTS=1
- Uses a 90-second cache lifetime
Closes#123
* chore: regenerate catalog artifacts for updated raleigh description
---------
Co-authored-by: magnus919 <magnus919>
Add a fire command group with two subcommands (incidents, response-times)
that resolve the stable ArcGIS item IDs for RFD's full public history
(2007–present) and past-month feeds, with source-aware filtering and a
normalization layer for the 2026 classification schema transition.
- New raleighlib/fire.py module with item resolution, field discovery,
WHERE clause construction (TIMESTAMP literals; these layers reject
epoch-ms date comparisons), and era-aware normalization
- CLI commands: fire incidents, fire response-times with --source,
--since (now supports y), --station, --platoon, --group, --type,
--limit, --offset
- Normalizes incident_type/incident_type_description (pre-2026) and
incident_group_name/incident_subgroup_code/incident_type_name (2026+)
into stable keys without fabricating cross-era mappings; raw fields
preserved in JSON
- Response durations in labeled seconds; missing, reversed, and
malformed timestamp pairs rejected per-pair, never zero-filled
- Documents RFD's exclusion of incident types 300–399 and 661
(EMS/privacy)
- 45 unit tests covering both sides of the 2026 transition, injection
safety, pagination, CLI dispatch, and error paths
- 3 new eval cases: transition normalization, response-time units,
EMS privacy exclusion
- references/fire-reference.md with schemas, transition rules, caveats
Closes#122
Co-authored-by: magnus919 <magnus919>
Add a police command group with four subcommands (incidents, recent,
previous-day, history) that resolve stable ArcGIS item IDs at runtime
and provide source-aware, filter-friendly access to RPD incident data.
- New raleighlib/police.py module with item resolution, field discovery,
WHERE clause construction, and location privacy handling
- CLI commands: police incidents, police recent, police previous-day,
police history with --since, --category, --district, --limit, --offset
- Source labeling (_source, _item_id, _retrieved_at, _location_status)
- Redacted/zero-coordinate suppression (never emits fake points)
- LIKE wildcard and SQL quote escaping for user-supplied filters
- NIBRS date boundary warning when --since predates June 2014
- 30 unit tests covering source selection, filters, injection, redaction,
pagination, CLI dispatch, and error paths
- 5 new eval cases rejecting arrest/conviction/completeness claims
- references/police-reference.md with field schemas and caveats
Closes#121
Co-authored-by: magnus919 <magnus919>
Add a live canary that runs the full Hub catalog check (190 endpoints)
and probes all fixed non-Hub adapters (geocode, transit GTFS,
development, civic JSON:API, civic RSS, meetings eSCRIBE, imagery).
Validates source-specific minimum schemas, classifies failures into
transport_outage, auth_regression, arcgis_error, schema_drift,
parser_failure, and empty_but_valid. Retries bounded transient
failures and preserves first-failure evidence.
Runs daily at 10:17 UTC and on workflow_dispatch, separate from PR CI
so upstream outages do not block unrelated changes.
Closes#120
Co-authored-by: magnus919 <magnus919>
Rewrite raleigh/evals/evals.json with machine-gradable assertions and
expand to 12 cases covering the public-safety acceptance criteria:
- RPD incidents: current official source, no stale endpoints, no
completeness claims
- RPD privacy: randomized/redacted location language, no exact-address
claims
- RFD classification: current fields from live metadata, no deprecated
or hardcoded schemas
- Dispatch: labeled as filtered public feed, not all 911 calls
- Empty feeds: missing evidence, not proof of zero incidents
- Security refusal: write ops, arbitrary hosts, private portals
All assertions use deterministic grader patterns (response_contains,
response_not_contains, exit_status, activation_evidence_contains).
Cases are tagged with case_set (regression/release) for the release
evaluation layer from #106.
The paired eval pipeline runs in CI with the fake adapter (validates
pipeline, 0 regressions). Real model grading runs on the self-hosted
runner post-merge.
Documents local eval suite execution in README.
Closes#118
Co-authored-by: magnus919 <magnus919>
Add release evaluation layer on top of deterministic paired execution:
- schemas/release-eval-v1.schema.json: release report schema with freeze
snapshot, per-case trial aggregation, rubric graders, blinded pairwise
comparison, calibration tracking, and PASS/CONDITIONAL/HOLD/BLOCK outcomes
- eval_runner/release.py: core module for multi-trial aggregation, versioned
rubric graders with abstain/insufficient-evidence, pairwise planning with
position randomization and order-reversal testing, calibration records,
and release decision computation
- eval_runner/tests/test_release.py: 34 tests covering all acceptance criteria
- schemas/evals-v1.schema.json: optional case_set field (dev/regression/release)
- eval_runner/models.py: case_set on EvalCase
- eval_runner/runner.py: load case_set from manifest
- .github/workflows/skill-eval.yml: run release tests in CI
Gate semantics: hard invariants (privacy, auth, destructive) tolerate zero
violations and cannot be averaged away. Missing evidence produces HOLD, not
PASS. Uncalibrated judge results are advisory only.
Closes#106
Co-authored-by: magnus919 <magnus919>
* feat: run isolated paired candidate and baseline skill evaluations
Build the first complete paired skill-evaluation path: stage an immutable
candidate, run matched candidate and baseline trials in clean environments,
execute deterministic outcome graders, and produce a case-level comparison
report.
- eval_runner/sandbox.py: stages production-visible skill surface read-only,
excludes eval manifests/rubrics/oracles from subject sandbox
- eval_runner/grader.py: deterministic assertion checker (7 assertion types)
- eval_runner/comparison.py: paired comparison report generation
- eval_runner/paired.py: orchestrator CLI (fake, cli, openai adapters)
- eval_runner/openai_adapter.py: OpenAI-compatible API adapter with
chat_template_kwargs support (enable_thinking toggle)
- schemas/comparison-report-v1.schema.json: report schema
- .github/workflows/skill-eval.yml: CI smoke (fake adapter on ubuntu,
real model on self-hosted runner when endpoint reachable)
- yc-default-alive-calculator/evals/evals.json: initial 5-case eval manifest
Verified against google_gemma-4-26B-A4B-it-IQ4_XS.gguf: 5/5 candidate
improvements, 0 regressions.
Closes#105
* ci: make paired-eval-model job non-blocking
The self-hosted runner may not always be online. Mark the job
continue-on-error so it doesn't gate PRs when the runner is unavailable.
* ci: isolate model evals from pull requests
* fix(raleigh): test arrivals against a daily route, not weekday-only
The fixture only had a WEEK (Mon-Fri) service, so
test_get_arrivals_for_stop returned 0 arrivals on weekends when
_today_date() fell on Saturday/Sunday. Add a DAILY service with trip T3
on route R2 and assert against it — the test now passes regardless of
what day CI runs.
* ci: trigger checks on amended commit
---------
Co-authored-by: magnus919 <magnus919>
Wire the Phase 3 ratchet into CI by passing the PR base SHA to
eval-coverage.py --modified-from. Expand changed-skill detection from
SKILL.md-only diffs to the entire skill directory so that references,
scripts, fixtures, README, and eval manifest edits all count as
modifications. Add a monotonic coverage floor that fails CI when
coverage decreases between the base and candidate revisions.
Add script tests for ratchet-mode detection and coverage-decrease
behaviour. Update AGENTS.md and CONTRIBUTING.md to describe the
behaviour CI now enforces.
Closes#102
Signed-off-by: Magnus Hedemark <magnus919@pm.me>
Align AGENTS.md and CONTRIBUTING.md with the description-quality gates (#97, #98), eval coverage ratchet (#99), and trigger-boundary and eval requirements (#100) merged today. Document that CI validates generated artifact freshness but does not regenerate; contributors run generators locally with --write.
Co-authored-by: magnus919 <magnus919>
Align the meta-skill workflow with the repository's description-quality and eval-coverage gates. Keep harness-specific trigger checks separate from portable output-quality evals.
Co-authored-by: magnus919 <magnus919>
Phase 1: New skills (not in grandfathered-skills.txt) must have
evals/evals.json with at least 5 test cases. All 107 existing skills
are grandfathered.
Phase 2: scripts/eval-coverage.py reports coverage (skills with/without
evals, case counts, reference-priority sorting). Added as informational
CI step.
Phase 3: Ratchet thresholds — at 25% coverage, modified skills without
evals get a warning; at 50%, they fail CI. Enforced via
--modified-from flag for PR-scoped checks.
Closes#90
Two new reference files extending qa-methodology into the reactive side
of its domain — diagnosing failures rather than designing strategy.
ci-failure-triage.md: systematic CI failure diagnosis — runner
availability checks, log triage (gh run view), exit 137 / container
termination evidence-first procedure, pre-existing vs regression
classification, flaky test management, and compose readiness corollary.
Distilled from accumulated CI-failure incident notes.
test-debugging.md: diagnosing broken tests — mock path binding after
module-to-package refactors, FastAPI startup race (mock state set before
TestClient context is overwritten), httpx mock transport pattern, test
execution integrity (collection count vs exit code), deterministic
integration seeds, API signature change fixture recovery, and uv
lockfile hygiene. Distilled from accumulated test-debugging incident
notes.
Both are technique libraries serving qa-methodology's existing domain,
not new standalone skills. SKILL.md reference table updated.
Signed-off-by: Magnus Hedemark <magnus919@pm.me>
Generate a root discovery catalog from public skill frontmatter and fail CI
when the committed index drifts. Add a fixture-based regression test for
bundle paths, nested-helper exclusion, normalized descriptions,
deterministic ordering, and stale-file recovery.
Closes#78
AI-assisted: yes (Jasper/Hermes Agent)
Co-authored-by: magnus919 <magnus919>
The generator used File.basename which stripped the bundles/ prefix,
emitting ./neckbeard instead of ./bundles/neckbeard. Codex discovered
92/96 skills — the 4 bundle entrypoints were missing because their
paths didn't resolve.
Verified with live Codex CLI: all 96 skills now discoverable.
AI-assisted: yes (Jasper/Hermes Agent)
* fix: SkillOpt Epoch 1 — neckbeard description trigger-verb-first
Move trigger verbs (fix/build/refactor/review/verify/release) to the front
of the description for better discoverability. Negative case (non-software
questions) now correctly rejected. Validation: 4/6 held-out tasks correct.
* fix: SkillOpt Epoch 2 — wire overlooked catalog specialists into routing
Add 6 stage-owning methodology skills to the routing table and SKILL.md
summary: secure-software-engineering, web-accessibility, qa-methodology,
product-design-and-ux, api-design-and-evolution, site-reliability-engineering.
Note product-methodology and c4-diagramming as narrower composers.
Baseline rollout showed security reviews, UI features, and regression-safety
questions all routed without their natural specialist. Validation: 6/6 held-out
routing tasks now route correctly (baseline 3/6). All additions are
agent-agnostic methodology skills; no Hermes/deployment/personal content.
* fix: SkillOpt Epoch 3 — align stages.md Stage 4 with expanded routing
Stage 4 execution flow now names the same specialists added to the routing
table in Epoch 2: secure-software-engineering and web-accessibility for
implementation, qa-methodology for verification, product-design-and-ux and
api-design-and-evolution for design, site-reliability-engineering for delivery.
Rollout confirmed the gap: an agent following stages.md alone would route a
security-sensitive change (untrusted input, trust-boundary crossing) with no
security specialist. Validation: PASS — stages.md now names the specialist.
* fix: SkillOpt final validation — remove Hermes-specific skill_view reference
Replace skill_view(name="neckbeard") with agent-agnostic "read SKILL.md"
in Quick Start. Public skill must not reference Hermes-specific APIs.
* chore: regenerate Claude marketplace for neckbeard description update
* feat: add Codex plugin packaging (single-plugin, metadata-only)
Adds .codex-plugin/plugin.json with a skills array listing all public
skills, plus .agents/plugins/marketplace.json for one-command install:
codex plugin marketplace add magnus919/agent-skills
codex plugin install magnus919
Same pattern as mattpocock/skills — one plugin, explicit skill paths,
no dist/, no curation, no duplication. Bundle-internal helpers excluded
by the shared glob. CI check mode fails if the manifest drifts.
Closes#79
AI-assisted: yes (Jasper/Hermes Agent)
* chore: trigger CI
* chore: regenerate Codex plugin manifest to include neckbeard bundle
* feat: add neckbeard, an evidence-driven SDLC skill bundle
A portable operating model for software delivery that routes a change through
framing, discovery, design, implementation, review, verification, delivery, and
learning. Chooses the smallest *safe* intervention (minimalism as a consequence
of understanding, not a reflex), proves it at the real delivery boundary, and
leaves an inspectable evidence ledger.
Design responds directly to the Ponytail/YAGNI benchmark critique: no persona,
no LOC-as-success-proxy, no universal performance claims. Composes the specialist
catalog (product-discovery, spec-driven-development, software-architecture-analysis,
systematic-debugging, technical-documentation, verification-methodology) via an
explicit routing table rather than duplicating it.
Ships a versioned evaluation harness (task schema, scoring rubric, baseline
protocol, runner, and 10 fixtures across all 9 task classes incl. adversarial and
no-change-needed cases) that measures SDLC outcomes, never LOC or brevity.
Closes#25
* chore: regenerate Claude marketplace for neckbeard
Adds .claude-plugin/marketplace.json exposing all 95 public skills as
installable plugins via /plugin marketplace add magnus919/agent-skills.
Metadata-only approach: each entry uses source './' + skills ['./<name>']
+ strict:false, so no per-skill plugin.json or directory restructuring is
needed. Bundle-internal helper skills are excluded; bundle entrypoints are
included.
- scripts/gen-claude-marketplace.rb: generates and validates the manifest
- CI step fails if marketplace.json drifts from the skill tree
- README: Claude Code install instructions
Closes#76
AI-assisted: yes (Jasper/Hermes Agent)
The catalog entries don't need individual install commands — the
general 'Hermes Agent' section already explains how skills load.
Per-skill snippets are noise that has to be maintained for every
new skill.
Accumulates conventional commits into a Release PR that bumps the
version and updates CHANGELOG.md. Nothing is tagged until a human
merges the Release PR.
Co-authored-by: Jasper <jasper@magnus919.com>