* feat(raleigh): add Raleigh-Wake ECC active incident feed adapter
Add an isolated, capability-limited adapter for the RWECC public
incident feed (incidents.rwecc.com/getdata). The adapter:
- Calls only the fixed public host and exact read-only endpoint
- Supports agency, incident type, and bounded result filters
- Preserves source identifiers without implying a historical archive
- Labels every response as a filtered active feed, not all 911 calls
- Fails/warns distinctly on empty, stale, malformed, and unavailable
responses; empty feed is not reported as proof of no incidents
- Deduplicates records and guards against schema drift
- Can be disabled independently via RALEIGH_DISABLE_INCIDENTS=1
- Uses a 90-second cache lifetime
Closes#123
* chore: regenerate catalog artifacts for updated raleigh description
---------
Co-authored-by: magnus919 <magnus919>
Add a fire command group with two subcommands (incidents, response-times)
that resolve the stable ArcGIS item IDs for RFD's full public history
(2007–present) and past-month feeds, with source-aware filtering and a
normalization layer for the 2026 classification schema transition.
- New raleighlib/fire.py module with item resolution, field discovery,
WHERE clause construction (TIMESTAMP literals; these layers reject
epoch-ms date comparisons), and era-aware normalization
- CLI commands: fire incidents, fire response-times with --source,
--since (now supports y), --station, --platoon, --group, --type,
--limit, --offset
- Normalizes incident_type/incident_type_description (pre-2026) and
incident_group_name/incident_subgroup_code/incident_type_name (2026+)
into stable keys without fabricating cross-era mappings; raw fields
preserved in JSON
- Response durations in labeled seconds; missing, reversed, and
malformed timestamp pairs rejected per-pair, never zero-filled
- Documents RFD's exclusion of incident types 300–399 and 661
(EMS/privacy)
- 45 unit tests covering both sides of the 2026 transition, injection
safety, pagination, CLI dispatch, and error paths
- 3 new eval cases: transition normalization, response-time units,
EMS privacy exclusion
- references/fire-reference.md with schemas, transition rules, caveats
Closes#122
Co-authored-by: magnus919 <magnus919>
Add a police command group with four subcommands (incidents, recent,
previous-day, history) that resolve stable ArcGIS item IDs at runtime
and provide source-aware, filter-friendly access to RPD incident data.
- New raleighlib/police.py module with item resolution, field discovery,
WHERE clause construction, and location privacy handling
- CLI commands: police incidents, police recent, police previous-day,
police history with --since, --category, --district, --limit, --offset
- Source labeling (_source, _item_id, _retrieved_at, _location_status)
- Redacted/zero-coordinate suppression (never emits fake points)
- LIKE wildcard and SQL quote escaping for user-supplied filters
- NIBRS date boundary warning when --since predates June 2014
- 30 unit tests covering source selection, filters, injection, redaction,
pagination, CLI dispatch, and error paths
- 5 new eval cases rejecting arrest/conviction/completeness claims
- references/police-reference.md with field schemas and caveats
Closes#121
Co-authored-by: magnus919 <magnus919>
Add a live canary that runs the full Hub catalog check (190 endpoints)
and probes all fixed non-Hub adapters (geocode, transit GTFS,
development, civic JSON:API, civic RSS, meetings eSCRIBE, imagery).
Validates source-specific minimum schemas, classifies failures into
transport_outage, auth_regression, arcgis_error, schema_drift,
parser_failure, and empty_but_valid. Retries bounded transient
failures and preserves first-failure evidence.
Runs daily at 10:17 UTC and on workflow_dispatch, separate from PR CI
so upstream outages do not block unrelated changes.
Closes#120
Co-authored-by: magnus919 <magnus919>
Rewrite raleigh/evals/evals.json with machine-gradable assertions and
expand to 12 cases covering the public-safety acceptance criteria:
- RPD incidents: current official source, no stale endpoints, no
completeness claims
- RPD privacy: randomized/redacted location language, no exact-address
claims
- RFD classification: current fields from live metadata, no deprecated
or hardcoded schemas
- Dispatch: labeled as filtered public feed, not all 911 calls
- Empty feeds: missing evidence, not proof of zero incidents
- Security refusal: write ops, arbitrary hosts, private portals
All assertions use deterministic grader patterns (response_contains,
response_not_contains, exit_status, activation_evidence_contains).
Cases are tagged with case_set (regression/release) for the release
evaluation layer from #106.
The paired eval pipeline runs in CI with the fake adapter (validates
pipeline, 0 regressions). Real model grading runs on the self-hosted
runner post-merge.
Documents local eval suite execution in README.
Closes#118
Co-authored-by: magnus919 <magnus919>
Add release evaluation layer on top of deterministic paired execution:
- schemas/release-eval-v1.schema.json: release report schema with freeze
snapshot, per-case trial aggregation, rubric graders, blinded pairwise
comparison, calibration tracking, and PASS/CONDITIONAL/HOLD/BLOCK outcomes
- eval_runner/release.py: core module for multi-trial aggregation, versioned
rubric graders with abstain/insufficient-evidence, pairwise planning with
position randomization and order-reversal testing, calibration records,
and release decision computation
- eval_runner/tests/test_release.py: 34 tests covering all acceptance criteria
- schemas/evals-v1.schema.json: optional case_set field (dev/regression/release)
- eval_runner/models.py: case_set on EvalCase
- eval_runner/runner.py: load case_set from manifest
- .github/workflows/skill-eval.yml: run release tests in CI
Gate semantics: hard invariants (privacy, auth, destructive) tolerate zero
violations and cannot be averaged away. Missing evidence produces HOLD, not
PASS. Uncalibrated judge results are advisory only.
Closes#106
Co-authored-by: magnus919 <magnus919>
* feat: run isolated paired candidate and baseline skill evaluations
Build the first complete paired skill-evaluation path: stage an immutable
candidate, run matched candidate and baseline trials in clean environments,
execute deterministic outcome graders, and produce a case-level comparison
report.
- eval_runner/sandbox.py: stages production-visible skill surface read-only,
excludes eval manifests/rubrics/oracles from subject sandbox
- eval_runner/grader.py: deterministic assertion checker (7 assertion types)
- eval_runner/comparison.py: paired comparison report generation
- eval_runner/paired.py: orchestrator CLI (fake, cli, openai adapters)
- eval_runner/openai_adapter.py: OpenAI-compatible API adapter with
chat_template_kwargs support (enable_thinking toggle)
- schemas/comparison-report-v1.schema.json: report schema
- .github/workflows/skill-eval.yml: CI smoke (fake adapter on ubuntu,
real model on self-hosted runner when endpoint reachable)
- yc-default-alive-calculator/evals/evals.json: initial 5-case eval manifest
Verified against google_gemma-4-26B-A4B-it-IQ4_XS.gguf: 5/5 candidate
improvements, 0 regressions.
Closes#105
* ci: make paired-eval-model job non-blocking
The self-hosted runner may not always be online. Mark the job
continue-on-error so it doesn't gate PRs when the runner is unavailable.
* ci: isolate model evals from pull requests
* fix(raleigh): test arrivals against a daily route, not weekday-only
The fixture only had a WEEK (Mon-Fri) service, so
test_get_arrivals_for_stop returned 0 arrivals on weekends when
_today_date() fell on Saturday/Sunday. Add a DAILY service with trip T3
on route R2 and assert against it — the test now passes regardless of
what day CI runs.
* ci: trigger checks on amended commit
---------
Co-authored-by: magnus919 <magnus919>
Wire the Phase 3 ratchet into CI by passing the PR base SHA to
eval-coverage.py --modified-from. Expand changed-skill detection from
SKILL.md-only diffs to the entire skill directory so that references,
scripts, fixtures, README, and eval manifest edits all count as
modifications. Add a monotonic coverage floor that fails CI when
coverage decreases between the base and candidate revisions.
Add script tests for ratchet-mode detection and coverage-decrease
behaviour. Update AGENTS.md and CONTRIBUTING.md to describe the
behaviour CI now enforces.
Closes#102
Signed-off-by: Magnus Hedemark <magnus919@pm.me>
Align AGENTS.md and CONTRIBUTING.md with the description-quality gates (#97, #98), eval coverage ratchet (#99), and trigger-boundary and eval requirements (#100) merged today. Document that CI validates generated artifact freshness but does not regenerate; contributors run generators locally with --write.
Co-authored-by: magnus919 <magnus919>
Align the meta-skill workflow with the repository's description-quality and eval-coverage gates. Keep harness-specific trigger checks separate from portable output-quality evals.
Co-authored-by: magnus919 <magnus919>
Phase 1: New skills (not in grandfathered-skills.txt) must have
evals/evals.json with at least 5 test cases. All 107 existing skills
are grandfathered.
Phase 2: scripts/eval-coverage.py reports coverage (skills with/without
evals, case counts, reference-priority sorting). Added as informational
CI step.
Phase 3: Ratchet thresholds — at 25% coverage, modified skills without
evals get a warning; at 50%, they fail CI. Enforced via
--modified-from flag for PR-scoped checks.
Closes#90
Two new reference files extending qa-methodology into the reactive side
of its domain — diagnosing failures rather than designing strategy.
ci-failure-triage.md: systematic CI failure diagnosis — runner
availability checks, log triage (gh run view), exit 137 / container
termination evidence-first procedure, pre-existing vs regression
classification, flaky test management, and compose readiness corollary.
Distilled from accumulated CI-failure incident notes.
test-debugging.md: diagnosing broken tests — mock path binding after
module-to-package refactors, FastAPI startup race (mock state set before
TestClient context is overwritten), httpx mock transport pattern, test
execution integrity (collection count vs exit code), deterministic
integration seeds, API signature change fixture recovery, and uv
lockfile hygiene. Distilled from accumulated test-debugging incident
notes.
Both are technique libraries serving qa-methodology's existing domain,
not new standalone skills. SKILL.md reference table updated.
Signed-off-by: Magnus Hedemark <magnus919@pm.me>
Generate a root discovery catalog from public skill frontmatter and fail CI
when the committed index drifts. Add a fixture-based regression test for
bundle paths, nested-helper exclusion, normalized descriptions,
deterministic ordering, and stale-file recovery.
Closes#78
AI-assisted: yes (Jasper/Hermes Agent)
Co-authored-by: magnus919 <magnus919>
The generator used File.basename which stripped the bundles/ prefix,
emitting ./neckbeard instead of ./bundles/neckbeard. Codex discovered
92/96 skills — the 4 bundle entrypoints were missing because their
paths didn't resolve.
Verified with live Codex CLI: all 96 skills now discoverable.
AI-assisted: yes (Jasper/Hermes Agent)