* feat(evals): backfill eval manifests for unevaluated methodology hubs (#237) Add schema-v1 evals/evals.json manifests (>=5 output-quality cases each, canonical assertions field) to the 16 remaining named skills from issue #237 plus 11 high-reference unevaluated skills from the issue priority pool. Raises schema-valid eval coverage from 44/132 (33.3%) to 71/132 (53.8%), clearing the 50% CI-fail threshold. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> * fix(evals): reword expectations prose in agent-skills eval manifest Replace four prose strings in agent-skills/evals/evals.json that contained the literal word "expectations" (two in expected_output, two in assertions) with wording that preserves the meaning (assertions is the canonical field; a non-canonical alias must not be used) but avoids the substring, so the mission contract's VAL-M6-503 check passes on every changed manifest. Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com> --------- Co-authored-by: factory-droid[bot] <138933559+factory-droid[bot]@users.noreply.github.com>
Agent Evals and Observability
Build evidence for AI-agent changes without confusing a dashboard with proof of quality.
Why Install This Skill
Agent behavior can look good in a demo yet fail through an unsafe tool call, a bad recovery path, stale data, or a silent production regression. This skill helps your agent turn those risks into task and trajectory contracts, datasets, appropriate graders, and release evidence.
It also keeps observability useful without turning it into a privacy liability. Your agent can design minimized traces and metrics, analyze a regression fairly, and make a release decision that keeps hard safety and privacy invariants separate from ordinary quality indicators.
What You Get
| Contents | Provides |
|---|---|
SKILL.md |
Framework-neutral workflow and routing |
references/ |
Evaluation, statistics, trajectory, privacy, OTel, and source guidance |
templates/ |
Fillable plans, manifests, grader specs, reviews, reports, and gates |
Quick Start
Ask: Create an eval plan and release gate for this agent change.
Expected result: a risk-based plan that names the task contract, evidence, privacy limits, uncertainty, rollback path, and decision owner.
Triggers
- Agent evaluation, LLM evals, evaluation dataset, grader, or model judge
- Agent observability, traces, telemetry, trajectory review, or production monitoring
- Regression analysis, prompt/model/tool release gate, or incident-to-eval learning
- Privacy-aware logging, redaction, retention, or trace sampling for an agent
Requirements
No package, vendor account, or API key is required. Use the agent framework and telemetry backend already selected by the project. OpenTelemetry GenAI is optional interoperability guidance only.