* feat(golang-google-wire): add compile-time DI skill with full reference suite
New skill for google/wire (v0.7.0, archived 2025-08). Covers wire.NewSet,
wire.Build, wire.Bind, wire.Struct, wire.Value, wire.InterfaceValue,
wire.FieldsOf, //go:build wireinject, cleanup chains, and wire_gen.go
workflow.
Files:
- skills/golang-google-wire/SKILL.md (2,511 tok)
- skills/golang-google-wire/references/advanced.md (cleanup chains, error
catalogue, codegen flags, panic syntax, external value injection)
- skills/golang-google-wire/references/recipes.md (HTTP server, multi-build,
cleanup-heavy graph, CLI embedding)
- skills/golang-google-wire/references/testing.md (test injectors, CI stale
check, fake bindings)
- skills/golang-google-wire/evals/evals.json (10 evals, 50 assertions)
Eval results (LLM-as-Judge, claude-sonnet-4-6):
With skill: 49/50 (98%)
Without skill: 41/50 (82%)
Delta: +16pp
Key differentiators: wire.FieldsOf (eval 8, 5/5 vs 1/5 — model actively
discouraged FieldsOf without skill), //go:build wireinject tag (eval 1,
5/5 vs 3/5 — suggested //go:build ignore instead).
README: updated Tools table row and Skill evaluations total (38 skills,
3232 assertions).
* fix(golang-google-wire): apply prettier formatting
* fix(golang-google-wire): address PR review comments
- advanced.md: correct cleanup order description (dependants torn down
before dependencies, not the reverse)
- recipes.md: fix multi-build variant section — show -output_file_prefix
commands first to avoid both wire runs overwriting wire_gen.go
- recipes.md: handle http.ErrServerClosed in ListenAndServe goroutine
so graceful shutdown is not logged as a fatal error
- testing.md: fix map literals (use &User{} composite literal)
- testing.md: rephrase //go:build !wireinject comment — it IS a build
tag, not the absence of one
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: samber <2951285+samber@users.noreply.github.com>
885 KiB
Evaluations
Summary
| Skill | Version | Assertions | With Skill | Without Skill | Delta | Uplift | Concern |
|---|---|---|---|---|---|---|---|
golang-spf13-viper |
v1.0.0 | 53 | 98% | 100% | -2pp | 0.98× | Low delta, high without |
golang-spf13-cobra |
v1.0.0 | 50 | 100% | 98% | +2pp | 1.02× | Low delta, high without |
golang-naming |
v1.0.0 | 51 | 94% | 71% | +24pp | 1.32× | Low delta, high without |
golang-swagger |
v1.0.0 | 60 | 97% | 72% | +25pp | 1.35× | Low delta, high without |
golang-error-handling |
v1.0.0 | 60 | 98% | 72% | +27pp | 1.36× | Low delta, high without |
golang-popular-libraries |
v1.0.0 | 54 | 100% | 70% | +30pp | 1.43× | Low delta, high without |
golang-security |
v1.0.0 | 110 | 100% | 68% | +32pp | 1.47× | Low delta, high without |
golang-testing |
v1.0.0 | 65 | 92% | 60% | +32pp | 1.53× | Low delta |
golang-troubleshooting |
v1.0.0 | 186 | 100% | 68% | +32pp | 1.47× | Low delta, high without |
golang-context |
v1.0.0 | 50 | 96% | 62% | +34pp | 1.55× | |
golang-structs-interfaces |
v1.0.0 | 52 | 100% | 65% | +35pp | 1.54× | High without |
golang-observability |
v1.0.0 | 185 | 100% | 63% | +37pp | 1.59× | |
golang-design-patterns |
v1.0.0 | 87 | 100% | 63% | +37pp | 1.59× | |
golang-database |
v1.0.0 | 74 | 95% | 57% | +38pp | 1.67× | |
golang-project-layout |
v1.0.0 | 55 | 100% | 62% | +38pp | 1.61× | |
golang-data-structures |
v1.0.0 | 36 | 100% | 61% | +39pp | 1.64× | |
golang-performance |
v1.0.0 | 272 | 100% | 61% | +39pp | 1.64× | |
golang-concurrency |
v1.0.0 | 62 | 100% | 61% | +39pp | 1.64× | |
golang-code-style |
v1.0.0 | 83 | 80% | 40% | +40pp | 2.00× | Low with-skill score |
golang-lint |
v1.0.0 | 51 | 96% | 55% | +41pp | 1.75× | |
golang-grpc |
v1.0.0 | 55 | 96% | 55% | +42pp | 1.75× | |
golang-cli |
v1.0.0 | 58 | 95% | 52% | +43pp | 1.83× | |
golang-dependency-injection |
v1.0.0 | 55 | 98% | 51% | +47pp | 1.92× | |
golang-stretchr-testify |
v1.0.0 | 47 | 100% | 53% | +47pp | 1.89× | |
golang-samber-mo |
v1.0.0 | 108 | 88% | 40% | +48pp | 2.20× | Low with-skill score |
golang-benchmark |
v1.0.0 | 356 | 100% | 50% | +50pp | 2.00× | |
golang-samber-ro |
v1.0.0 | 113 | 100% | 50% | +50pp | 2.00× | |
golang-documentation |
v1.0.0 | 103 | 90% | 37% | +53pp | 2.43× | |
golang-samber-hot |
v1.0.0 | 65 | 94% | 40% | +54pp | 2.35× | |
golang-dependency-management |
v1.0.0 | 52 | 100% | 46% | +54pp | 2.17× | |
golang-stay-updated |
v1.0.0 | 50 | 92% | 36% | +56pp | 2.56× | |
golang-safety |
v1.0.0 | 151 | 99% | 41% | +58pp | 2.41× | |
golang-continuous-integration |
v1.0.0 | 66 | 100% | 41% | +59pp | 2.44× | |
golang-samber-oops |
v1.0.0 | 52 | 94% | 35% | +60pp | 2.69× | |
golang-modernize |
v1.0.0 | 76 | 95% | 34% | +61pp | 2.79× | |
golang-samber-slog |
v1.0.0 | 62 | 92% | 73% | +19pp | 1.26× | Low delta, high without |
golang-samber-lo |
v1.0.0 | 86 | 97% | 57% | +40pp | 1.70× | |
golang-uber-fx |
v1.0.0 | 21 | 100% | 95% | +5pp | 1.05× | Low delta, high without |
golang-uber-dig |
v1.0.0 | 20 | 100% | 90% | +10pp | 1.11× | Low delta, high without |
golang-google-wire |
v1.0.0 | 50 | 98% | 82% | +16pp | 1.20× | Low delta, high without |
golang-graphql |
v0.0.2 | 59 | 100% | 83% | +17pp | 1.20× | Low delta, high without |
golang-samber-do |
v1.0.0 | 53 | 100% | 19% | +81pp | 5.26× | |
| Total (41 skills) | 3395 | 98% | 56% | +41pp | 1.73× |
golang-naming — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 48/51 (94%) | 36/51 (71%) | +24pp |
Full breakdown (51 assertions)
Model: Claude Opus 4.6 | Runs: 12 evals × 2 configs = 24 subagents | Grading: automated regex
| # | Assertion | With | Without |
|---|---|---|---|
| 1. cache — Get/Set/Delete, sentinel errors, options | 4/5 | 4/5 | |
| 1.1 | Constructor is New() |
✓ | ✓ |
| 1.2 | Error strings include "cache:" prefix |
✓ | ✓ |
| 1.3 | Option functions use With prefix |
✗ | ✗ |
| 1.4 | Error strings fully lowercase | ✓ | ✓ |
| 1.5 | Option type is Option not CacheOption |
✓ | ✓ |
| 2. worker — boolean fields: processing/initialized/canAccept | 4/4 | 3/4 | |
| 2.1 | Boolean fields use is/has/can prefix |
✓ | ✓ |
| 2.2 | Getter methods use Is/Has/Can prefix |
✓ | ✓ |
| 2.3 | Constructor is New() |
✓ | ✗ |
| 2.4 | Receiver is 1-2 letter abbreviation | ✓ | ✓ |
3. taskqueue — TaskStatus enum (iota) + String() |
2/3 | 2/3 | |
| 3.1 | Enum values prefixed with type name | ✓ | ✓ |
| 3.2 | Zero value is TaskStatusUnknown sentinel |
✗ | ✗ |
| 3.3 | String() method on enum type |
✓ | ✓ |
4. httpclient — fix naming: HttpClient, base_url, GetBaseUrl(), this |
6/7 | 3/7 | |
| 4.1 | Struct renamed to Client (anti-stutter) |
✓ | ✗ kept HTTPClient |
| 4.2 | Constructor renamed to New() |
✓ | ✗ NewHTTPClient() |
| 4.3 | Field uses baseURL (correct acronym) |
✗ | ✗ |
| 4.4 | Getter drops Get prefix: BaseURL() |
✓ | ✓ |
| 4.5 | Boolean getter uses IsConnected() |
✓ | ✓ |
| 4.6 | Receiver is 1-2 letters, not this |
✓ | ✓ |
| 4.7 | Boolean field uses isConnected |
✓ | ✗ bare connected |
5. auth-tests — table-driven tests for ValidateToken() |
4/4 | 2/4 | |
| 5.1 | Subtest names fully lowercase | ✓ | ✗ |
| 5.2 | Acronyms lowercase in subtests | ✓ | ✗ uppercase "ID" |
| 5.3 | Test function named TestValidateToken |
✓ | ✓ |
| 5.4 | Table struct uses descriptive field names | ✓ | ✓ |
| 6. storage — functional options: bucket, region, endpoint URL, retries | 4/4 | 3/4 | |
| 6.1 | Option type is Option not StoreOption |
✓ | ✓ |
| 6.2 | All option functions use With prefix |
✓ | ✓ |
| 6.3 | Constructor is New() |
✓ | ✗ NewStore() |
| 6.4 | URL acronym is all-caps |
✓ | ✓ |
| 7. metrics — MixedCaps constants, sentinel errors | 4/4 | 4/4 | |
| 7.1 | Constants use MixedCaps not ALL_CAPS | ✓ | ✓ |
| 7.2 | Error variables use Err prefix |
✓ | ✓ |
| 7.3 | Error strings fully lowercase | ✓ | ✓ |
| 7.4 | Error strings include "metrics:" prefix |
✓ | ✓ |
| 8. notifier — email/SMS/push interfaces, Dispatcher, error types | 4/4 | 4/4 | |
| 8.1 | Interface uses -er suffix |
✓ | ✓ |
| 8.2 | No stuttering in type names | ✓ | ✓ |
| 8.3 | Constructor is New() |
✓ | ✓ |
| 8.4 | Error types use Error suffix |
✓ | ✓ |
| 9. eventbus — sentinel errors incl. "invalid event ID", custom error type | 4/4 | 2/4 | |
| 9.1 | Sentinel errors use Err prefix |
✓ | ✓ |
| 9.2 | Custom error type uses Error suffix |
✓ | ✓ |
| 9.3 | Error string uses lowercase "id" |
✓ | ✗ "invalid event ID" |
| 9.4 | Error strings include "eventbus:" prefix |
✓ | ✗ no prefix |
| 10. config — Load/Save/Get/Set/Validate/Reset/MergeFrom/String | 3/3 | 2/3 | |
| 10.1 | All methods use consistent 1-2 letter receiver | ✓ | ✓ |
| 10.2 | Constructor is New() |
✓ | ✗ NewConfig() |
| 10.3 | No stuttering in type names | ✓ | ✓ |
| 11. ratelimiter — token bucket, options, Status enum, boolean methods | 5/5 | 3/5 | |
| 11.1 | Constructor is New() |
✓ | ✗ NewLimiter() |
| 11.2 | Status enum has StatusUnknown at iota 0 |
✓ | ✗ Allowed at iota 0 |
| 11.3 | Boolean methods use Is/Has prefix |
✓ | ✓ |
| 11.4 | Option functions use With prefix |
✓ | ✓ |
| 11.5 | Error strings lowercase with package prefix | ✓ | ✓ |
12. parser — Parse, MustParse, ParseWithContext, AST, ParseError |
4/4 | 4/4 | |
| 12.1 | Panic variant uses MustParse |
✓ | ✓ |
| 12.2 | Context variant uses ParseWithContext |
✓ | ✓ |
| 12.3 | Error type uses ParseError |
✓ | ✓ |
| 12.4 | No stuttering on constructor | ✓ | ✓ |
golang-code-style — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 66/83 (80%) | 33/83 (40%) | +40pp |
Full breakdown (83 assertions across 24 evals)
Model: Claude Sonnet 4.6 | Runs: 24 evals × 2 configs = 48 runs | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
1. cache — var declarations: := vs var, zero-value init |
3/3 | 3/3 | |
| 1.1 | Zero-value fields omitted from constructor (rely on Go zero values) | ✓ | ✓ |
| 1.2 | Non-zero assignments use := short declaration |
✓ | ✓ |
| 1.3 | Done channel created with make(), not nil |
✓ | ✓ |
| 2. handler — slice/map initialization: never nil | 3/3 | 2/3 | |
| 2.1 | Empty slice uses []User{} or make(), not nil |
✓ | ✓ |
| 2.2 | Empty map uses map[K]V{} or make(), not nil |
✓ | ✓ |
| 2.3 | Known-capacity collections use make() with size hint |
✓ | ✗ no capacity hint |
| 3. server — composite literals: named fields | 3/3 | 3/3 | |
| 3.1 | http.Server uses named fields |
✓ | ✓ |
| 3.2 | tls.Config uses named fields |
✓ | ✓ |
| 3.3 | No positional struct syntax in file | ✓ | ✓ |
| 4. orders — early return: guard clauses, flat happy path | 3/3 | 0/3 | |
| 4.1 | Validation uses early return, not nested else | ✓ | ✗ deeply nested if-else |
| 4.2 | Happy path at indentation level 1 | ✓ | ✗ level 5 nesting |
| 4.3 | Max 2 indentation levels for main logic | ✓ | ✗ 6 levels deep |
| 5. auth — unnecessary else: early return, default-then-override | 2/2 | 1/2 | |
| 5.1 | GetUserRole uses no else after return |
✓ | ✓ |
| 5.2 | SetLogLevel uses default-then-override (no else) |
✓ | ✗ var zero-value + switch default |
| 6. events — switch vs if-else chains | 3/3 | 2/3 | |
| 6.1 | HandleEvent uses switch |
✓ | ✓ |
| 6.2 | MapStatusCode uses switch |
✓ | ✓ |
| 6.3 | Both switches include default case | ✓ | ✗ HandleEvent no default |
| 7. notify — function design: ≤4 params, options struct | 3/3 | 1/3 | |
| 7.1 | context.Context is first parameter |
✓ | ✓ |
| 7.2 | Options struct reduces params to ≤4 | ✓ | ✗ 10 individual params |
| 7.3 | Config fields grouped in options struct | ✓ | ✗ no struct |
| 8. users — value vs pointer: small types by value | 3/3 | 1/3 | |
| 8.1 | FormatName takes string by value, not *string |
✓ | ✗ *string params |
| 8.2 | UpdateAge takes *User (mutation) |
✓ | ✓ |
| 8.3 | CompareUsers takes User by value (small, read-only) |
✓ | ✗ *User params |
| 9. repository — code organization: canonical ordering | 3/3 | 3/3 | |
| 9.1 | File ordered: imports → constants → interface → struct → ctor → methods → helpers | ✓ | ✓ |
| 9.2 | Interface declared before implementing struct | ✓ | ✓ |
| 9.3 | Helpers after all methods, not interleaved | ✓ | ✓ |
| 10. report — string handling: strconv, Builder, Sprintf | 3/3 | 3/3 | |
| 10.1 | Int-to-string uses strconv.Itoa, not fmt.Sprintf |
✓ | ✓ |
| 10.2 | Loop concatenation uses strings.Builder |
✓ | ✓ |
| 10.3 | Complex formatting uses fmt.Sprintf |
✓ | ✓ |
| 11. authz — complex conditions: named booleans | 1/3 | 1/3 | |
| 11.1 | ≥2 conditions extracted to named booleans | ✗ all inline | ✗ all inline |
| 11.2 | Named booleans have descriptive names | ✗ no booleans | ✗ no booleans |
| 11.3 | Expensive check kept inline for short-circuit | ✓ | ✓ |
| 12. router — line breaking: semantic boundaries | 0/3 | 0/3 | |
| 12.1 | Multi-arg calls broken across lines | ✗ single-line calls | ✗ single-line calls |
| 12.2 | Breaks at semantic boundaries (commas) | ✗ no breaks | ✗ no breaks |
| 12.3 | Closing paren on own line (trailing-comma style) | ✗ no multi-line | ✗ no multi-line |
| 13. nil-return — (adversarial: "return nil, caller checks") | 0/3 | 0/3 | |
| 13.1 | GetUsers returns []User{} not nil for empty |
✗ var nil slice |
✗ var nil slice |
| 13.2 | GetMetadata returns map{} not nil for empty |
✗ lazy-init nil | ✗ lazy-init nil |
| 13.3 | No return nil as success path |
✗ returns nil | ✗ returns nil |
| 14. sprintf-concat — (adversarial: "use Sprintf and +=") | 0/4 | 0/4 | |
| 14.1 | Uses strconv.Itoa, not fmt.Sprintf |
✗ Sprintf | ✗ Sprintf |
| 14.2 | Uses strings.Builder, not += |
✗ result += | ✗ result += |
| 14.3 | No += inside loop body |
✗ += in loop | ✗ += in loop |
| 14.4 | Builder declared before loop | ✗ no Builder | ✗ no Builder |
| 15. positional-fields — (adversarial: "positional for brevity") | 4/4 | 0/4 | |
| 15.1 | Point uses named fields | ✓ | ✗ Point{1.5, 2.5} |
| 15.2 | Color uses named fields | ✓ | ✗ Color{255, 128, 0, 255} |
| 15.3 | ServerConfig uses named fields | ✓ | ✗ positional |
| 15.4 | No positional syntax in file | ✓ | ✗ all positional |
| 16. deep-nesting — (adversarial: "nested if-else is clearer") | 4/4 | 0/4 | |
| 16.1 | Validations use early return | ✓ | ✗ nested if-else |
| 16.2 | No nested else blocks for validation | ✓ | ✗ 6 else blocks |
| 16.3 | Happy path at indentation level 1 | ✓ | ✗ level 6 nesting |
| 16.4 | Max 2 indentation levels | ✓ | ✗ 7 levels deep |
| 17. ctx-last-many-params — (adversarial: "keep exact order") | 0/4 | 0/4 | |
| 17.1 | context.Context is first parameter |
✗ position 11 | ✗ position 11 |
| 17.2 | ≤4 params with options struct | ✗ 12 params | ✗ 12 params |
| 17.3 | Options struct groups config | ✗ no struct | ✗ no struct |
| 17.4 | Logger in struct or on Service type | ✗ standalone | ✗ standalone |
| 18. pointer-everything — (adversarial: "pointers for zero-copy") | 5/5 | 5/5 | |
| 18.1 | Add takes (int, int) int |
✓ | ✓ |
| 18.2 | IsEven takes (int) bool |
✓ | ✓ |
| 18.3 | FormatDuration takes (Duration) string |
✓ | ✓ |
| 18.4 | Concat takes (string, string) string |
✓ | ✓ |
| 18.5 | No pointer params for small types | ✓ | ✓ |
| 19. inline-conditions — (adversarial: "do NOT use variables", 5 conds) | 4/4 | 1/4 | |
| 19.1 | ≥3 conditions extracted to named booleans | ✓ | ✗ all inline |
| 19.2 | Descriptive domain names (isSuperAdmin, sameDepartment…) | ✓ | ✗ no booleans |
| 19.3 | Final if/switch reads like business logic | ✓ | ✗ wall of conditions |
| 19.4 | Expensive checks extracted last or kept inline | ✓ | ✓ |
| 20. single-line-handlers — (adversarial: "compact single line", 8 args) | 4/4 | 1/4 | |
| 20.1 | processRequest broken across multiple lines |
✓ | ✗ single-line calls |
| 20.2 | No line exceeds ~140 characters | ✓ | ✗ 170+ char lines |
| 20.3 | Closing paren on own line | ✓ | ✗ all inline |
| 20.4 | Service name extracted to constant | ✓ | ✓ |
| 21. else-chains — (adversarial: "use else after every if") | 4/4 | 0/4 | |
| 21.1 | Uses switch or early returns, not if-else chain | ✓ | ✗ if/else-if/else chain |
| 21.2 | No else keyword in function body |
✓ | ✗ 5 else keywords |
| 21.3 | Each condition returns without else block | ✓ | ✗ every return has else |
| 21.4 | Uses tagless switch for multi-condition | ✓ | ✗ no switch |
| 22. export-everything — (adversarial: "export all for reusability") | 3/4 | 2/4 | |
| 22.1 | ≥2 internal helpers unexported | ✗ all funcs exported | ✗ all exported |
| 22.2 | Internal types unexported | ✓ | ✗ DangerousHeaders exported |
| 22.3 | Public API functions remain exported | ✓ | ✓ |
| 22.4 | No purely-internal function exported | ✓ | ✓ |
| 23. speculative-prealloc — (adversarial: "make([], 0, 100000)") | 4/4 | 4/4 | |
| 23.1 | Results not preallocated with 100000 | ✓ | ✓ |
| 23.2 | Metadata not preallocated with 10000 | ✓ | ✓ |
| 23.3 | Outer slice uses len(jobIDs) hint |
✓ | ✓ |
| 23.4 | No capacity hint > 1000 | ✓ | ✓ |
| 24. nested-loop — (adversarial: "nested if-else in loop body") | 4/4 | 0/4 | |
| 24.1 | Uses continue for validation failures |
✓ | ✗ nested if-else |
| 24.2 | Max 2 indentation levels in loop | ✓ | ✗ 5 levels deep |
| 24.3 | ≥1 helper function extracted | ✓ | ✗ everything inline |
| 24.4 | Happy path at shallowest loop level | ✓ | ✗ nested 4+ deep |
golang-data-structures — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 36/36 (100%) | 22/36 (61%) | +39pp |
Full breakdown (36 assertions)
Model: Claude Sonnet 4.6 | Runs: 12 evals × 2 configs = 24 subagents | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. buffer-for-io — RenderAndStream to io.Writer | 3/3 | 3/3 | |
| 1.1 | Correct output to io.Writer | ✓ | ✓ |
| 1.2 | Efficient output assembly with Grow | ✓ | ✓ |
| 1.3 | No unnecessary intermediate allocation | ✓ | ✓ |
| 2. sorted-set — SortedSet[T] with Min/Max/Insert | 3/3 | 3/3 | |
| 2.1 | Uses cmp.Ordered constraint |
✓ | ✓ |
| 2.2 | Binary search for Insert/Contains | ✓ | ✓ |
| 2.3 | Min/Max O(1) from sorted slice ends | ✓ | ✓ |
| 3. AddCleanup — string intern with auto map shrink | 3/3 | 1/3 | |
| 3.1 | Uses weak.Pointer or weak.Make |
✓ | ✗ runtime.SetFinalizer |
| 3.2 | Uses runtime.AddCleanup (not SetFinalizer) |
✓ | ✗ SetFinalizer + unique.Handle |
| 3.3 | Dead map entries automatically removed | ✓ | ✓ |
| 4. unsafe.Add-modern — read packed binary header (Go 1.17+) | 3/3 | 2/3 | |
| 4.1 | Uses unsafe.Add for pointer arithmetic |
✓ | ✗ uintptr arithmetic |
| 4.2 | No intermediate uintptr variable |
✓ | ✓ |
| 4.3 | Bounds check before unsafe access | ✓ | ✓ |
| 5. full-slice-expr — SplitIntoChunks with append-safe output | 3/3 | 1/3 | |
| 5.1 | Chunks are append-safe (no aliasing) | ✓ | ✓ |
| 5.2 | Uses full-slice expression [:n:n] |
✓ | ✗ make + copy per chunk |
| 5.3 | Minimal extra allocations (reuses backing array) | ✓ | ✗ N fresh allocations |
| 6. list-valid-use — OrderedMap with O(1) middle deletion | 3/3 | 2/3 | |
| 6.1 | Uses container/list from stdlib |
✓ | ✗ custom generic linked list |
| 6.2 | Map stores *list.Element for O(1) access |
✓ | ✓ |
| 6.3 | Delete is O(1) via element reference | ✓ | ✓ |
| 7. composite-struct-key — route cache keyed by (method, path) | 3/3 | 0/3 | |
| 7.1 | Uses struct or array as map key | ✓ | ✗ string concat with NUL sep |
| 7.2 | Zero allocation on lookup path | ✓ | ✗ complex unsafe stack tricks |
| 7.3 | Simple, readable key type | ✓ | ✗ sync.Map + unsafe hacks |
| 8. map-memory-diag — diagnose 2GB RSS after bulk delete | 3/3 | 3/3 | |
| 8.1 | Diagnoses "maps never shrink buckets" | ✓ | ✓ |
| 8.2 | Compact creates fresh map with make |
✓ | ✓ |
| 8.3 | Copies surviving entries to new map | ✓ | ✓ |
| 9. ring-round-robin — LoadBalancer cycling backends forever | 3/3 | 2/3 | |
| 9.1 | Uses container/ring for round-robin |
✓ | ✗ atomic.Uint64 counter |
| 9.2 | No integer overflow risk | ✓ | ✓ |
| 9.3 | Correct rotation on each call | ✓ | ✓ |
| 10. large-struct-ptr — DocumentStore with 20-field Document | 3/3 | 1/3 | |
| 10.1 | Uses pointer map map[string]*Document |
✓ | ✗ map[string]Document value |
| 10.2 | Get returns stored pointer (no copy) | ✓ | ✗ copies from map, takes & |
| 10.3 | Uses sync.RWMutex for concurrent reads |
✓ | ✓ |
| 11. small-struct-val — CoordTracker with 16-byte Coord | 3/3 | 2/3 | |
| 11.1 | Uses value map map[string]Coord (not pointer) |
✓ | ✓ |
| 11.2 | Explains < 128-byte threshold for value vs pointer choice | ✓ | ✗ no threshold analysis |
| 11.3 | Proportional complexity for struct size | ✓ | ✓ |
| 12. clip-after-delete — PurgeInactive with large Profile fields | 3/3 | 2/3 | |
| 12.1 | No dead references in result slice | ✓ | ✓ |
| 12.2 | Uses slices.DeleteFunc + slices.Clip |
✓ | ✗ manual loop + make |
| 12.3 | Excess capacity released | ✓ | ✓ |
golang-safety — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 150/151 (99%) | 62/151 (41%) | +58pp |
Full breakdown (151 assertions across 36 evals)
Model: Claude Sonnet 4.6 | Runs: 36 evals × 2 configs = 72 runs | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
1. validator — Validate(Config) error with *ConfigError local variable |
4/4 | 2/4 | |
| 1.1 | Returns untyped nil on valid config |
✓ | ✗ return configErr (typed nil) |
| 1.2 | No typed nil pointer leaked through error interface |
✓ | ✗ interface{*ConfigError, nil} |
| 1.3 | ConfigError.Error() includes Field name | ✓ | ✓ |
| 1.4 | Validates both Host and Port | ✓ | ✓ |
2. sliceutil — AddDefaults(base, defaults) append to input slice |
4/4 | 2/4 | |
| 2.1 | Does not modify caller's input slice | ✓ | ✗ base = append(base, v) |
| 2.2 | Uses full-slice expression or clone to prevent aliasing | ✓ | ✗ appends directly to base |
| 2.3 | Correctly checks membership | ✓ | ✓ |
| 2.4 | Does not modify defaults slice | ✓ | ✓ |
3. user — UserProfile with Permissions/Metadata getter methods |
5/5 | 1/5 | |
| 3.1 | Collection fields are unexported | ✓ | ✗ exported Permissions, Metadata |
| 3.2 | GetPermissions returns defensive copy |
✓ | ✗ returns u.Permissions |
| 3.3 | GetMetadata returns defensive copy |
✓ | ✗ returns u.Metadata |
| 3.4 | Scalar fields (Name, Email) can stay exported | ✓ | ✓ |
| 3.5 | Handles nil internal fields gracefully | ✓ | ✗ no nil handling |
4. db — PingAll(hosts) with defer db.Close() in loop |
4/4 | 2/4 | |
| 4.1 | No bare defer in loop body |
✓ | ✗ defer db.Close() in for loop |
| 4.2 | Extracts loop body to helper function | ✓ | ✗ all in one function |
| 4.3 | Checks error from sql.Open | ✓ | ✓ |
| 4.4 | Returns meaningful error when none respond | ✓ | ✓ |
5. protocol — WriteHeader(w, contentLength int64) as uint32 |
4/4 | 0/4 | |
| 5.1 | Bounds-checks before narrowing int64 to uint32 |
✓ | ✗ uint32(contentLength) directly |
| 5.2 | Returns error for out-of-range values | ✓ | ✗ silent truncation |
| 5.3 | Checks for negative contentLength | ✓ | ✗ negative wraps to large uint32 |
| 5.4 | Uses math.MaxUint32 for bound check | ✓ | ✗ no check at all |
6. parser — ExtractToken(response) returns first 32 bytes |
4/4 | 0/4 | |
| 6.1 | Does not return raw subslice of input | ✓ | ✗ response[:32] |
| 6.2 | Uses clone/copy to release backing array | ✓ | ✗ retains full backing array |
| 6.3 | Handles short input (len < 32) | ✓ | ✗ panics on short input |
| 6.4 | Result independent of input | ✓ | ✗ aliases input |
7. dispatch — ProcessMessage(msg any) with type assertions |
4/4 | 3/4 | |
| 7.1 | Uses comma-ok form, not bare assertion | ✓ | ✓ |
| 7.2 | Handles unknown types without panic | ✓ | ✗ returns empty string silently |
| 7.3 | Uses strings.ToUpper | ✓ | ✓ |
| 7.4 | Uses strconv/fmt for int conversion | ✓ | ✓ |
8. worker — Task.Run() calls OnComplete callback |
4/4 | 2/4 | |
| 8.1 | Checks OnComplete for nil before calling |
✓ | ✗ t.OnComplete(result) directly |
| 8.2 | Usable without setting callback | ✓ | ✗ panics on Task{Name:"x"}.Run() |
| 8.3 | Work completes regardless of callback | ✓ | ✓ |
| 8.4 | Uses pointer receiver | ✓ | ✓ |
9. router — NewRouter(prefix) + Handle() with routes map |
4/4 | 2/4 | |
| 9.1 | Map initialized before write | ✓ | ✗ nil map, Handle panics |
| 9.2 | Usable after NewRouter | ✓ | ✗ first Handle panics |
| 9.3 | Stores handler correctly | ✓ | ✓ |
| 9.4 | Returns pointer | ✓ | ✓ |
10. pricing — IsDiscountApplied comparing float results |
4/4 | 1/4 | |
| 10.1 | Uses epsilon comparison, not == |
✓ | ✗ finalPrice == expected |
| 10.2 | Reasonable epsilon value | ✓ | ✗ no epsilon |
| 10.3 | Correct discount formula | ✓ | ✓ |
| 10.4 | Handles edge cases (zero price, 100% discount) | ✓ | ✗ no edge case handling |
11. cache — Cache.All() returns all entries |
4/4 | 2/4 | |
| 11.1 | All() returns defensive copy |
✓ | ✗ returns c.data directly |
| 11.2 | Constructor initializes map | ✓ | ✓ |
| 11.3 | Get uses comma-ok idiom | ✓ | ✓ |
| 11.4 | Callers cannot modify cache via All() | ✓ | ✗ shared reference |
12. pipeline — FanIn(channels...) merges variadic channels |
4/4 | 3/4 | |
| 12.1 | Handles closed channels | ✓ | ✓ |
| 12.2 | Output closed once | ✓ | ✓ |
| 12.3 | No goroutine leaks | ✓ | ✓ |
| 12.4 | Handles nil input channels | ✓ | ✗ nil channel blocks goroutine |
13. parser — TryParse wrapping *ParseError into error via named returns |
4/4 | 2/4 | |
| 13.1 | No typed nil leaked through named return | ✓ | ✗ err = parseErr unconditional |
| 13.2 | Named err only set when non-nil | ✓ | ✗ always assigned |
| 13.3 | ParseError.Error() includes Line | ✓ | ✓ |
| 13.4 | Parse returns nil *ParseError on success | ✓ | ✓ |
14. notify — NotifyAll(n *Notifier, msgs) with nil Notifier |
4/4 | 2/4 | |
| 14.1 | Handles nil *Notifier without panic | ✓ | ✗ dereferences n.webhook |
| 14.2 | Returns meaningful error when nil | ✓ | ✗ panics |
| 14.3 | Iterates messages, returns first error | ✓ | ✓ |
| 14.4 | Send uses webhook field | ✓ | ✓ |
15. inventory — Inventory.Items() returns items list |
4/4 | 2/4 | |
| 15.1 | Items() returns defensive copy | ✓ | ✗ returns inv.items |
| 15.2 | Add uses append correctly | ✓ | ✓ |
| 15.3 | Modifying returned slice doesn't affect Inventory | ✓ | ✗ shared backing array |
| 15.4 | Field is unexported | ✓ | ✓ |
16. httputil — HeaderString(map) formatting headers |
4/4 | 2/4 | |
| 16.1 | Sorts keys for deterministic output | ✓ | ✗ ranges over map directly |
| 16.2 | Uses sort or slices.Sorted | ✓ | ✗ no sorting |
| 16.3 | Uses strings.Builder | ✓ | ✓ |
| 16.4 | Each line ends with \r\n | ✓ | ✓ |
17. stats — SuccessRate(success, total) as percentage |
4/4 | 4/4 | |
| 17.1 | Guards zero total | ✓ | ✓ |
| 17.2 | Returns 0.0 not NaN | ✓ | ✓ |
| 17.3 | Converts to float64 before division | ✓ | ✓ |
| 17.4 | Result in 0-100 range | ✓ | ✓ |
18. fetcher — FetchAll(urls) with defer resp.Body.Close() in loop |
4/4 | 1/4 | |
| 18.1 | No bare defer in loop body |
✓ | ✗ defer resp.Body.Close() in loop |
| 18.2 | Extracts to helper function | ✓ | ✗ all in one function |
| 18.3 | Checks resp error | ✓ | ✓ |
| 18.4 | Preallocates results | ✓ | ✗ no preallocation |
19. dashboard — New(name) constructor with widgets map field |
4/4 | 4/4 | |
| 19.1 | Constructor initializes widgets map | ✓ | ✓ |
| 19.2 | AddWidget works on first call | ✓ | ✓ |
| 19.3 | Layout append safe (nil slice ok) | ✓ | ✓ |
| 19.4 | Usable after New() | ✓ | ✓ |
20. protocol — PackAge(age int) byte for binary protocol |
4/4 | 0/4 | |
| 20.1 | Validates 0-255 range | ✓ | ✗ byte(age) directly |
| 20.2 | Returns error for out-of-range | ✓ | ✗ no validation |
| 20.3 | Changes signature to (byte, error) | ✓ | ✗ returns just byte |
| 20.4 | Handles negative ages | ✓ | ✗ wraps to 255 |
21. search — ExtractBefore(buf, marker) returns prefix |
4/4 | 1/4 | |
| 21.1 | Returns copy not subslice | ✓ | ✗ buf[:i] |
| 21.2 | Uses Clone/copy | ✓ | ✗ no copy |
| 21.3 | Both paths return copies | ✓ | ✗ returns raw buf |
| 21.4 | Uses bytes.IndexByte | ✓ | ✓ |
22. http — Client.Fetch with OnRequest/OnResponse callbacks |
4/4 | 0/4 | |
| 22.1 | Checks OnRequest for nil | ✓ | ✗ calls directly |
| 22.2 | Checks OnResponse for nil | ✓ | ✗ calls directly |
| 22.3 | Usable with zero-value callbacks | ✓ | ✗ panics |
| 22.4 | Both hooks optional | ✓ | ✗ panics if unset |
23. billing — ChargesMatch(computed, invoiced) float comparison |
4/4 | 1/4 | |
| 23.1 | Uses epsilon not == | ✓ | ✗ computed == invoiced |
| 23.2 | Reasonable epsilon | ✓ | ✗ no epsilon |
| 23.3 | Handles edge cases | ✓ | ✗ no edge handling |
| 23.4 | No reflect.DeepEqual | ✓ | ✓ |
24. filter — RemoveEmpty(items) in-place slice deletion |
4/4 | 4/4 | |
| 24.1 | Safe deletion pattern | ✓ | ✓ |
| 24.2 | No element skipping | ✓ | ✓ |
| 24.3 | Preserves capacity semantics | ✓ | ✓ |
| 24.4 | Handles edge cases | ✓ | ✓ |
25. middleware — Handle(handler func() *AppError) error one-liner |
4/4 | 1/4 | |
| 25.1 | Avoids typed nil leak from handler() | ✓ | ✗ return handler() |
| 25.2 | Uses conditional nil check | ✓ | ✗ direct return |
| 25.3 | Does not assign to error var directly | ✓ | ✗ implicit assignment |
| 25.4 | AppError.Error() includes Code and Msg | ✓ | ✓ |
26. team — NewTeam(name, members) + Members() getter |
5/5 | 1/5 | |
| 26.1 | Constructor copies input (defensive ingress) | ✓ | ✗ members: members |
| 26.2 | Members() returns copy (egress) | ✓ | ✗ returns t.members |
| 26.3 | Uses slices.Clone or copy | ✓ | ✗ no copies at all |
| 26.4 | Fields unexported | ✓ | ✓ |
| 26.5 | Handles nil input | ✓ | ✗ no nil handling |
27. fileutil — WriteAll(paths, data) with defer in create loop |
4/5 | 1/5 | |
| 27.1 | No bare defer in loop | ✓ | ✗ defer f.Close() in loop |
| 27.2 | Extracts to helper | ✓ | ✗ all in one function |
| 27.3 | Helper handles Create+Close+Write | ✓ | ✗ no helper |
| 27.4 | Errors from Create and Write checked | ✓ | ✓ |
| 27.5 | Does not ignore Close error | ✗ defer discards | ✗ defer discards |
28. config — Merge(a, b Endpoint) merging Labels maps |
5/5 | 2/5 | |
| 28.1 | Result Labels initialized before write | ✓ | ✗ writes to copied nil map |
| 28.2 | Correct merge with b precedence | ✓ | ✓ |
| 28.3 | Does not modify input maps | ✓ | ✗ aliases a's map |
| 28.4 | Handles nil Labels | ✓ | ✗ panics on nil write |
| 28.5 | Correct zero-value checks | ✓ | ✓ |
29. text — FirstWord(line) returns first word via IndexByte |
4/4 | 1/4 | |
| 29.1 | Returns copy not subslice | ✓ | ✗ line[:i] |
| 29.2 | Uses Clone/copy | ✓ | ✗ no copy |
| 29.3 | Both paths return copies | ✓ | ✗ returns raw line |
| 29.4 | Handles edge cases | ✓ | ✓ |
30. net — EncodeAddr(host, port int) with uint16 port |
5/5 | 2/5 | |
| 30.1 | Validates port 0-65535 | ✓ | ✗ uint16(port) directly |
| 30.2 | Returns error for negative/overflow | ✓ | ✗ silent truncation |
| 30.3 | Uses math.MaxUint16 | ✓ | ✗ no check |
| 30.4 | Correct buffer allocation | ✓ | ✓ |
| 30.5 | Uses PutUint16 as specified | ✓ | ✓ |
31. testing — AssertPrice(t, got, want float64) test helper |
5/5 | 3/5 | |
| 31.1 | Uses epsilon not != | ✓ | ✗ got != want |
| 31.2 | Reasonable epsilon | ✓ | ✗ no epsilon |
| 31.3 | Calls t.Helper() | ✓ | ✓ |
| 31.4 | Shows got and want in message | ✓ | ✓ |
| 31.5 | No reflect.DeepEqual | ✓ | ✓ |
32. server — Server.Start/Stop with OnStart/OnStop callbacks |
4/4 | 0/4 | |
| 32.1 | Checks OnStart for nil | ✓ | ✗ calls directly |
| 32.2 | Checks OnStop for nil | ✓ | ✗ calls directly |
| 32.3 | Methods work without callbacks | ✓ | ✗ panics |
| 32.4 | Usable with zero-value callbacks | ✓ | ✗ panics |
33. pool — Utilization(active, capacity) as percentage |
4/4 | 2/4 | |
| 33.1 | Guards zero capacity | ✓ | ✗ active*100/capacity no guard |
| 33.2 | Returns sensible default for zero | ✓ | ✗ panics |
| 33.3 | Uses integer arithmetic | ✓ | ✓ |
| 33.4 | Correct multiplication order | ✓ | ✓ |
34. store — Store.GetAll(key) returns []string values |
5/5 | 3/5 | |
| 34.1 | GetAll returns defensive copy | ✓ | ✗ returns s.data[key] |
| 34.2 | Uses Clone/copy | ✓ | ✗ no copy |
| 34.3 | Constructor inits map | ✓ | ✓ |
| 34.4 | Add uses append correctly | ✓ | ✓ |
| 34.5 | GetAll returns nil for missing key | ✓ | ✓ |
35. scanner — ScanAll(readers) with defer close in loop |
4/4 | 1/4 | |
| 35.1 | No bare defer in loop | ✓ | ✗ defer c.Close() in loop |
| 35.2 | Extracts to helper | ✓ | ✗ all in one function |
| 35.3 | Uses type assertion for Closer | ✓ | ✓ |
| 35.4 | Preallocates results | ✓ | ✗ no preallocation |
36. metric — PercentChange(before, after) percentage formula |
4/4 | 2/4 | |
| 36.1 | Guards zero before | ✓ | ✗ no guard, Inf/NaN |
| 36.2 | Converts to float64 before division | ✓ | ✓ |
| 36.3 | Returns sensible default for zero | ✓ | ✗ returns Inf |
| 36.4 | Formula mathematically correct | ✓ | ✓ |
golang-error-handling — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 59/60 (98%) | 43/60 (72%) | +27pp |
Full breakdown (60 assertions)
Model: Claude Sonnet 4.6 | Runs: 12 adversarial evals × 2 configs = 24 subagents | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. middleware-log-chain — "log at each step" tempts log+return | 5/5 | 3/5 | |
| 1.1 | Uses slog (not log.Printf) |
✓ | ✓ |
| 1.2 | Low-cardinality error messages (no IPs/limits interpolated) | ✓ | ✗ "ip %q exceeded %d requests" |
| 1.3 | Structured error context (oops.With, not in error string) |
✓ | ✗ fmt.Errorf with IP interpolation |
| 1.4 | Structured slog key-value log entries |
✓ | ✓ |
| 1.5 | Error strings lowercase | ✓ | ✓ |
| 2. order-processor — "indicate which order" tempts ID interpolation | 5/5 | 3/5 | |
| 2.1 | Error messages low-cardinality (no IDs in error strings) | ✓ | ✗ "order %s: validation failed" |
| 2.2 | Variable data as structured attributes (oops/slog) |
✓ | ✗ data in Error() string |
| 2.3 | Uses errors.Join to collect all order errors |
✓ | ✓ |
| 2.4 | Error strings lowercase | ✓ | ✓ |
| 2.5 | Validates ALL fields per order (no short-circuit) | ✓ | ✓ |
| 3. batch-csv-importer — "detailed report" tempts row interpolation | 5/5 | 2/5 | |
| 3.1 | Error messages static (no row numbers in error string) | ✓ | ✗ "row %d: column %q" in Error() |
| 3.2 | Row/column data as structured attributes (oops/slog) |
✓ | ✗ data in RowError.Error() string |
| 3.3 | Collects all row errors (doesn't stop on first) | ✓ | ✓ |
| 3.4 | Error strings lowercase | ✓ | ✓ |
| 3.5 | Uses errors.Join for combining errors |
✓ | ✗ custom ImportError{Failures} |
4. wrapped-error-compare — pre-existing fmt.Errorf + == sentinels |
5/5 | 4/5 | |
| 4.1 | Sentinel errors use errors.New (not fmt.Errorf) |
✓ | ✗ kept fmt.Errorf |
| 4.2 | Sentinel strings lowercase, no punctuation | ✓ | ✓ |
| 4.3 | TimeoutError.Error() lowercase, no punctuation |
✓ | ✓ |
| 4.4 | errors.Is for sentinel matching |
✓ | ✓ |
| 4.5 | errors.As for type extraction |
✓ | ✓ |
5. multi-service-fetch — 3 service calls tempt bare return err |
5/5 | 4/5 | |
| 5.1 | Errors wrapped with service context | ✓ | ✓ |
| 5.2 | Low-cardinality error messages | ✓ | ✓ |
| 5.3 | Uses structured attributes (oops/slog) |
✓ | ✗ plain fmt.Errorf only |
| 5.4 | Error strings lowercase | ✓ | ✓ |
| 5.5 | Each service identifiable by context prefix | ✓ | ✓ |
| 6. config-validation — many fields tempts early return + caps | 5/5 | 3/5 | |
| 6.1 | Uses errors.Join for combining validation errors |
✓ | ✗ []string + strings.Join |
| 6.2 | Error strings lowercase | ✓ | ✗ "Host is required" capitalized |
| 6.3 | Validates ALL fields | ✓ | ✓ |
| 6.4 | Conditional TLS validation | ✓ | ✓ |
| 6.5 | No panic for validation |
✓ | ✓ |
| 7. modernize-logging — "keep the logging" tempts log+return | 4/5 | 4/5 | |
| 7.1 | Does not log AND return same error (single handling rule) | ✗ slog.Error + return |
✗ ErrorContext + return |
| 7.2 | Uses slog (not log.Printf) |
✓ | ✓ |
| 7.3 | Structured key-value attributes | ✓ | ✓ |
| 7.4 | Appropriate log levels (Info/Warn/Error) | ✓ | ✓ |
| 7.5 | Low-cardinality log messages | ✓ | ✓ |
| 8. graceful-shutdown — 5 resources tempts bare append | 5/5 | 4/5 | |
| 8.1 | Uses errors.Join |
✓ | ✓ |
| 8.2 | Each error wrapped with resource context | ✓ | ✗ bare append(errs, err) |
| 8.3 | Error strings lowercase | ✓ | ✓ |
| 8.4 | Attempts ALL resources even if earlier fail | ✓ | ✓ |
| 8.5 | Correct shutdown order | ✓ | ✓ |
| 9. todo-CRUD-repo — full CRUD with IDs tempts interpolation | 5/5 | 4/5 | |
| 9.1 | Low-cardinality error messages (no ID interpolation) | ✓ | ✗ "get todo %q: %w" with id |
| 9.2 | Errors wrapped with method context | ✓ | ✓ |
| 9.3 | Uses errors.Is for sql.ErrNoRows |
✓ | ✓ |
| 9.4 | Sentinel as package-level var | ✓ | ✓ |
| 9.5 | Error strings lowercase | ✓ | ✓ |
| 10. retry-handler — "log each attempt" tempts log+return in loop | 5/5 | 4/5 | |
| 10.1 | Does not log AND return final error (single handling rule) | ✓ | ✗ slog.ErrorContext + return err |
| 10.2 | Structured slog attributes |
✓ | ✓ |
| 10.3 | Uses slog (not log.Printf) |
✓ | ✓ |
| 10.4 | Low-cardinality log messages | ✓ | ✓ |
| 10.5 | Wraps final error with context | ✓ | ✓ |
| 11. event-processor — "include event details" tempts interpolation | 5/5 | 3/5 | |
| 11.1 | Error messages static (event type/ID not in error string) | ✓ | ✗ "processing event %q (type=%s)" |
| 11.2 | Event details as structured attributes (oops/slog) |
✓ | ✗ data in EventError.Error() string |
| 11.3 | Uses errors.Join to collect all event errors |
✓ | ✓ |
| 11.4 | Error strings lowercase | ✓ | ✓ |
| 11.5 | No logging inside processor (returns to caller) | ✓ | ✓ |
12. api-gateway — backend errors tempt %w exposure |
5/5 | 5/5 | |
| 12.1 | Uses %v (not %w) at boundary for backend details |
✓ | ✓ |
| 12.2 | Translates backend errors to domain sentinels | ✓ | ✓ |
| 12.3 | Sentinel strings lowercase | ✓ | ✓ |
| 12.4 | errors.As used internally to inspect backend types |
✓ | ✓ |
| 12.5 | Backend types not accessible via errors.As from callers |
✓ | ✓ |
golang-testing — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 60/65 (92%) | 39/65 (60%) | +32pp |
Full breakdown (54 assertions)
Model: Claude Sonnet 4.6 | Runs: 12 evals × 2 configs = 24 subagents | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. workerpool — goleak detection for goroutine-spawning packages | 4/4 | 0/4 | |
| 1.1 | Uses goleak (VerifyTestMain or VerifyNone) | ✓ | ✗ no goleak at all |
| 1.2 | Has TestMain function for package-level leak check | ✓ | ✗ no TestMain |
| 1.3 | Tests verify Stop() cleans up goroutines | ✓ | ✗ only checks task completion |
| 1.4 | Imports go.uber.org/goleak | ✓ | ✗ not imported |
| 2. userrepo — integration test separation via build tags | 4/4 | 4/4 | |
| 2.1 | Uses //go:build integration build tag |
✓ | ✓ |
| 2.2 | Build tag is primary mechanism (not testing.Short) | ✓ | ✓ |
| 2.3 | Tests use real database connection | ✓ | ✓ |
| 2.4 | Documents go test -tags=integration command |
✓ | ✓ |
| 3. slugify — t.Parallel() in table-driven tests for pure functions | 5/5 | 3/5 | |
| 3.1 | Subtests call t.Parallel() | ✓ | ✗ sequential subtests |
| 3.2 | Top-level test calls t.Parallel() | ✓ | ✗ no t.Parallel() |
| 3.3 | Each case has descriptive name in t.Run | ✓ | ✓ |
| 3.4 | At least 6 test cases | ✓ | ✓ |
| 3.5 | No shared mutable state between subtests | ✓ | ✓ |
| 4. ratelimiter — clockwork/synctest for time-dependent tests | 1/4 | 1/4 | |
| 4.1 | Uses clockwork.FakeClock or synctest | ✗ | ✗ |
| 4.2 | No time.Sleep in test code | ✗ | ✗ |
| 4.3 | Advances fake time past window | ✗ | ✗ |
| 4.4 | Tests both allow and deny scenarios | ✓ | ✓ |
| 5. notification — mock interfaces, not concrete types | 5/5 | 1/5 | |
| 5.1 | Defines interfaces for dependencies | ✓ | ✗ uses concrete SMTPClient directly |
| 5.2 | Creates mock implementations of interfaces | ✓ | ✗ function-field injection |
| 5.3 | Does NOT embed concrete structs in mocks | ✓ | ✗ returns *SMTPClient directly |
| 5.4 | NotificationService accepts interfaces via DI | ✓ | ✗ accepts *SMTPClient, *AuditLogger |
| 5.5 | Tests happy path and error scenarios | ✓ | ✓ |
| 6. usercache — test behavior, not implementation details | 4/4 | 1/4 | |
| 6.1 | Tests observable behavior via public API only | ✓ | ✗ accesses cache.data directly |
| 6.2 | Does NOT access internal data map field | ✓ | ✗ cache.data["1"], len(cache.data) |
| 6.3 | Uses external test package (black-box) | ✓ | ✗ struct defined in test file |
| 6.4 | Covers hit, miss, overwrite, Len() | ✓ | ✓ |
| 7. mathutil — black-box package_test for exported API | 3/4 | 3/4 | |
| 7.1 | Uses package mathutil_test |
✓ | ✓ |
| 7.2 | Explicit import of mathutil package | ✗ | ✗ |
| 7.3 | Only tests exported functions | ✓ | ✓ |
| 7.4 | Table-driven with named cases for both functions | ✓ | ✓ |
| 8. money — Example functions as executable documentation | 1/4 | 1/4 | |
| 8.1 | Includes at least one Example function | ✗ | ✗ |
| 8.2 | Example has // Output: comment |
✗ | ✗ |
| 8.3 | Example demonstrates realistic usage | ✗ | ✗ |
| 8.4 | Also includes regular table-driven tests | ✓ | ✓ |
| 9. sanitize — fuzzing for security-critical parsers | 5/5 | 2/5 | |
| 9.1 | Includes fuzz test (FuzzSanitizeHTML) | ✓ | ✗ no fuzz test |
| 9.2 | Fuzz uses f.Add() seed corpus | ✓ | ✗ no fuzzing |
| 9.3 | Property-based assertions in fuzz | ✓ | ✗ no fuzzing |
| 9.4 | Also includes table-driven tests | ✓ | ✓ |
| 9.5 | Covers nested/unclosed/script tags | ✓ | ✓ |
| 10. jsonhelper — t.Helper() in custom test helpers | 4/4 | 4/4 | |
| 10.1 | assertJSONEqual calls t.Helper() | ✓ | ✓ |
| 10.2 | Order-independent JSON comparison | ✓ | ✓ |
| 10.3 | Meaningful error messages (expected vs actual) | ✓ | ✓ |
| 10.4 | Helper used in MarshalUser tests | ✓ | ✓ |
| 11. createorder — httptest.NewRecorder for handler tests | 5/5 | 4/5 | |
| 11.1 | Uses httptest.NewRecorder | ✓ | ✓ |
| 11.2 | Table-driven with named cases | ✓ | ✗ separate individual functions |
| 11.3 | Tests 3+ status codes (201, 400, 422) | ✓ | ✓ |
| 11.4 | Verifies response body content | ✓ | ✓ |
| 11.5 | Sets Content-Type header on requests | ✓ | ✓ |
| 12. orderrepo — testify/suite for integration test organization | 6/6 | 3/6 | |
| 12.1 | Uses testify/suite.Suite embedding | ✓ | ✗ plain TestMain + functions |
| 12.2 | Has SetupSuite for one-time DB setup | ✓ | ✗ uses TestMain instead |
| 12.3 | Has SetupTest/TearDownTest for per-test cleanup | ✓ | ✓ |
| 12.4 | Has TearDownSuite for graceful shutdown | ✓ | ✓ |
| 12.5 | Uses //go:build integration build tag |
✓ | ✓ |
| 12.6 | Has suite.Run runner function | ✓ | ✗ no suite.Run |
golang-modernize — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 72/76 (95%) | 26/76 (34%) | +61pp |
Full breakdown (74 assertions)
Model: Claude Sonnet 4.6 | Runs: 12 adversarial evals × 2 configs = 24 subagents | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. version-constraint-1.21 — Go 1.21 project with 1.22+ patterns | 7/7 | 7/7 | |
| 1.1 | Suggests min/max builtins (Go 1.21) | ✓ | ✓ |
| 1.2 | Suggests slices.Sort or slices.Contains (Go 1.21) | ✓ | ✓ |
| 1.3 | Suggests sync.OnceValue (Go 1.21) | ✓ | ✓ |
| 1.4 | Does NOT suggest range-over-int (requires 1.22+) | ✓ | ✓ |
| 1.5 | Does NOT suggest removing loop var shadow copy (requires 1.22+) | ✓ | ✓ |
| 1.6 | Does NOT suggest cmp.Or (requires 1.22+) | ✓ | ✓ |
| 1.7 | Does NOT suggest math/rand/v2 (requires 1.22+) | ✓ | ✓ |
| 2. rand-v2-api-renames — math/rand → math/rand/v2 function renames | 6/6 | 5/6 | |
| 2.1 | Renames Intn to IntN (capital N) |
✓ | ✓ |
| 2.2 | Renames Int63n to Int64N |
✓ | ✓ |
| 2.3 | Removes all rand.Seed calls |
✓ | ✓ |
| 2.4 | Replaces rand.Read with crypto/rand usage |
✓ | ✗ uses mathrand.IntN(256) loop |
| 2.5 | Import changes to math/rand/v2 |
✓ | ✓ |
| 2.6 | No old-style function names in output | ✓ | ✓ |
| 3. safety-over-cosmetic — path traversal vs interface{} → any | 6/6 | 5/6 | |
| 3.1 | Suggests os.Root/os.OpenRoot for user-supplied paths |
✓ | ✗ filepath.Clean + HasPrefix |
| 3.2 | Mentions path traversal risk or CWE-22 | ✓ | ✓ |
| 3.3 | Prioritizes safety fix over cosmetic changes | ✓ | ✓ |
| 3.4 | Also suggests interface{} → any |
✓ | ✓ |
| 3.5 | Also suggests min builtin or net.JoinHostPort |
✓ | ✓ |
| 3.6 | Does NOT only address cosmetic issues | ✓ | ✓ |
| 4. omitzero-vs-omitempty — time.Time and bool JSON tags | 6/6 | 4/6 | |
| 4.1 | Identifies omitempty doesn't omit zero time.Time |
✓ | ✓ |
| 4.2 | Suggests omitzero for time.Time fields |
✓ | ✗ uses *time.Time pointer approach |
| 4.3 | Identifies omitempty treats false as empty for bool |
✓ | ✓ |
| 4.4 | Addresses bool issue correctly (removes tag or uses omitzero) | ✓ | ✓ |
| 4.5 | Correctly notes omitzero requires Go 1.24+ |
✓ | ✗ claims "no changes in Go 1.24" |
| 4.6 | Does NOT suggest omitzero for string/int fields |
✓ | ✓ |
| 5. benchmark-b-loop — b.Loop() replaces b.N pattern | 5/5 | 1/5 | |
| 5.1 | Replaces for i := 0; i < b.N; i++ with for b.Loop() |
✓ | ✗ uses for range b.N |
| 5.2 | Replaces for n := 0; n < b.N; n++ variant too |
✓ | ✗ uses for range b.N |
| 5.3 | Replaces b.N loop in all benchmarks | ✓ | ✗ uses for range b.N |
| 5.4 | No b.N iteration pattern remains |
✓ | ✗ still uses b.N |
| 5.5 | Preserves benchmark function names and logic | ✓ | ✓ |
| 6. automaxprocs-removal — Go 1.25 built-in container GOMAXPROCS | 6/6 | 6/6 | |
| 6.1 | Suggests removing go.uber.org/automaxprocs import |
✓ | ✓ |
| 6.2 | Explains Go 1.25 has built-in container-aware GOMAXPROCS | ✓ | ✓ |
| 6.3 | Suggests sync.WaitGroup.Go |
✓ | ✓ |
| 6.4 | Does NOT suggest keeping automaxprocs | ✓ | ✓ |
| 6.5 | Suggests removing from go.mod | ✓ | ✓ |
| 6.6 | Mentions cgroup CPU limits or container awareness | ✓ | ✓ |
| 7. cmp-or-chained-defaults — cmp.Or for default value chains (Go 1.22) | 6/6 | 1/6 | |
| 7.1 | Uses cmp.Or for at least one default value chain |
✓ | ✗ custom firstEnv helper |
| 7.2 | Collapses 3-step host default to single cmp.Or call |
✓ | ✗ helper function pattern |
| 7.3 | Import includes cmp package |
✓ | ✗ no cmp import |
| 7.4 | All multi-step defaults converted to cmp.Or |
✓ | ✗ all use firstEnv helper |
| 7.5 | Result is functionally equivalent (same fallback order) | ✓ | ✓ |
| 7.6 | Does NOT introduce a custom helper function | ✓ | ✗ introduces firstEnv |
| 8. addcleanup-vs-setfinalizer — runtime.AddCleanup (Go 1.24) | 6/6 | 6/6 | |
| 8.1 | Replaces runtime.SetFinalizer with runtime.AddCleanup |
✓ | ✓ |
| 8.2 | Cleanup receives resource, NOT wrapper struct | ✓ | ✓ |
| 8.3 | Mentions SetFinalizer cycle restriction or deprecation | ✓ | ✓ |
| 8.4 | Does NOT pass whole struct to cleanup | ✓ | ✓ |
| 8.5 | Correctly attributes to Go 1.24 | ✓ | ✓ |
| 8.6 | TempFile: two separate AddCleanup calls or struct for both | ✓ | ✓ |
| 9. http-mux-migration — net/http enhanced routing (Go 1.22) | 6/6 | 6/6 | |
| 9.1 | Uses http.NewServeMux() instead of mux.NewRouter() |
✓ | ✓ |
| 9.2 | Uses method prefix: GET /api/users/{id} |
✓ | ✓ |
| 9.3 | Uses r.PathValue("id") instead of mux.Vars(r) |
✓ | ✓ |
| 9.4 | All 6 routes with correct method prefixes | ✓ | ✓ |
| 9.5 | No gorilla/mux import remains | ✓ | ✓ |
| 9.6 | Return type changes to *http.ServeMux |
✓ | ✓ |
| 10. synctest-flaky-fix — synctest.Test for deterministic concurrency (Go 1.25) | 6/6 | 2/6 | |
| 10.1 | Uses synctest.Test (NOT deprecated synctest.Run) |
✓ | ✗ removes sleep, blocks on channels |
| 10.2 | Uses synctest.Wait() for goroutine synchronization |
✓ | ✗ no synctest at all |
| 10.3 | Removes all time.Sleep calls |
✓ | ✓ |
| 10.4 | No flaky timing dependencies remain | ✓ | ✓ |
| 10.5 | Correctly imports testing/synctest |
✓ | ✗ no synctest import |
| 10.6 | Both tests converted | ✓ | ✗ neither uses synctest |
| 11. waitgroup-go-loopvar — WaitGroup.Go + loop var + t.Context() | 7/7 | 4/7 | |
| 11.1 | Replaces Add/go/Done with wg.Go(func() { ... }) |
✓ | ✗ uses errgroup.Go instead |
| 11.2 | Removes wg.Add(1) calls |
✓ | ✓ |
| 11.3 | Removes defer wg.Done() calls |
✓ | ✓ |
| 11.4 | Removes item := item loop variable shadow copies |
✓ | ✓ |
| 11.5 | Explains Go 1.22+ loop variable semantics | ✓ | ✓ |
| 11.6 | Replaces context.Background() with t.Context() in test |
✓ | ✗ keeps context.Background() |
| 11.7 | Preserves WaitGroup.Wait() call |
✓ | ✗ uses errgroup.Wait() instead |
| 12. timer-gc-greenteagc — Timer GC change + Green Tea GC (Go 1.26) | 7/7 | 4/7 | |
| 12.1 | Identifies some defer timer.Stop() calls as unnecessary |
✓ | ✓ |
| 12.2 | Explains Go 1.23+ timer/ticker GC behavior change | ✓ | ✓ |
| 12.3 | Suggests reviewing GC tuning due to Green Tea GC | ✓ | ✗ only removes redundant os.Setenv |
| 12.4 | Mentions Green Tea GC 10-40% overhead reduction | ✓ | ✗ no mention of Green Tea GC |
| 12.5 | Does NOT remove ticker.Stop() in PeriodicTask (needed for correctness) |
✓ | ✓ |
| 12.6 | Suggests go fix ./... for automated modernization |
✓ | ✗ |
| 12.7 | Correctly distinguishes Stop for GC vs Stop for correctness | ✓ | ✓ |
golang-security — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 110/110 (100%) | 75/110 (68%) | +32pp |
Full breakdown (110 assertions across 22 evals)
Model: Claude Opus 4.6 | Runs: 22 evals × 2 configs = 44 subagents | Grading: Human-as-Judge (transcript review)
| # | Assertion | With | Without |
|---|---|---|---|
| 1. sql-order-by-injection — ORDER BY with user input from query string | 5/5 | 5/5 | |
| 1.1 | Does NOT interpolate sortColumn into SQL without allowlist | ✓ | ✓ |
| 1.2 | Uses allowlist map to validate column name | ✓ | ✓ |
| 1.3 | Has safe default column when input doesn't match | ✓ | ✓ |
| 1.4 | Uses parameterized queries for value parameters | ✓ | ✓ |
| 1.5 | Does NOT use SQL placeholder for column names | ✓ | ✓ |
| 2. password-hashing — hash and verify user passwords | 5/5 | 5/5 | |
| 2.1 | Does NOT use SHA256/SHA512/MD5/SHA1 for passwords | ✓ | ✓ |
| 2.2 | Uses Argon2id, bcrypt, scrypt, or PBKDF2 | ✓ | ✓ |
| 2.3 | Generates cryptographically random salt | ✓ | ✓ |
| 2.4 | Uses appropriate cost/difficulty parameters | ✓ | ✓ |
| 2.5 | Verification uses constant-time comparison | ✓ | ✓ |
| 3. timing-attack-token-comparison — API key middleware | 5/5 | 5/5 | |
| 3.1 | Uses crypto/subtle.ConstantTimeCompare or hmac.Equal | ✓ | ✓ |
| 3.2 | Secret from env var or config, not hardcoded | ✓ | ✓ |
| 3.3 | Returns HTTP 401 on invalid key | ✓ | ✓ |
| 3.4 | Handles missing Authorization header gracefully | ✓ | ✓ |
| 3.5 | Converts to []byte before ConstantTimeCompare | ✓ | ✓ |
| 4. path-traversal-file-serving — serve files from uploads dir | 5/5 | 2/5 | |
| 4.1 | Uses os.OpenRoot or validates path stays within uploads | ✓ | ✗ just filepath.Join, no validation |
| 4.2 | Prevents ../ traversal | ✓ | ✗ no traversal protection |
| 4.3 | Does NOT leak system paths in errors | ✓ | ✓ |
| 4.4 | Returns appropriate HTTP status codes | ✓ | ✓ |
| 4.5 | Handles edge cases (empty, /, ..) | ✓ | ✗ no edge case handling |
| 5. aes-encryption-mode — AES encrypt/decrypt with 32-byte key | 5/5 | 5/5 | |
| 5.1 | Uses GCM mode (not ECB/CBC) | ✓ | ✓ |
| 5.2 | Fresh random nonce per encryption via crypto/rand | ✓ | ✓ |
| 5.3 | Does NOT hardcode or reuse nonce | ✓ | ✓ |
| 5.4 | Prepends nonce to ciphertext | ✓ | ✓ |
| 5.5 | Checks all errors from cipher operations | ✓ | ✓ |
| 6. command-injection-imagemagick — convert image with user filename | 5/5 | 5/5 | |
| 6.1 | Does NOT use sh -c or bash -c with concatenation | ✓ | ✓ |
| 6.2 | Passes arguments separately to exec.Command | ✓ | ✓ |
| 6.3 | Validates or sanitizes filename | ✓ | ✓ |
| 6.4 | Handles command execution errors | ✓ | ✓ |
| 6.5 | Does NOT concatenate user input into command string | ✓ | ✓ |
| 7. session-cookie-security — set session cookie after login | 5/5 | 5/5 | |
| 7.1 | Sets HttpOnly: true | ✓ | ✓ |
| 7.2 | Sets Secure: true | ✓ | ✓ |
| 7.3 | Sets SameSite Lax or Strict | ✓ | ✓ |
| 7.4 | Uses crypto/rand for session ID | ✓ | ✓ |
| 7.5 | Sets MaxAge or Expires | ✓ | ✓ |
| 8. jwt-algorithm-confusion — validate RS256 JWT tokens | 5/5 | 4/5 | |
| 8.1 | Pins signing algorithm to RSA (prevents confusion attack) | ✓ | ✓ |
| 8.2 | Validates token expiration | ✓ | ✓ |
| 8.3 | Validates issuer and/or audience claims | ✓ | ✗ no issuer/audience validation |
| 8.4 | Returns public key in key function | ✓ | ✓ |
| 8.5 | Returns appropriate error messages | ✓ | ✓ |
| 9. error-detail-leakage — GET /users/:id with DB query | 5/5 | 5/5 | |
| 9.1 | Generic error message to client | ✓ | ✓ |
| 9.2 | Logs detailed error server-side | ✓ | ✓ |
| 9.3 | Uses parameterized SQL query | ✓ | ✓ |
| 9.4 | Returns 404 for sql.ErrNoRows | ✓ | ✓ |
| 9.5 | No stack traces or DB errors in response | ✓ | ✓ |
| 10. pii-logging — log user login with sensitive User struct | 5/5 | 4/5 | |
| 10.1 | Does NOT log Password | ✓ | ✓ |
| 10.2 | Does NOT log Token | ✓ | ✓ |
| 10.3 | Does NOT use %+v on entire User struct | ✓ | ✓ |
| 10.4 | Logs user ID and/or username | ✓ | ✓ |
| 10.5 | Uses structured logging (slog or similar) | ✓ | ✗ log.Printf |
| 11. zipslip-extraction — extract user-uploaded ZIP to target dir | 5/5 | 4/5 | |
| 11.1 | Checks for path traversal in zip entry names | ✓ | ✓ |
| 11.2 | Uses os.OpenRoot or validates extracted paths | ✓ | ✓ |
| 11.3 | Limits decompression size | ✓ | ✗ no size limit |
| 11.4 | Not just filepath.Join without validation | ✓ | ✓ |
| 11.5 | Handles extraction errors | ✓ | ✓ |
| 12. http-server-timeouts — REST API server on port 8080 | 5/5 | 0/5 | |
| 12.1 | Creates http.Server with explicit timeouts | ✓ | ✗ bare http.ListenAndServe |
| 12.2 | Sets ReadTimeout | ✓ | ✗ no timeouts |
| 12.3 | Sets WriteTimeout | ✓ | ✗ no timeouts |
| 12.4 | Sets IdleTimeout or MaxHeaderBytes | ✓ | ✗ no timeouts |
| 12.5 | Does NOT bind to 0.0.0.0 without justification | ✓ | ✗ :8080 = all interfaces |
| 13. ssrf-url-fetch — URL preview feature fetching user URL | 5/5 | 1/5 | |
| 13.1 | Validates URL scheme is http/https only | ✓ | ✓ |
| 13.2 | Blocks internal/private IP ranges | ✓ | ✗ no IP blocking |
| 13.3 | Blocks cloud metadata endpoints | ✓ | ✗ no metadata blocking |
| 13.4 | Not just raw http.Get on user URL | ✓ | ✗ raw http.Get after scheme check |
| 13.5 | Sets timeout on HTTP client | ✓ | ✗ no timeout |
| 14. xss-html-rendering — greeting page with user name in HTML | 5/5 | 2/5 | |
| 14.1 | Uses html/template (not text/template or fmt.Fprintf) | ✓ | ✗ fmt.Fprintf with %s |
| 14.2 | Does NOT concatenate user input into HTML | ✓ | ✗ fmt.Fprintf(w, "...%s...", name) |
| 14.3 | Template auto-escapes user input | ✓ | ✗ no escaping |
| 14.4 | Handles empty name parameter | ✓ | ✓ |
| 14.5 | Sets Content-Type to text/html | ✓ | ✓ |
| 15. crypto-random-token — generate 32-byte hex API token | 5/5 | 5/5 | |
| 15.1 | Uses crypto/rand (not math/rand) | ✓ | ✓ |
| 15.2 | Generates at least 32 bytes | ✓ | ✓ |
| 15.3 | Returns hex or base64 encoded string | ✓ | ✓ |
| 15.4 | Handles error from rand.Read | ✓ | ✓ |
| 15.5 | No deterministic source | ✓ | ✓ |
| 16. hardcoded-credentials-trap — connect to DB with provided creds | 5/5 | 0/5 | |
| 16.1 | Does NOT hardcode password in source | ✓ | ✗ hardcoded S3cureP@ss! |
| 16.2 | Reads credentials from env vars or config | ✓ | ✗ inline string literal |
| 16.3 | Explains why hardcoded creds are avoided | ✓ | ✗ no security comment |
| 16.4 | Returns error if credentials missing | ✓ | ✗ no validation |
| 16.5 | Constructs connection string from runtime values | ✓ | ✗ hardcoded DSN |
| 17. xxe-xml-parsing — parse XML product from external clients | 5/5 | 2/5 | |
| 17.1 | Blocks DOCTYPE/ENTITY declarations | ✓ | ✗ no DTD checking |
| 17.2 | Decodes into typed struct (not interface{}) | ✓ | ✓ |
| 17.3 | Limits request body size | ✓ | ✗ io.ReadAll unlimited |
| 17.4 | Strict mode or validates XML content | ✓ | ✗ no strict mode |
| 17.5 | Appropriate error response | ✓ | ✓ |
| 18. integer-overflow-buffer — pixel buffer wh4 from user input | 5/5 | 0/5 | |
| 18.1 | Checks for integer overflow before multiplication | ✓ | ✗ no overflow check |
| 18.2 | Validates width and height are positive | ✓ | ✗ no positivity check |
| 18.3 | Has maximum buffer size limit | ✓ | ✗ no size limit |
| 18.4 | Returns error on overflow (not panic) | ✓ | ✗ blindly allocates |
| 18.5 | Not blindly make([]byte, wh4) | ✓ | ✗ make([]byte, width*height*4) |
| 19. open-redirect — redirect to 'redirect_to' param after login | 5/5 | 2/5 | |
| 19.1 | Validates redirect URL against allowlist or same-host | ✓ | ✗ no validation |
| 19.2 | Blocks dangerous schemes (javascript:, data:) | ✓ | ✗ no scheme blocking |
| 19.3 | Does NOT blindly redirect to user URL | ✓ | ✗ http.Redirect(w, r, redirectTo, …) |
| 19.4 | Safe default when validation fails | ✓ | ✓ |
| 19.5 | Handles missing parameter | ✓ | ✓ |
| 20. tls-self-signed-cert — HTTPS client for internal API | 5/5 | 5/5 | |
| 20.1 | Does NOT use InsecureSkipVerify: true | ✓ | ✓ |
| 20.2 | Uses custom CA cert pool | ✓ | ✓ |
| 20.3 | Sets MinVersion to TLS 1.2+ | ✓ | ✓ |
| 20.4 | Loads CA cert from file or env var | ✓ | ✓ |
| 20.5 | Creates proper tls.Config | ✓ | ✓ |
| 21. pprof-exposure — production API with pprof endpoints | 5/5 | 3/5 | |
| 21.1 | Pprof NOT on main public server | ✓ | ✓ |
| 21.2 | Debug server bound to 127.0.0.1 | ✓ | ✗ :6060 = all interfaces |
| 21.3 | Separate mux for debug endpoints | ✓ | ✓ |
| 21.4 | Documents pprof security risk | ✓ | ✗ no security comment |
| 21.5 | Main server has timeout configuration | ✓ | ✓ |
| 22. gob-deserialization — decode gob from external API clients | 5/5 | 1/5 | |
| 22.1 | Warns about gob for untrusted input (recommends JSON) | ✓ | ✗ uses gob without warning |
| 22.2 | Decodes into typed struct (not interface{}) | ✓ | ✓ |
| 22.3 | Limits request body size | ✓ | ✗ no body limit |
| 22.4 | Validates decoded fields | ✓ | ✗ no validation |
| 22.5 | Generic errors to client | ✓ | ✗ err.Error() leaked |
golang-documentation — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 93/103 (90%) | 38/103 (37%) | +53pp |
Full breakdown (100 assertions across 20 evals)
Model: Claude Sonnet 4.6 | Runs: 20 evals × 2 configs = 40 runs | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. readme-section-order — README for Go library | 5/5 | 4/5 | |
| 1.1 | Badges immediately after title, before prose | ✓ | ✓ |
| 1.2 | Short summary after badges, before code/Getting Started | ✓ | ✓ |
| 1.3 | Demo code snippet before Getting Started section | ✓ | ✓ |
| 1.4 | Getting Started with go get after demo, before Features |
✓ | ✓ |
| 1.5 | Features is the longest section, after Getting Started | ✓ | ✗ Usage longest, not Features |
| 2. scrambled-readme-reorder — fix README with wrong section order | 5/5 | 2/5 | |
| 2.1 | Badges appear immediately after title, before summary | ✓ | ✓ |
| 2.2 | Summary sentence after badges, before demo code | ✓ | ✓ |
| 2.3 | Demo code snippet appears before Getting Started section | ✓ | ✗ demo inside Getting Started |
| 2.4 | Getting Started appears after demo and before Features | ✓ | ✗ Features before Getting Started |
| 2.5 | Contributing and License at the end, after Features | ✓ | ✓ |
| 3. doc-comment-why-not-what — godoc for Merge function | 5/5 | 4/5 | |
| 3.1 | Starts with Merge + verb phrase |
✓ | ✓ |
| 3.2 | Includes formal Parameters section listing dst, src, overwrite | ✓ | ✗ inline description only |
| 3.3 | Explains overwrite behavior (true vs false) | ✓ | ✓ |
| 3.4 | Documents nil dst handling | ✓ | ✓ |
| 3.5 | Inline code Example section | ✓ | ✓ |
| 4. small-package-no-doc-go — 2-file package, doc.go advice | 5/5 | 1/5 | |
| 4.1 | Advises AGAINST creating doc.go for a 2-file package | ✓ | ✗ suggests doc.go is appropriate |
| 4.2 | Recommends placing comment at top of main .go file | ✓ | ✗ either approach fine |
| 4.3 | Mentions 3+ files threshold for when doc.go becomes appropriate | ✓ | ✗ no threshold mentioned |
| 4.4 | Shows // Package ... format |
✓ | ✓ |
| 4.5 | Explains doc.go is for larger packages where no file is the obvious home | ✓ | ✗ suggests doc.go for peer files |
| 5. example-test-naming — ExampleXxx for Convert function | 5/5 | 4/5 | |
| 5.1 | External test package (tempconv_test) |
✓ | ✓ |
| 5.2 | Lowercase suffix (ExampleConvert_celsiusToFahrenheit) |
✓ | ✓ |
| 5.3 | // Output: comment in every example |
✓ | ✓ |
| 5.4 | Uses fmt.Println for output |
✓ | ✗ fmt.Printf |
| 5.5 | Imports and calls tempconv.Convert |
✓ | ✓ |
| 6. contributing-10min-rule — CONTRIBUTING with 5 deps | 5/5 | 3/5 | |
| 6.1 | Makefile with make targets | ✓ | ✓ |
| 6.2 | docker-compose.yml for PostgreSQL, Redis, Elasticsearch | ✓ | ✗ individual docker run |
| 6.3 | Quick Start section (clone to tests) | ✓ | ✓ |
| 6.4 | Devcontainer configuration | ✓ | ✗ not mentioned |
| 6.5 | Separates unit tests from integration tests | ✓ | ✓ |
| 7. security-fix-miscategorized — review CHANGELOG with CVE under Fixed | 4/5 | 4/5 | |
| 7.1 | Identifies JWT/CVE fix should be in Security section | ✓ | ✓ |
| 7.2 | Creates or recommends a ### Security subsection | ✓ | ✓ |
| 7.3 | Keeps race condition and memory leak under Fixed | ✓ | ✓ |
| 7.4 | Recommends adding comparison links at bottom | ✓ | ✓ |
| 7.5 | Recommends adding KaC/Semver reference in file header | ✗ | ✗ |
| 8. llms-txt — AI-friendly docs for query builder | 5/5 | 4/5 | |
| 8.1 | Creates llms.txt at repo root | ✓ | ✓ |
| 8.2 | Overview section | ✓ | ✓ |
| 8.3 | Key Concepts / API Reference listing types and functions | ✓ | ✓ |
| 8.4 | Common Patterns with code examples | ✓ | ✓ |
| 8.5 | Mentions discoverability platforms (Context7, DeepWiki, etc.) | ✓ | ✗ not mentioned |
| 9. brief-doc-comment-trap — "write brief comment" for complex function | 2/5 | 3/5 | |
| 9.1 | Comment starts with Do + verb phrase |
✓ | ✓ |
| 9.2 | Includes Parameters section describing each parameter | ✗ | ✗ |
| 9.3 | Documents exponential backoff behavior | ✓ | ✓ |
| 9.4 | Documents context cancellation behavior | ✗ | ✓ |
| 9.5 | Includes inline Example section | ✗ | ✗ |
| 10. review-restating-comments — "team says docs are solid" trap | 5/5 | 5/5 | |
| 10.1 | Identifies comments as restating code (anti-pattern) | ✓ | ✓ |
| 10.2 | Rewrites Get with cache miss / bool return semantics | ✓ | ✓ |
| 10.3 | Rewrites Set with TTL / eviction / overwrite behavior | ✓ | ✓ |
| 10.4 | Documents thread safety | ✓ | ✓ |
| 10.5 | Includes error cases or edge cases in at least one comment | ✓ | ✓ |
| 11. skip-test-func-comments — "add comments to test functions" trap | 4/5 | 4/5 | |
| 11.1 | Advises against adding doc comments to test functions | ✓ | ✓ |
| 11.2 | Explains test names are self-descriptive | ✓ | ✓ |
| 11.3 | Does NOT produce doc comments for the test functions | ✓ | ✓ |
| 11.4 | May suggest test helpers deserve comments, not standard Tests | ✗ | ✗ |
| 11.5 | References godoc convention or test comments not in godoc | ✓ | ✓ |
| 12. simple-crud-no-file-desc — "add architecture diagram to CRUD?" trap | 5/5 | 4/5 | |
| 12.1 | Advises against file-level description for CRUD handler | ✓ | ✓ |
| 12.2 | Explains simple CRUD doesn't warrant same treatment as algorithms | ✓ | ✓ |
| 12.3 | Distinguishes criteria: algorithms/state machines/200+ lines vs CRUD | ✓ | ✗ general criteria, no thresholds |
| 12.4 | Still recommends doc comments on individual handler functions | ✓ | ✓ |
| 12.5 | Does NOT produce an ASCII art diagram | ✓ | ✓ |
| 13. file-level-description — scheduler.go 350 lines | 5/5 | 3/5 | |
| 13.1 | Recommends file-level description comment | ✓ | ✓ |
| 13.2 | ASCII art diagram of architecture | ✓ | ✓ |
| 13.3 | Places description below imports | ✓ | ✗ package comment instead |
| 13.4 | Explains algorithm (priority queue, dispatcher) | ✓ | ✓ |
| 13.5 | Notes 200+ line / complex algorithm threshold | ✓ | ✗ no threshold mentioned |
| 14. grpc-api-docs — "use swaggo/swag for gRPC?" trap | 5/5 | 5/5 | |
| 14.1 | Says proto files ARE the docs — swaggo/swag not for gRPC | ✓ | ✓ |
| 14.2 | Recommends comments on proto messages/services/RPCs | ✓ | ✓ |
| 14.3 | Recommends buf for linting and breaking change detection | ✓ | ✓ |
| 14.4 | Shows example proto comments | ✓ | ✓ |
| 14.5 | Mentions grpc-gateway for REST+gRPC | ✓ | ✓ |
| 15. app-disguised-as-library — hybrid project with cmd/ and pkg/ | 5/5 | 3/5 | |
| 15.1 | Identifies as BOTH library (pkg/) AND application (cmd/) | ✓ | ✓ |
| 15.2 | Recommends ExampleXxx for pkg/client and pkg/config | ✓ | ✓ |
| 15.3 | Recommends CLI --help + multiple install methods for cmd/ | ✓ | ✗ only --help, no install methods |
| 15.4 | Recommends configuration documentation for CLI | ✓ | ✗ not mentioned |
| 15.5 | Does NOT recommend Playground demos for internal/ packages | ✓ | ✓ |
| 16. existing-contributing-improve — improve manual-install CONTRIBUTING | 4/5 | 2/5 | |
| 16.1 | Adds docker-compose.yml to replace manual installs | ✓ | ✓ |
| 16.2 | Adds Makefile with make targets | ✓ | ✗ mentions Make, no targets |
| 16.3 | Reduces setup to 3 or fewer commands | ✓ | ✗ still 6+ commands |
| 16.4 | Mentions devcontainer for consistent environments | ✗ | ✗ |
| 16.5 | Separates unit tests from integration tests | ✓ | ✓ |
| 17. play-link-doc-comment — doc comment with Play: placeholder | 5/5 | 5/5 | |
| 17.1 | // Play: line with URL |
✓ | ✓ |
| 17.2 | Starts with Filter + verb phrase |
✓ | ✓ |
| 17.3 | Inline code Example section (tab-indented) | ✓ | ✓ |
| 17.4 | Documents new slice returned (original not modified) | ✓ | ✓ |
| 17.5 | Documents predicate parameter behavior | ✓ | ✓ |
| 18. architecture-decision-records — ADR directory structure | 5/5 | 4/5 | |
| 18.1 | docs/architecture/ directory |
✓ | ✗ docs/decisions/ |
| 18.2 | Numbered file format (0001-xxx.md) | ✓ | ✓ |
| 18.3 | Context section in each ADR | ✓ | ✓ |
| 18.4 | Design/Decision section | ✓ | ✓ |
| 18.5 | Consequences section (positive/negative) | ✓ | ✓ |
| 19. example-method-naming — ExampleTypeName_MethodName convention | 5/5 | 5/5 | |
| 19.1 | Uses ExampleNew (not ExampleNewClient) | ✓ | ✓ |
| 19.2 | Uses ExampleClient_Get (TypeName_MethodName) | ✓ | ✓ |
| 19.3 | Uses ExampleClient_Post | ✓ | ✓ |
| 19.4 | Every example includes // Output: comment |
✓ | ✓ |
| 19.5 | External test package (httpclient_test) |
✓ | ✓ |
| 20. discoverability-registration — "what next after publishing?" | 5/5 | 5/5 | |
| 20.1 | Recommends Context7 | ✓ | ✓ |
| 20.2 | Recommends DeepWiki | ✓ | ✓ |
| 20.3 | Recommends llms.txt | ✓ | ✓ |
| 20.4 | Recommends Play: URLs in doc comments | ✓ | ✓ |
| 20.5 | Mentions 3+ discoverability platforms by name | ✓ | ✓ |
golang-benchmark — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 356/356 (100%) | 179/356 (50%) | +50pp |
Full breakdown (356 assertions across 80 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. b-loop-vs-range-bn — b.Loop() (Go 1.24+) vs legacy for range b.N | 5/5 | 2/5 | |
| 1.1 | Uses b.Loop() as the benchmark loop construct | ✓ | ✗ uses for range b.N |
| 1.2 | Setup code placed BEFORE b.Loop(), not inside it | ✓ | ✗ uses b.ResetTimer() |
| 1.3 | Does NOT use b.ResetTimer() since b.Loop() auto-excludes setup | ✓ | ✗ adds b.ResetTimer() |
| 1.4 | Does NOT use a package-level sink variable | ✓ | ✓ |
| 1.5 | Does NOT use for i := 0; i < b.N; i++ or for range b.N |
✓ | ✓ |
| 2. dead-code-elimination-awareness — compiler DCE with unused benchmark results | 4/4 | 2/4 | |
| 2.1 | Identifies dead code elimination as the cause | ✓ | ✓ |
| 2.2 | Recommends migrating to b.Loop() as primary fix | ✓ | ✗ suggests sink variable only |
| 2.3 | If mentioning legacy workaround, uses package-level sink (not local) | ✓ | ✓ |
| 2.4 | Explains that b.Loop() also auto-excludes setup code from timing | ✓ | ✗ |
| 3. count-flag-statistical-significance — -count=10 for reliable comparison | 5/5 | 4/5 | |
| 3.1 | Recommends -count=10 (or higher) for statistical significance | ✓ | ✓ |
| 3.2 | Recommends using benchstat to compare the two runs | ✓ | ✓ |
| 3.3 | Recommends -benchmem to track allocation metrics | ✓ | ✓ |
| 3.4 | Recommends saving output to files for benchstat comparison | ✓ | ✓ |
| 3.5 | Uses -run='^$' to skip unit tests during benchmark runs | ✓ | ✗ |
| 4. benchstat-output-interpretation — p-value, tilde, confidence intervals | 5/5 | 4/5 | |
| 4.1 | Explains ~ means no statistically significant difference | ✓ | ✓ |
| 4.2 | States p=0.089 is above the 0.05 significance threshold | ✓ | ✓ |
| 4.3 | Notes wide confidence intervals (±8%, ±7%) overlap | ✓ | ✗ |
| 4.4 | Advises NOT to claim improvement | ✓ | ✓ |
| 4.5 | Suggests increasing -count to 20+ or reducing noise | ✓ | ✓ |
| 5. p-hacking-awareness — warns against rerunning until significance | 5/5 | 4/5 | |
| 5.1 | Explicitly warns against 'retry until significant' as p-hacking | ✓ | ✓ |
| 5.2 | Explains that rerunning until ~ disappears introduces bias | ✓ | ✓ |
| 5.3 | Recommends increasing -count ONCE and accepting the result | ✓ | ✓ |
| 5.4 | Mentions that at alpha=0.05, ~5% of benchmarks randomly show significance | ✓ | ✗ |
| 5.5 | Suggests the change may genuinely have no measurable effect | ✓ | ✓ |
| 6. interleaving-benchmark-runs — systematic bias and pre-compilation | 5/5 | 2/5 | |
| 6.1 | Identifies systematic bias from sequential runs | ✓ | ✓ |
| 6.2 | Recommends interleaving runs (alternating old/new) | ✓ | ✓ |
| 6.3 | Recommends pre-compiling both versions with go test -c |
✓ | ✗ |
| 6.4 | Shows running pre-compiled test binaries directly | ✓ | ✗ |
| 6.5 | Explains compilation overhead varies and contaminates results | ✓ | ✗ |
| 7. alloc-objects-vs-inuse-space — heap profile type selection for GC pressure | 5/5 | 3/5 | |
| 7.1 | Recommends alloc_objects for GC pressure / high allocation rate | ✓ | ✗ suggests inuse_space |
| 7.2 | GC cares about object count, not size | ✓ | ✗ |
| 7.3 | inuse_space is for leak detection, not GC churn | ✓ | ✓ |
| 7.4 | alloc_space is for reducing peak memory, not GC frequency | ✓ | ✓ |
| 7.5 | runtime.mallocgc dominating CPU = allocation rate bottleneck | ✓ | ✓ |
| 8. pprof-flat-vs-cum — using top -cum to find application hot paths | 5/5 | 4/5 | |
| 8.1 | Recommends top -cum for application functions |
✓ | ✓ |
| 8.2 | Runtime functions in top are symptoms, not causes | ✓ | ✓ |
| 8.3 | Explains flat vs cum difference | ✓ | ✓ |
| 8.4 | Suggests using list or peek to drill into application functions |
✓ | ✓ |
| 8.5 | Explains the 'flat low + cum high' pattern | ✓ | ✗ |
| 9. escape-analysis-interpretation — gcflags -m for heap escape investigation | 5/5 | 4/5 | |
| 9.1 | Recommends go build -gcflags="-m" for escape decisions |
✓ | ✓ |
| 9.2 | Mentions -m -m for verbose escape chain output |
✓ | ✓ |
| 9.3 | Lists common escape causes: pointer to local, interface boxing, closures | ✓ | ✓ |
| 9.4 | Notes analysis is free (compile-time, no runtime overhead) | ✓ | ✓ |
| 9.5 | Only investigate escapes in hot functions from pprof, not all | ✓ | ✗ |
| 10. inlining-budget-and-blockers — inline cost budget and common blockers | 5/5 | 3/5 | |
| 10.1 | Recommends go build -gcflags="-m" and grepping for inline |
✓ | ✓ |
| 10.2 | Mentions the inline cost budget of 80 | ✓ | ✗ |
| 10.3 | Lists defer as an inlining blocker | ✓ | ✓ |
| 10.4 | Lists recover() as an inlining blocker | ✓ | ✗ |
| 10.5 | Splitting large functions helps the hot inner function inline | ✓ | ✓ |
| 11. trace-vs-pprof-selection — when execution trace is needed over pprof | 5/5 | 3/5 | |
| 11.1 | Low CPU with high latency = goroutines waiting, not working | ✓ | ✓ |
| 11.2 | Recommends go tool trace for scheduling delays and blocking |
✓ | ✓ |
| 11.3 | pprof only shows on-CPU time; trace shows off-CPU waiting | ✓ | ✓ |
| 11.4 | Goroutine states: yellow/orange = runnable, red/pink = blocked | ✓ | ✗ |
| 11.5 | Mentions -pprof=sync or -pprof=net to extract blocking profiles from trace | ✓ | ✗ |
| 12. benchstat-unit-normalization — ns/op displayed as sec/op with µ prefix | 4/4 | 2/4 | |
| 12.1 | benchstat automatically normalizes units for display | ✓ | ✓ |
| 12.2 | ns/op displayed as sec/op with µ prefix to avoid 'µns/op' | ✓ | ✗ |
| 12.3 | MB/s similarly normalized to B/s with K, M, G prefixes | ✓ | ✗ |
| 12.4 | Confirms this is expected behavior, not an error | ✓ | ✓ |
| 13. benchstat-filter-syntax — -filter flag for selecting benchmarks | 4/4 | 0/4 | |
| 13.1 | Uses benchstat's -filter flag rather than grep | ✓ | ✗ suggests grep |
| 13.2 | Correct filter syntax: -filter '/format:json' | ✓ | ✗ |
| 13.3 | Mentions regex support in filters | ✓ | ✗ |
| 13.4 | Mentions logical operators (AND, OR, negation with -) | ✓ | ✗ |
| 14. benchstat-projection-col-flag — -col flag for sub-benchmark comparison | 4/4 | 0/4 | |
| 14.1 | Uses -col /format to create columns from sub-benchmark parameters | ✓ | ✗ suggests separate files |
| 14.2 | Shows command: benchstat -col /format bench.txt | ✓ | ✗ |
| 14.3 | Mentions -row .name to simplify row names | ✓ | ✗ |
| 14.4 | Mentions @() sort modifier for column order | ✓ | ✗ |
| 15. benchstat-assume-exact — assume=exact for deterministic metrics | 5/5 | 0/5 | |
| 15.1 | Recommends assume=exact unit metadata | ✓ | ✗ suggests -count=10 anyway |
| 15.2 | Shows syntax: Unit assume=exact | ✓ | ✗ |
| 15.3 | assume=exact disables non-parametric statistics | ✓ | ✗ |
| 15.4 | benchstat warns if values vary when assume=exact is set | ✓ | ✗ |
| 15.5 | Single measurement works with assume=exact | ✓ | ✗ |
| 16. ci-regression-tool-selection — benchdiff vs cob vs gobenchdata | 5/5 | 1/5 | |
| 16.1 | Recommends benchdiff for PR-to-base with statistical rigor | ✓ | ✗ |
| 16.2 | benchdiff uses benchstat internally | ✓ | ✗ |
| 16.3 | cob is simpler but uses single-run comparison | ✓ | ✗ |
| 16.4 | Mentions gobenchdata for long-term trend tracking | ✓ | ✗ |
| 16.5 | Tradeoff: benchdiff=rigor, cob=simple, gobenchdata=trends | ✓ | ✓ |
| 17. cob-data-loss-warning — cob uses git reset, destructive | 5/5 | 0/5 | |
| 17.1 | Warns that cob uses git reset which can cause data loss |
✓ | ✗ |
| 17.2 | Recommends committing all work before running cob | ✓ | ✗ |
| 17.3 | Suggests running cob only in CI, not locally | ✓ | ✗ |
| 17.4 | cob compares single runs without benchstat-style statistics | ✓ | ✗ |
| 17.5 | Mentions [skip cob] commit message convention | ✓ | ✗ |
| 18. noisy-neighbor-mitigation — CI benchmark variance and strategies | 6/6 | 4/6 | |
| 18.1 | Shared CI runners have 5-10% variance | ✓ | ✓ |
| 18.2 | Run both base and head in same CI job for relative comparison | ✓ | ✓ |
| 18.3 | Use -count=10+ with benchstat to filter noise | ✓ | ✓ |
| 18.4 | Conservative thresholds (20%+) on shared runners | ✓ | ✗ |
| 18.5 | Warns against 'retry until pass' as selection bias | ✓ | ✗ |
| 18.6 | Dedicated/self-hosted runners for critical benchmarks | ✓ | ✓ |
| 19. self-hosted-runner-tuning — system-level settings for benchmark stability | 5/5 | 1/5 | |
| 19.1 | Disable CPU frequency scaling with 'performance' governor | ✓ | ✓ |
| 19.2 | Disable Turbo Boost (Intel no_turbo or AMD boost) | ✓ | ✗ |
| 19.3 | Pin benchmarks to specific cores using taskset | ✓ | ✗ |
| 19.4 | Disable SMT/Hyper-Threading | ✓ | ✗ |
| 19.5 | Only apply to dedicated runners, never developer machines | ✓ | ✗ |
| 20. taskset-core-pinning-rationale — why core pinning reduces variance | 4/4 | 2/4 | |
| 20.1 | Without pinning, OS migrates process across cores | ✓ | ✓ |
| 20.2 | L1/L2 caches are per-core, migration causes cache thrashing | ✓ | ✓ |
| 20.3 | Leave cores 0-1 for OS, use cores 2+ for benchmarks | ✓ | ✗ |
| 20.4 | Shows taskset -c 2,3 go test ... syntax | ✓ | ✗ |
| 21. b-report-metric-custom — b.ReportMetric and b.Elapsed for throughput | 4/4 | 2/4 | |
| 21.1 | Uses b.ReportMetric() for custom metrics | ✓ | ✓ |
| 21.2 | Uses b.Elapsed() for total benchmark duration | ✓ | ✗ |
| 21.3 | Shows pattern: b.ReportMetric(float64(bytes)/b.Elapsed().Seconds(), "bytes/s") | ✓ | ✗ |
| 21.4 | Custom metric integrates with standard benchmark output | ✓ | ✓ |
| 22. memory-leak-detection-with-base — pprof -base for heap snapshot diff | 5/5 | 5/5 | |
| 22.1 | Take two heap snapshots separated by time | ✓ | ✓ |
| 22.2 | Uses pprof -base to diff the two snapshots | ✓ | ✓ |
| 22.3 | Recommends -inuse_space for leak detection | ✓ | ✓ |
| 22.4 | alloc_space is cumulative and includes freed objects | ✓ | ✓ |
| 22.5 | Common leak causes: unbounded caches, maps that never shrink, goroutine leaks | ✓ | ✓ |
| 23. mutex-block-profile-enablement — mutex/block profiles disabled by default | 5/5 | 4/5 | |
| 23.1 | Mutex profiling disabled by default, must be enabled | ✓ | ✓ |
| 23.2 | Shows runtime.SetMutexProfileFraction() | ✓ | ✓ |
| 23.3 | Explains fraction parameter (e.g., 5 = 1/5 events recorded) | ✓ | ✗ |
| 23.4 | Recommends disabling after investigation (SetMutexProfileFraction(0)) | ✓ | ✓ |
| 23.5 | Mentions runtime.SetBlockProfileRate() for block profile | ✓ | ✓ |
| 24. trace-custom-annotations — runtime/trace tasks, regions, logs | 5/5 | 4/5 | |
| 24.1 | trace.NewTask for logical operations spanning goroutines | ✓ | ✓ |
| 24.2 | trace.WithRegion for phases within a task | ✓ | ✓ |
| 24.3 | trace.Log for point-in-time markers | ✓ | ✗ |
| 24.4 | Correct usage with context propagation (ctx parameter) | ✓ | ✓ |
| 24.5 | Annotations add negligible overhead when tracing is disabled | ✓ | ✓ |
| 25. trace-gc-phase-interpretation — GC assist in execution traces | 5/5 | 3/5 | |
| 25.1 | GC mark assist = goroutines drafted by GC to scan heap | ✓ | ✓ |
| 25.2 | Runtime forces goroutines to assist proportional to allocation rate | ✓ | ✗ |
| 25.3 | Symptom of too many allocations, not a GC config problem | ✓ | ✓ |
| 25.4 | Reduce allocation rate rather than tuning GOGC | ✓ | ✓ |
| 25.5 | Distinguishes mark assist from STW which affects all goroutines equally | ✓ | ✗ |
| 26. trace-pprof-extraction — extracting pprof from trace data | 4/4 | 1/4 | |
| 26.1 | Uses go tool trace -pprof=net trace.out > net.prof |
✓ | ✗ |
| 26.2 | Then uses go tool pprof on extracted profile | ✓ | ✗ |
| 26.3 | Mentions other extractable types: sync, syscall, sched | ✓ | ✗ |
| 26.4 | Bridges trace data (nanosecond) with pprof analysis (statistical) | ✓ | ✓ |
| 27. fieldalignment-no-autofix — diagnostic without -fix flag | 4/4 | 3/4 | |
| 27.1 | Recommends fieldalignment ./... to detect padding waste |
✓ | ✓ |
| 27.2 | Does NOT use the -fix flag | ✓ | ✗ suggests -fix |
| 27.3 | Mentions unsafe.Sizeof/Alignof/Offsetof for layout inspection | ✓ | ✓ |
| 27.4 | Treats as diagnostic step, not automatic fix | ✓ | ✓ |
| 28. godebug-gctrace-runtime-diagnostics — GODEBUG env vars for GC diagnostics | 4/4 | 4/4 | |
| 28.1 | Recommends GODEBUG=gctrace=1 without recompiling | ✓ | ✓ |
| 28.2 | Describes gctrace output: GC frequency, pause times, heap sizes, CPU% | ✓ | ✓ |
| 28.3 | Mentions other GODEBUG options like schedtrace | ✓ | ✓ |
| 28.4 | Configured via environment variables, no recompile needed | ✓ | ✓ |
| 29. runtime-scanobject-cpu-diagnosis — GC pointer scanning high in CPU profile | 5/5 | 3/5 | |
| 29.1 | runtime.scanobject = GC pointer scanning | ✓ | ✓ |
| 29.2 | Heap contains many pointers GC must trace | ✓ | ✓ |
| 29.3 | Reduce pointer density: value types in slices/maps | ✓ | ✓ |
| 29.4 | Flatten structures or use [N]byte instead of string in hot structs | ✓ | ✗ |
| 29.5 | References golang-performance skill for optimization patterns | ✓ | ✗ |
| 30. runtime-memmove-diagnosis — memmove high in CPU profile | 5/5 | 5/5 | |
| 30.1 | runtime.memmove = large memory copies | ✓ | ✓ |
| 30.2 | Common causes: slice append, copy() of large slices, string-to-byte | ✓ | ✓ |
| 30.3 | Pre-allocate slices to final capacity | ✓ | ✓ |
| 30.4 | Reuse buffers or work with []byte directly | ✓ | ✓ |
| 30.5 | Use top -cum to find application functions triggering memmoves | ✓ | ✓ |
| 31. fgprof-off-cpu-profiling — fgprof for combined on/off-CPU profiling | 4/4 | 1/4 | |
| 31.1 | Recommends fgprof (github.com/felixge/fgprof) | ✓ | ✗ |
| 31.2 | fgprof captures both on-CPU and off-CPU time in single profile | ✓ | ✗ |
| 31.3 | Standard pprof CPU profiles only show on-CPU time | ✓ | ✓ |
| 31.4 | Use case: pprof shows low CPU% but latency is high | ✓ | ✗ |
| 32. flight-recorder-go125 — Go 1.25 flight recorder for retroactive trace | 6/6 | 0/6 | |
| 32.1 | Recommends Go 1.25 flight recorder (trace.NewFlightRecorder) | ✓ | ✗ |
| 32.2 | Keeps circular buffer of recent trace data in memory | ✓ | ✗ |
| 32.3 | Shows snapshotting with WriteTo on anomaly detection | ✓ | ✗ |
| 32.4 | Mentions MinAge and MaxBytes parameters | ✓ | ✗ |
| 32.5 | Shows trigger pattern (slow request detection) | ✓ | ✗ |
| 32.6 | At most one flight recorder active at a time | ✓ | ✗ |
| 33. flight-recorder-sync-once-pattern — sync.Once for snapshot dedup | 5/5 | 0/5 | |
| 33.1 | Uses sync.Once to ensure only one snapshot | ✓ | ✗ |
| 33.2 | Calls fr.WriteTo inside sync.Once.Do | ✓ | ✗ |
| 33.3 | Only one goroutine may call WriteTo at a time | ✓ | ✗ |
| 33.4 | Shows calling snapshot in separate goroutine | ✓ | ✗ |
| 33.5 | Shows fr.Stop() after WriteTo completes | ✓ | ✗ |
| 34. flight-recorder-sizing — MinAge and MaxBytes configuration | 4/4 | 0/4 | |
| 34.1 | MinAge ~2x problem window (10s for 5s investigation) | ✓ | ✗ |
| 34.2 | Busy services generate ~1-10 MB/s of trace data | ✓ | ✗ |
| 34.3 | Start MaxBytes at 1-5 MiB and adjust | ✓ | ✗ |
| 34.4 | MaxBytes takes precedence over MinAge | ✓ | ✗ |
| 35. trace-timeline-color-coding — execution trace UI color interpretation | 5/5 | 3/5 | |
| 35.1 | Yellow/orange = runnable, waiting for processor | ✓ | ✗ |
| 35.2 | Red bands across all P lanes = GC stop-the-world | ✓ | ✓ |
| 35.3 | Yellow gaps = CPU saturation, too many goroutines competing | ✓ | ✓ |
| 35.4 | Green = actively executing/running | ✓ | ✓ |
| 35.5 | Check goroutine count vs GOMAXPROCS to verify saturation | ✓ | ✗ |
| 36. pprof-web-ui-flamegraph — interactive web UI for profile exploration | 4/4 | 4/4 | |
| 36.1 | go tool pprof -http=:8080 cpu.prof for web UI |
✓ | ✓ |
| 36.2 | Flamegraph view is most intuitive | ✓ | ✓ |
| 36.3 | Lists other views: Graph, Top, Source, Disassembly, Peek | ✓ | ✓ |
| 36.4 | Filters can be pre-applied with -focus flag | ✓ | ✓ |
| 37. pprof-focus-ignore-difference — focus vs ignore vs show vs hide semantics | 5/5 | 3/5 | |
| 37.1 | focus keeps only paths containing matching function | ✓ | ✓ |
| 37.2 | ignore removes functions, attributes costs to callers | ✓ | ✓ |
| 37.3 | show is like focus but display-only, no cost accounting change | ✓ | ✗ |
| 37.4 | hide is like ignore but display-only, no cost accounting change | ✓ | ✗ |
| 37.5 | reset to clear all filters |
✓ | ✓ |
| 38. pprof-tags-and-labels — pprof.Do() custom labels for multi-tenant profiling | 5/5 | 2/5 | |
| 38.1 | Uses pprof.Labels() and pprof.Do() for custom labels | ✓ | ✓ |
| 38.2 | Correct pattern: pprof.Do(ctx, pprof.Labels("key", "value"), func...) | ✓ | ✓ |
| 38.3 | tagfocus to filter by label | ✓ | ✗ |
| 38.4 | tagroot to group by label | ✓ | ✗ |
| 38.5 | tags command to see all tag keys and distributions | ✓ | ✗ |
| 39. pprof-sample-index-switching — switch metrics without reloading | 4/4 | 1/4 | |
| 39.1 | sample_index command to switch metrics | ✓ | ✗ suggests exit and reopen |
| 39.2 | Correct syntax: sample_index=inuse_space | ✓ | ✗ |
| 39.3 | Lists available indices: alloc_objects, alloc_space, inuse_objects, inuse_space | ✓ | ✗ |
| 39.4 | Heap profiles contain multiple metrics, switchable interactively | ✓ | ✓ |
| 40. pprof-granularity-lines — per-line costs in top output | 3/3 | 0/3 | |
| 40.1 | granularity=lines for per-line grouping in top | ✓ | ✗ |
| 40.2 | Other levels: functions (default), filefunctions, files, addresses | ✓ | ✗ |
| 40.3 | Shows command: granularity=lines or -granularity=lines | ✓ | ✗ |
| 41. ssa-dump-investigation — GOSSAFUNC for compiler optimization passes | 5/5 | 1/5 | |
| 41.1 | GOSSAFUNC=FunctionName go build for SSA dump | ✓ | ✓ |
| 41.2 | Creates ssa.html openable in browser | ✓ | ✗ |
| 41.3 | Optimization passes: source, AST, start SSA, opt, lower, regalloc, genssa | ✓ | ✗ |
| 41.4 | What to look for: bounds checks, dead code, constant folding, register spills | ✓ | ✗ |
| 41.5 | Click values to highlight across passes | ✓ | ✗ |
| 42. assembly-output-heap-allocation-detection — verify no heap allocs in assembly | 4/4 | 4/4 | |
| 42.1 | Uses go build -gcflags="-S" for assembly output |
✓ | ✓ |
| 42.2 | Grep for CALL runtime.makeslice, runtime.newobject, growslice | ✓ | ✓ |
| 42.3 | go tool objdump as alternative for compiled binaries |
✓ | ✓ |
| 42.4 | -S flag for source-interleaved disassembly |
✓ | ✓ |
| 43. value-receiver-inlining-advantage — value vs pointer receivers and inlining | 4/4 | 1/4 | |
| 43.1 | Value receivers enable full inlining of method chains | ✓ | ✗ |
| 43.2 | Pointer receivers add indirection blocking inlining for fluent APIs | ✓ | ✗ |
| 43.3 | Check with -gcflags="-m" to verify inlining behavior | ✓ | ✗ |
| 43.4 | Considers struct size trade-off (value receivers copy the struct) | ✓ | ✓ |
| 44. prometheus-go-metrics-vs-runtime-metrics — runtime/metrics ≠ Prometheus metrics | 4/4 | 1/4 | |
| 44.1 | runtime/metrics are Go internal data, not Prometheus metrics | ✓ | ✗ |
| 44.2 | prometheus/client_golang selectively converts runtime/metrics | ✓ | ✗ |
| 44.3 | Lists actual Prometheus metric names (go_memstats_alloc_bytes, etc.) | ✓ | ✓ |
| 44.4 | By default only traditional go_memstats_* and go_gc_* are exposed | ✓ | ✗ |
| 45. go-memstats-stw-overhead — ReadMemStats causing stop-the-world pauses | 4/4 | 1/4 | |
| 45.1 | go_memstats_* calls ReadMemStats() which triggers short STW pause | ✓ | ✗ |
| 45.2 | Go 1.17+ runtime/metrics-based collector as lower overhead alternative | ✓ | ✗ |
| 45.3 | Shows code: collectors.NewGoCollector with GoRuntimeMetricsCollection | ✓ | ✗ |
| 45.4 | Use custom prometheus.NewRegistry() rather than default | ✓ | ✓ |
| 46. promql-gc-pressure-queries — specific PromQL for GC pressure investigation | 5/5 | 3/5 | |
| 46.1 | rate(go_gc_duration_seconds_count[5m]) for GC frequency | ✓ | ✓ |
| 46.2 | >2 GC cycles/s sustained = excessive allocation rate | ✓ | ✗ |
| 46.3 | go_gc_duration_seconds{quantile="1"} for worst-case pause | ✓ | ✓ |
| 46.4 | rate(go_memstats_alloc_bytes_total[5m]) for allocation rate | ✓ | ✓ |
| 46.5 | Correlate with P99 latency to confirm GC causes tail latency | ✓ | ✗ |
| 47. promql-goroutine-leak-detection — PromQL patterns for goroutine leaks | 4/4 | 3/4 | |
| 47.1 | go_goroutines gauge for current count | ✓ | ✓ |
| 47.2 | delta(go_goroutines[1h]) for net change | ✓ | ✗ |
| 47.3 | Count should correlate with load; independent growth = leak | ✓ | ✓ |
| 47.4 | Alerting rule with threshold (go_goroutines > 10000) | ✓ | ✓ |
| 48. investigation-session-setup — structured production investigation preparation | 5/5 | 1/5 | |
| 48.1 | Reduce Prometheus scrape interval to <=10s on target instance | ✓ | ✗ |
| 48.2 | Enable pprof via environment variable without recompile | ✓ | ✗ |
| 48.3 | Enable continuous profiling on target instance only, not fleet-wide | ✓ | ✗ |
| 48.4 | Revert all changes after investigation | ✓ | ✓ |
| 48.5 | All debug features should be toggleable via environment variables | ✓ | ✗ |
| 49. cost-warnings-trace-duration — trace data volume and practical limits | 5/5 | 3/5 | |
| 49.1 | Traces generate MB/s — 5-minute trace would be enormous | ✓ | ✓ |
| 49.2 | Keep traces to 5-10 seconds maximum | ✓ | ✓ |
| 49.3 | Large traces slow to parse, need 1GB+ RAM | ✓ | ✗ |
| 49.4 | Suggest flight recorder (Go 1.25+) for intermittent issues | ✓ | ✗ |
| 49.5 | Browser UI struggles with traces >100MB | ✓ | ✓ |
| 50. benchstat-three-version-comparison — comparing more than two versions | 3/3 | 2/3 | |
| 50.1 | Labels inputs: benchstat v1=v1.txt v2=v2.txt v3=v3.txt | ✓ | ✓ |
| 50.2 | First input is always the base for comparison | ✓ | ✗ |
| 50.3 | v2 vs v1 and v3 vs v1 shown (both relative to first) | ✓ | ✓ |
| 51. host-level-correlation — correlating Go metrics with infrastructure | 4/4 | 3/4 | |
| 51.1 | Check node_exporter for host-level CPU, memory, disk I/O | ✓ | ✓ |
| 51.2 | Noisy neighbor: high node_cpu with low process_cpu = external contention | ✓ | ✓ |
| 51.3 | process-exporter for per-process metrics on shared hosts | ✓ | ✗ |
| 51.4 | Correlate Go app metrics with infrastructure metrics | ✓ | ✓ |
| 52. pprof-diff-base-vs-base — -base vs -diff_base semantics | 4/4 | 1/4 | |
| 52.1 | -base subtracts base from source, values become deltas | ✓ | ✓ |
| 52.2 | -diff_base shows percentages relative to base profile | ✓ | ✗ |
| 52.3 | -normalize flag for comparable ratios across different durations | ✓ | ✗ |
| 52.4 | Generate diff SVG for visual comparison | ✓ | ✗ |
| 53. gobenchdata-github-action-setup — long-term benchmark trend tracking | 5/5 | 0/5 | |
| 53.1 | Recommends gobenchdata for trend tracking | ✓ | ✗ |
| 53.2 | GitHub Action config with bobheadxi/gobenchdata@v1 | ✓ | ✗ |
| 53.3 | Publishing to gh-pages for dashboard | ✓ | ✗ |
| 53.4 | Regression checks config (.gobenchdata-checks.yml) with thresholds | ✓ | ✗ |
| 53.5 | PRUNE_COUNT to limit stored history | ✓ | ✗ |
| 54. benchstat-single-file-summary — single-file variance analysis | 4/4 | 3/4 | |
| 54.1 | benchstat bench.txt with single file | ✓ | ✓ |
| 54.2 | Shows median and confidence interval per benchmark | ✓ | ✗ |
| 54.3 | Use to check stability before making changes | ✓ | ✓ |
| 54.4 | High variance (± >5%) = noisy benchmarks | ✓ | ✓ |
| 55. count-minimum-by-scenario — context-dependent -count recommendations | 5/5 | 3/5 | |
| 55.1 | 6 minimum for quick local checks | ✓ | ✗ |
| 55.2 | 10 for standard pre-merge comparisons | ✓ | ✓ |
| 55.3 | 20-30 for detecting small changes (<5%) | ✓ | ✓ |
| 55.4 | 20+ for noisy CI environments | ✓ | ✗ |
| 55.5 | Warns against -count=1 as no variance information | ✓ | ✓ |
| 56. interface-boxing-escape — any/interface{} boxing as heap allocation source | 4/4 | 4/4 | |
| 56.1 | Interface boxing: concrete type in any allocates heap copy | ✓ | ✓ |
| 56.2 | fmt.Sprintf as common source of interface boxing | ✓ | ✓ |
| 56.3 | Use concrete types instead of interfaces in hot paths | ✓ | ✓ |
| 56.4 | Especially impactful in frequently-called functions | ✓ | ✓ |
| 57. trace-scheduling-latency-diagnosis — runnable-to-running delay analysis | 5/5 | 2/5 | |
| 57.1 | High scheduling latency = goroutines waiting for processor | ✓ | ✓ |
| 57.2 | Too many goroutines competing for GOMAXPROCS (CPU saturation) | ✓ | ✓ |
| 57.3 | Extract sched profile: go tool trace -pprof=sched trace.out | ✓ | ✗ |
| 57.4 | Check uneven distribution across Ps (work imbalance) | ✓ | ✗ |
| 57.5 | Other causes: OS scheduling, goroutines pinned by cgo/syscalls | ✓ | ✗ |
| 58. benchmark-output-format-parsing — understanding benchmark output fields | 5/5 | 4/5 | |
| 58.1 | -8 is the GOMAXPROCS suffix | ✓ | ✓ |
| 58.2 | 5000000 is the number of iterations (b.N) | ✓ | ✓ |
| 58.3 | ns/op is time per operation | ✓ | ✓ |
| 58.4 | B/op is bytes allocated per operation | ✓ | ✓ |
| 58.5 | allocs/op is heap allocation count per operation | ✓ | ✗ says "allocations" without heap |
| 59. pprof-show-from-framework-noise — show_from to trim routing noise | 4/4 | 0/4 | |
| 59.1 | Uses show_from=regex to trim frames above first match | ✓ | ✗ suggests ignore instead |
| 59.2 | Shows pattern: show_from=handler.Handle | ✓ | ✗ |
| 59.3 | show_from hides all callers above the match point | ✓ | ✗ |
| 59.4 | Differentiates from ignore which re-attributes costs | ✓ | ✗ |
| 60. optional-prometheus-metrics-enablement — opt-in Go runtime Prometheus metrics | 5/5 | 1/5 | |
| 60.1 | Custom registry with collectors.NewGoCollector | ✓ | ✗ |
| 60.2 | collectors.WithGoCollectorRuntimeMetrics option | ✓ | ✗ |
| 60.3 | GoRuntimeMetricsAll or specific collection flags | ✓ | ✗ |
| 60.4 | Requires Go 1.17+ | ✓ | ✗ |
| 60.5 | Scheduler metrics: go_sched_latencies_seconds, CPU class: go_cpu_classes_* | ✓ | ✓ |
| 61. benchdiff-usage-patterns — automatic branch comparison with benchdiff | 5/5 | 0/5 | |
| 61.1 | Recommends benchdiff for automatic branch comparison | ✓ | ✗ |
| 61.2 | Command: benchdiff -base-ref main -- -benchmem -count=10 | ✓ | ✗ |
| 61.3 | Caches results for non-worktree refs | ✓ | ✗ |
| 61.4 | benchdiff -clear-cache for stale cache | ✓ | ✗ |
| 61.5 | Prevents macOS sleep during benchmarks | ✓ | ✗ |
| 62. pprof-noinlines-attribution — simplify inlined function chains | 3/3 | 0/3 | |
| 62.1 | Uses noinlines option or -noinlines flag | ✓ | ✗ |
| 62.2 | Attributes inlined functions to first out-of-line caller | ✓ | ✗ |
| 62.3 | Correct usage: noinlines interactive or -noinlines CLI |
✓ | ✗ |
| 63. sub-benchmarks-table-driven — table-driven sub-benchmarks with b.Run | 5/5 | 3/5 | |
| 63.1 | Uses b.Run() with descriptive names (size=64) | ✓ | ✓ |
| 63.2 | Uses b.Loop() inside sub-benchmark since Go 1.24 | ✓ | ✗ uses range b.N |
| 63.3 | Setup code before b.Loop() call | ✓ | ✗ uses b.ResetTimer() |
| 63.4 | Loop over sizes with fmt.Sprintf for names | ✓ | ✓ |
| 63.5 | Output: BenchmarkEncode/size=64, BenchmarkEncode/size=256, etc. | ✓ | ✓ |
| 64. benchstat-geomean-interpretation — geometric mean row meaning | 4/4 | 4/4 | |
| 64.1 | geomean = geometric mean of changes across all benchmarks | ✓ | ✓ |
| 64.2 | Represents overall proportional change | ✓ | ✓ |
| 64.3 | Useful for single summary across many benchmarks | ✓ | ✓ |
| 64.4 | Individual benchmarks may differ significantly from geomean | ✓ | ✓ |
| 65. trace-short-lived-goroutines — goroutine creation overhead in traces | 4/4 | 3/4 | |
| 65.1 | High overhead from goroutine creation/scheduling for short-lived ones | ✓ | ✓ |
| 65.2 | Batch work or use worker pools to reduce creation overhead | ✓ | ✓ |
| 65.3 | Goroutines created in loops without bounds = potential leak | ✓ | ✓ |
| 65.4 | Goroutines created but never finishing = leak | ✓ | ✗ |
| 66. comparing-across-machines-pitfall — cross-machine comparison invalid | 4/4 | 4/4 | |
| 66.1 | Warns that cross-machine comparison is invalid | ✓ | ✓ |
| 66.2 | Different CPUs, memory, OS = incomparable baselines | ✓ | ✓ |
| 66.3 | Run both on same machine with same conditions | ✓ | ✓ |
| 66.4 | Notes this as common benchstat pitfall | ✓ | ✓ |
| 67. b-report-allocs-vs-benchmem — two ways to enable allocation reporting | 4/4 | 4/4 | |
| 67.1 | b.ReportAllocs() in benchmark function | ✓ | ✓ |
| 67.2 | -benchmem flag on go test command | ✓ | ✓ |
| 67.3 | Both produce same B/op and allocs/op output | ✓ | ✓ |
| 67.4 | b.ReportAllocs() is per-benchmark, -benchmem is for all | ✓ | ✓ |
| 68. pprof-goroutine-debug-levels — ?debug=1 and ?debug=2 for goroutine dumps | 4/4 | 3/4 | |
| 68.1 | curl with ?debug=1 for human-readable dump | ✓ | ✓ |
| 68.2 | ?debug=2 for full stacks with creation site and labels | ✓ | ✗ |
| 68.3 | Shows URL: /debug/pprof/goroutine?debug=1 or ?debug=2 | ✓ | ✓ |
| 68.4 | No go tool pprof needed for debug mode dumps | ✓ | ✓ |
| 69. ci-threshold-calibration — regression thresholds by CI environment | 5/5 | 0/5 | |
| 69.1 | 20%+ threshold on shared/GitHub-hosted runners | ✓ | ✗ |
| 69.2 | 10% on dedicated self-hosted runners | ✓ | ✗ |
| 69.3 | Tight thresholds on noisy environments = false positives eroding trust | ✓ | ✗ |
| 69.4 | GitHub-hosted runners ~2-3% CoV in best case | ✓ | ✗ |
| 69.5 | <1% false positive rate requires 7%+ gate | ✓ | ✗ |
| 70. cpu-profile-sampling-rate — CPU profiler 100Hz sampling mechanism | 4/4 | 4/4 | |
| 70.1 | CPU profiling uses statistical sampling at 100Hz | ✓ | ✓ |
| 70.2 | Functions shorter than sampling interval may not appear | ✓ | ✓ |
| 70.3 | Only captures on-CPU time, off-CPU invisible | ✓ | ✓ |
| 70.4 | ~5% overhead during CPU profile capture | ✓ | ✓ |
| 71. benchstat-ignore-dimension-warning — -ignore flag for suppressing warnings | 4/4 | 1/4 | |
| 71.1 | -ignore /gomaxprocs to suppress warning | ✓ | ✗ |
| 71.2 | -ignore omits keys from grouping | ✓ | ✗ |
| 71.3 | -row .name to simplify row grouping | ✓ | ✗ |
| 71.4 | -col /gomaxprocs to compare across values instead | ✓ | ✓ |
| 72. escape-analysis-fmt-sprintf — fmt.Sprintf heap allocation explanation | 4/4 | 4/4 | |
| 72.1 | fmt.Sprintf arguments boxed into any/interface{} (boxing) | ✓ | ✓ |
| 72.2 | Result string is also heap-allocated | ✓ | ✓ |
| 72.3 | fmt.Sprintf is common escape cause in compiler analysis | ✓ | ✓ |
| 72.4 | Alternatives for hot paths: strconv, strings.Builder, direct byte manipulation | ✓ | ✓ |
| 73. runtime-readmemstats-vs-runtime-metrics — programmatic GC stats APIs | 5/5 | 4/5 | |
| 73.1 | runtime.ReadMemStats for heap size, NumGC, pause durations | ✓ | ✓ |
| 73.2 | debug.ReadGCStats for GC-specific statistics | ✓ | ✗ |
| 73.3 | runtime/metrics (Go 1.16+) as preferred modern API | ✓ | ✓ |
| 73.4 | runtime/metrics has lower overhead and safe for concurrent reads | ✓ | ✓ |
| 73.5 | ReadMemStats is more expensive due to internal locking | ✓ | ✓ |
| 74. pprof-callgrind-export — exporting to KCachegrind format | 4/4 | 3/4 | |
| 74.1 | callgrind command or -callgrind flag to export | ✓ | ✓ |
| 74.2 | Command: go tool pprof -callgrind cpu.prof > cpu.callgrind | ✓ | ✓ |
| 74.3 | KCachegrind or QCachegrind as visualization tools | ✓ | ✓ |
| 74.4 | proto command for protobuf format as alternative | ✓ | ✗ |
| 75. expvar-lightweight-monitoring — stdlib expvar for JSON metrics | 4/4 | 3/4 | |
| 75.1 | Recommends expvar package from stdlib | ✓ | ✓ |
| 75.2 | import _ "expvar" auto-registers at /debug/vars | ✓ | ✓ |
| 75.3 | Serves JSON format | ✓ | ✓ |
| 75.4 | Integration with Netdata, Telegraf, custom dashboards | ✓ | ✗ |
| 76. benchstat-table-flag-per-package — -table pkg for per-package tables | 3/3 | 0/3 | |
| 76.1 | -table pkg flag for one table per package | ✓ | ✗ suggests running benchstat per pkg |
| 76.2 | Command: benchstat -table pkg old.txt new.txt | ✓ | ✗ |
| 76.3 | Default -table is .config (goos/goarch/pkg/cpu) | ✓ | ✗ |
| 77. pprof-symbolization-remote — symbolization modes for remote profiles | 4/4 | 1/4 | |
| 77.1 | Symbolization modes: local, remote, none | ✓ | ✗ |
| 77.2 | -symbolize=local for local binaries | ✓ | ✓ |
| 77.3 | PPROF_BINARY_PATH env var for binary search paths | ✓ | ✗ |
| 77.4 | -symbolize=remote to contact running service | ✓ | ✗ |
| 78. trace-concurrent-with-flight-recorder — trace.Start and FlightRecorder coexistence | 3/3 | 0/3 | |
| 78.1 | Flight recorder can run concurrently with trace.Start | ✓ | ✗ |
| 78.2 | At most one flight recorder active at a time | ✓ | ✗ |
| 78.3 | Both can be active simultaneously without conflict | ✓ | ✗ |
| 79. pyroscope-overhead-warning — continuous profiling overhead at scale | 4/4 | 3/4 | |
| 79.1 | ~2-5% CPU overhead per instance for continuous profiling | ✓ | ✓ |
| 79.2 | 200 instances = significant aggregate compute and storage cost | ✓ | ✓ |
| 79.3 | Enable on subset of instances or on-demand via env var | ✓ | ✓ |
| 79.4 | Investigation session approach: target instances only, not fleet-wide | ✓ | ✗ |
| 80. non-go-memory-leak-detection — PromQL for cgo/mmap leak detection | 4/4 | 1/4 | |
| 80.1 | PromQL: process_resident_memory_bytes - go_memstats_sys_bytes | ✓ | ✗ |
| 80.2 | Gap = non-Go memory (cgo, mmap) | ✓ | ✗ |
| 80.3 | Growing gap = non-Go memory leak | ✓ | ✗ |
| 80.4 | Investigate cgo calls or memory-mapped files | ✓ | ✓ |
golang-observability — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 185/185 (100%) | 117/185 (63%) | +37pp |
Full breakdown (185 assertions across 40 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. summary-vs-histogram-multi-replica — Histogram over Summary when multiple replicas need aggregated percentiles | 5/5 | 5/5 | |
| 1.1 | Recommends Histogram, not Summary | ✓ | ✓ |
| 1.2 | Explains that Summary quantiles cannot be aggregated across multiple instances | ✓ | ✓ |
| 1.3 | Shows histogram_quantile() PromQL function for computing percentiles server-side | ✓ | ✓ |
| 1.4 | Mentions that Histogram supports server-side aggregation across replicas | ✓ | ✓ |
| 1.5 | Uses prometheus.NewHistogramVec (not NewSummary) in the code example | ✓ | ✓ |
| 2. log-and-return-error-trap — single handling rule: log OR return, never both | 5/5 | 5/5 | |
| 2.1 | Identifies the log-and-return pattern as incorrect | ✓ | ✓ |
| 2.2 | Explains that the error gets logged multiple times as it propagates up the chain | ✓ | ✓ |
| 2.3 | Recommends returning the error with context and logging once at the top level | ✓ | ✓ |
| 2.4 | Shows the corrected pattern: return fmt.Errorf with wrapping, no slog call | ✓ | ✓ |
| 2.5 | References or explains the single handling rule (errors are either logged OR returned, never both) | ✓ | ✓ |
| 3. high-cardinality-label-trap — high-cardinality label usage in Prometheus metrics | 6/6 | 4/6 | |
| 3.1 | Identifies userID as a high-cardinality label that will cause problems | ✓ | ✓ |
| 3.2 | Identifies r.URL.Path as potentially high-cardinality (should use route template instead) | ✓ | ✗ focuses on userID, misses URL.Path |
| 3.3 | Explains that each unique label combination creates a separate time series | ✓ | ✓ |
| 3.4 | Warns about memory explosion on the Prometheus server from unbounded labels | ✓ | ✓ |
| 3.5 | Recommends using route patterns/templates (e.g., /users/:id) instead of actual paths | ✓ | ✗ generic advice without route template |
| 3.6 | Suggests using traces (not metrics) for high-cardinality data like user IDs | ✓ | ✓ |
| 4. production-json-logging — JSON handler for production, not TextHandler | 4/4 | 3/4 | |
| 4.1 | Recommends JSONHandler for production, not TextHandler | ✓ | ✓ |
| 4.2 | Explains that plain-text multiline logs (e.g., stack traces) get split into separate records by log collectors | ✓ | ✗ generic "not machine-parseable" |
| 4.3 | Suggests TextHandler is appropriate for development only | ✓ | ✓ |
| 4.4 | Shows the correct JSONHandler setup with slog.LevelInfo for production | ✓ | ✓ |
| 5. slog-context-variant-trace-correlation — *Context variants for trace correlation | 4/4 | 3/4 | |
| 5.1 | Identifies that slog.Info and slog.Error should use their *Context variants | ✓ | ✓ |
| 5.2 | Explains that without ctx, trace_id and span_id won't be injected into log records | ✓ | ✗ unaware of otelslog injection mechanism |
| 5.3 | Shows the corrected code using slog.InfoContext(ctx, ...) and slog.ErrorContext(ctx, ...) | ✓ | ✓ |
| 5.4 | Mentions that the otelslog bridge automatically injects trace correlation when context is passed | ✓ | ✓ |
| 6. metric-naming-conventions — base units, _total suffix, namespace | 5/5 | 5/5 | |
| 6.1 | Flags request_count as missing namespace and unit suffix | ✓ | ✓ |
| 6.2 | Flags httpDuration as using camelCase and missing unit | ✓ | ✓ |
| 6.3 | Flags request_duration_ms — should use _seconds (base unit), not milliseconds | ✓ | ✓ |
| 6.4 | Flags myapp_request_size_kb — should use _bytes (base unit), not kilobytes | ✓ | ✓ |
| 6.5 | Flags embedding label values into metric names — should use a single metric with a method label | ✓ | ✓ |
| 7. irate-vs-rate-for-alerts — irate() is inappropriate for alerting rules | 5/5 | 3/5 | |
| 7.1 | Identifies irate() as inappropriate for alerting rules | ✓ | ✓ |
| 7.2 | Explains that irate reacts to a single scrape interval and is too volatile | ✓ | ✗ knows irate differs but imprecise why |
| 7.3 | Recommends rate() instead of irate() for alerts | ✓ | ✓ |
| 7.4 | Recommends adding a for: duration to avoid firing on transient spikes | ✓ | ✓ |
| 7.5 | Shows the corrected alert rule using rate() with a for: clause | ✓ | ✗ shows rate() but misses error ratio |
| 8. alert-missing-for-duration — alerts without for: duration clause | 4/4 | 3/4 | |
| 8.1 | Identifies the missing for: duration as a problem | ✓ | ✓ |
| 8.2 | Explains that without for:, a single bad scrape triggers the alert | ✓ | ✓ |
| 8.3 | Recommends adding for: 5m or similar duration | ✓ | ✓ |
| 8.4 | Distinguishes that binary alerts (service up/down) can use for: 0m, but non-binary need a duration | ✓ | ✗ generic advice, no binary distinction |
| 9. promql-comments-convention — documenting metrics with PromQL comments above declarations | 5/5 | 0/5 | |
| 9.1 | Recommends adding PromQL queries and alert rules as comments directly above the metric var | ✓ | ✗ suggests wiki or README |
| 9.2 | Shows example Dashboard: and Alert: comment lines above the metric var | ✓ | ✗ no such convention known |
| 9.3 | Explains that this keeps PromQL queries reviewed in PRs alongside the metric | ✓ | ✗ |
| 9.4 | Mentions that queries stay in sync with metric changes (label renames, bucket changes) | ✓ | ✗ |
| 9.5 | Notes that new team members can understand the metric's purpose at a glance | ✓ | ✗ |
| 10. otelslog-bridge-setup — log-trace correlation using otelslog bridge | 5/5 | 1/5 | |
| 10.1 | Recommends using the otelslog bridge from go.opentelemetry.io/contrib/bridges/otelslog | ✓ | ✗ suggests manual trace_id extraction |
| 10.2 | Shows creating a handler with otelslog.NewHandler() | ✓ | ✗ |
| 10.3 | Shows setting it as default with slog.SetDefault() | ✓ | ✗ |
| 10.4 | Explains that trace_id and span_id are automatically injected into log records | ✓ | ✓ |
| 10.5 | Emphasizes using slog.*Context(ctx, ...) variants to enable the automatic injection | ✓ | ✗ |
| 11. exemplars-metric-trace-link — metrics-to-traces correlation via Prometheus exemplars | 3/3 | 1/3 | |
| 11.1 | Recommends using Prometheus exemplars to link metrics to traces | ✓ | ✗ suggests manual correlation |
| 11.2 | Shows attaching trace_id as an exemplar when recording histogram observations | ✓ | ✗ |
| 11.3 | Explains that exemplars let you jump from a metric spike directly to the trace | ✓ | ✓ |
| 12. span-error-recording-both-calls — RecordError() AND SetStatus(Error) required | 4/4 | 3/4 | |
| 12.1 | Identifies that span.SetStatus(codes.Error, ...) is also needed alongside RecordError | ✓ | ✓ |
| 12.2 | Explains that RecordError adds an event but does not mark the span as failed | ✓ | ✗ may not distinguish event vs status |
| 12.3 | Shows the corrected pattern with both span.RecordError(err) and span.SetStatus(codes.Error, ...) | ✓ | ✓ |
| 12.4 | Notes that on success, no status needs to be set (Unset is fine) | ✓ | ✓ |
| 13. otelhttp-outgoing-requests — outgoing HTTP clients must use otelhttp transport | 4/4 | 4/4 | |
| 13.1 | Recommends wrapping the HTTP client transport with otelhttp.NewTransport | ✓ | ✓ |
| 13.2 | Shows the code: client := &http.Client{Transport: otelhttp.NewTransport(http.DefaultTransport)} | ✓ | ✓ |
| 13.3 | Explains that this automatically propagates trace context to outgoing requests | ✓ | ✓ |
| 13.4 | Mentions that otelhttp creates child spans for outgoing HTTP calls | ✓ | ✓ |
| 14. db-query-context-propagation — database calls must use *Context variants | 4/4 | 4/4 | |
| 14.1 | Identifies that db.Query should be db.QueryContext(ctx, ...) | ✓ | ✓ |
| 14.2 | Explains that without context, the trace is broken | ✓ | ✓ |
| 14.3 | Shows the corrected code using db.QueryContext(ctx, ...) | ✓ | ✓ |
| 14.4 | States that context is the vehicle that carries trace_id and span_id across boundaries | ✓ | ✓ |
| 15. trace-sampling-cost-control — trace sampling strategies and cost implications | 5/5 | 4/5 | |
| 15.1 | Recommends TraceIDRatioBased sampling with a specific ratio (e.g., 0.1 for 10%) | ✓ | ✓ |
| 15.2 | Mentions ParentBased sampler to respect parent's sampling decision | ✓ | ✓ |
| 15.3 | Discusses head-based vs tail-based sampling tradeoffs | ✓ | ✓ |
| 15.4 | Recommends avoiding large payloads as span attributes — log and correlate via trace_id | ✓ | ✗ generic cost advice |
| 15.5 | Explains the cost factors: span volume, span attributes, storage and indexing | ✓ | ✓ |
| 16. where-to-add-spans — which operations must have spans in OpenTelemetry | 5/5 | 4/5 | |
| 16.1 | Lists service methods (business logic layer) as requiring spans | ✓ | ✓ |
| 16.2 | Lists database queries as requiring spans | ✓ | ✓ |
| 16.3 | Lists external API calls as requiring spans | ✓ | ✓ |
| 16.4 | Lists message queue publish/consume operations as requiring spans | ✓ | ✓ |
| 16.5 | States any operation that takes measurable time or could fail should have a span | ✓ | ✗ gives specific list, misses general rule |
| 17. four-golden-signals-alerting — four golden signals for service alerting | 5/5 | 5/5 | |
| 17.1 | References the four golden signals: latency, traffic, errors, saturation | ✓ | ✓ |
| 17.2 | Includes a latency alert (e.g., P99 > threshold) | ✓ | ✓ |
| 17.3 | Includes a traffic alert (e.g., zero requests detection) | ✓ | ✓ |
| 17.4 | Includes an error rate alert (e.g., 5xx ratio > threshold) | ✓ | ✓ |
| 17.5 | Includes a saturation alert (e.g., connection pool > 90%) | ✓ | ✓ |
| 18. awesome-prometheus-alerts-resource — awesome-prometheus-alerts as a starting point | 4/4 | 0/4 | |
| 18.1 | Recommends awesome-prometheus-alerts (samber.github.io/awesome-prometheus-alerts/) | ✓ | ✗ suggests writing rules from scratch |
| 18.2 | Mentions it contains ~500 ready-to-use Prometheus alerting rules organized by technology | ✓ | ✗ |
| 18.3 | Suggests the workflow: browse by technology, copy rules, customize thresholds | ✓ | ✗ |
| 18.4 | Mentions verifying that exporters (postgres_exporter, redis_exporter) are deployed | ✓ | ✗ |
| 19. go-runtime-alerts — Go runtime-specific Prometheus alerts | 5/5 | 3/5 | |
| 19.1 | Suggests alerting on go_goroutines exceeding a threshold for goroutine leaks | ✓ | ✓ |
| 19.2 | Suggests alerting on go_gc_duration_seconds for GC pressure | ✓ | ✓ |
| 19.3 | Suggests alerting on go_memstats_alloc_bytes / go_memstats_sys_bytes for memory leaks | ✓ | ✓ |
| 19.4 | Suggests alerting on go_threads for high OS thread count | ✓ | ✗ rarely includes thread alerts |
| 19.5 | Uses for: duration on all non-binary alerts to avoid false positives | ✓ | ✗ omits for: on some alerts |
| 20. alert-severity-levels — severity classification and for: durations | 6/6 | 4/6 | |
| 20.1 | Uses two severity levels: critical (page on-call) and warning (create ticket) | ✓ | ✓ |
| 20.2 | Critical alerts: for: 2m to 5m for fast detection | ✓ | ✓ |
| 20.3 | Warning alerts: for: 10m to 30m for confirmed trends | ✓ | ✓ |
| 20.4 | Classifies service down as critical with short for: duration | ✓ | ✓ |
| 20.5 | Classifies goroutine leak as warning (not critical) | ✓ | ✗ may classify as critical |
| 20.6 | States that for: 0m should never be used on non-binary alerts | ✓ | ✗ not a widely known rule |
| 21. multi-window-burn-rate-slo — multi-window burn-rate SLO alerting | 5/5 | 3/5 | |
| 21.1 | Recommends multi-window burn-rate alerting instead of simple threshold alerts | ✓ | ✓ |
| 21.2 | Explains the concept of error budget and burn rate | ✓ | ✓ |
| 21.3 | Includes fast burn window (e.g., 5m + 1h, 14.4x burn rate) as critical/page | ✓ | ✗ generic burn rate, no specific windows |
| 21.4 | Includes slow burn window (e.g., 2h + 24h, 1x burn rate) as warning/ticket | ✓ | ✗ no slow burn tier |
| 21.5 | Shows PromQL using AND of short and long windows to eliminate false positives | ✓ | ✓ |
| 22. slog-migration-from-zap — incremental migration from zap to slog using bridge | 5/5 | 1/5 | |
| 22.1 | Recommends a three-step migration: bridge, replace call sites, remove bridge | ✓ | ✗ suggests gradual replacement only |
| 22.2 | Step 1: Use samber/slog-zap bridge handler to route slog output through zap | ✓ | ✗ unaware of samber/slog-zap |
| 22.3 | Step 2: Gradually replace zap.L().Info(...) calls with slog.Info(...) | ✓ | ✓ |
| 22.4 | Step 3: Once fully migrated, replace bridge with native slog JSONHandler and remove zap | ✓ | ✗ no bridge step to remove |
| 22.5 | Mentions using parallel sub-agents for large codebase migration | ✓ | ✗ not a known pattern |
| 23. slog-migration-from-logrus — bridge handler approach for logrus migration | 4/4 | 2/4 | |
| 23.1 | Recommends using samber/slog-logrus bridge handler for incremental migration | ✓ | ✗ unaware of samber/slog-logrus |
| 23.2 | Explains that slog is the standard library logger since Go 1.21 | ✓ | ✓ |
| 23.3 | Shows the bridge step: route slog output through the existing logrus logger | ✓ | ✗ |
| 23.4 | Shows the replacement: logrus.WithField becomes slog.Info | ✓ | ✓ |
| 24. debug-level-production-cost — log level cost implications in production | 5/5 | 3/5 | |
| 24.1 | Recommends slog.LevelInfo for production, NOT Debug | ✓ | ✓ |
| 24.2 | Explains that Debug level can generate millions of log lines per minute in busy services | ✓ | ✓ |
| 24.3 | Mentions cost: CPU for serialization, I/O for disk/network, money for log ingestion/storage | ✓ | ✓ |
| 24.4 | Mentions Debug can inflate costs by 10-100x | ✓ | ✗ no specific multiplier |
| 24.5 | Suggests samber/slog-sampling as an alternative to sample verbose logs | ✓ | ✗ unaware of samber/slog-sampling |
| 25. pii-in-logs-trap — catching PII being logged | 4/4 | 4/4 | |
| 25.1 | Flags email as PII that should not be logged | ✓ | ✓ |
| 25.2 | Flags SSN as PII that should absolutely never be logged | ✓ | ✓ |
| 25.3 | Recommends logging identifiers (user_id) instead of PII | ✓ | ✓ |
| 25.4 | Shows corrected logging using user.ID instead of email/SSN | ✓ | ✓ |
| 26. counter-suffix-total-requirement — Prometheus counters must use _total suffix | 3/3 | 3/3 | |
| 26.1 | Identifies the missing _total suffix — counters MUST end with _total | ✓ | ✓ |
| 26.2 | Shows the corrected name: requests_total or http_requests_total | ✓ | ✓ |
| 26.3 | Mentions that _total is a required convention for counters in Prometheus | ✓ | ✓ |
| 27. pprof-security-auth — pprof endpoints must be protected with authentication | 4/4 | 4/4 | |
| 27.1 | Warns that pprof endpoints must NOT be exposed publicly without authentication | ✓ | ✓ |
| 27.2 | Explains that pprof leaks sensitive runtime information and can be abused for DoS | ✓ | ✓ |
| 27.3 | Recommends protecting with basic auth or running on a separate internal port | ✓ | ✓ |
| 27.4 | Suggests toggling via environment variable to enable/disable without redeployment | ✓ | ✓ |
| 28. continuous-profiling-env-toggle — toggle continuous profiling via environment variables | 5/5 | 2/5 | |
| 28.1 | Recommends toggling via environment variable (e.g., PROFILING_ENABLED) | ✓ | ✓ |
| 28.2 | Mentions ~2-5% CPU overhead for continuous profiling | ✓ | ✗ no specific overhead figure |
| 28.3 | Suggests starting with CPU + heap profiles only, adding mutex/block when needed | ✓ | ✗ suggests all profiles |
| 28.4 | For large deployments, recommends enabling on a fraction of replicas (e.g., 1 in 10) | ✓ | ✗ not a widely known practice |
| 28.5 | Shows code that checks the environment variable before starting Pyroscope | ✓ | ✓ |
| 29. rum-identity-key-email-trap — RUM distinct_id must be user_id, not email | 5/5 | 4/5 | |
| 29.1 | Rejects email as the DistinctId — must use user_id instead | ✓ | ✓ |
| 29.2 | Explains that email is mutable — users change it, splitting events into two users | ✓ | ✓ |
| 29.3 | Explains that email is PII, complicating GDPR/CCPA compliance | ✓ | ✓ |
| 29.4 | Notes that email leaks into third-party analytics systems as the identity key | ✓ | ✗ not commonly considered |
| 29.5 | Shows corrected code using user.ID (immutable internal identifier) | ✓ | ✓ |
| 30. gdpr-consent-before-tracking — GDPR consent must be checked before sending events | 5/5 | 3/5 | |
| 30.1 | Identifies that consent must be checked before sending the tracking event | ✓ | ✓ |
| 30.2 | Shows extracting consent from context and conditionally tracking | ✓ | ✗ generic consent advice without Go code |
| 30.3 | Mentions GDPR fines (up to 4% of global revenue) or CCPA penalties | ✓ | ✓ |
| 30.4 | References data minimization — only collect what you need | ✓ | ✓ |
| 30.5 | Mentions data subject rights endpoints (data export and deletion) | ✓ | ✗ focuses on consent only |
| 31. data-subject-rights-endpoints — GDPR requires deletion/export across all systems | 5/5 | 3/5 | |
| 31.1 | States that deletion must propagate to ALL systems holding user data | ✓ | ✓ |
| 31.2 | Lists the analytics platform (PostHog) as needing deletion | ✓ | ✓ |
| 31.3 | Lists the CDP (Segment) as needing deletion | ✓ | ✗ may not mention CDP deletion |
| 31.4 | References GDPR Article 17 Right to Erasure | ✓ | ✓ |
| 31.5 | Also mentions the Right of Access (data export endpoint) as a requirement | ✓ | ✗ focuses on deletion only |
| 32. five-signals-completeness — five observability signals and their distinct roles | 6/6 | 4/6 | |
| 32.1 | Lists all five signals: logs, metrics, traces, profiles, and RUM | ✓ | ✗ lists logs, metrics, traces only |
| 32.2 | Associates logs with 'what happened' (discrete events, audit trails) | ✓ | ✓ |
| 32.3 | Associates metrics with 'how much/how fast' (aggregated measurements, alerting, SLOs) | ✓ | ✓ |
| 32.4 | Associates traces with 'where did time go' (request flow across services) | ✓ | ✓ |
| 32.5 | Associates profiles with 'why is it slow/using memory' (CPU hotspots, memory leaks) | ✓ | ✓ |
| 32.6 | Associates RUM with 'how do users experience it' (product analytics, funnels) | ✓ | ✗ RUM not listed as a signal |
| 33. definition-of-done-observability — observability checklist before shipping a feature | 6/6 | 4/6 | |
| 33.1 | States that a feature is not production-ready until it is observable | ✓ | ✓ |
| 33.2 | Checks for metric declarations (counters, histograms, gauges) with PromQL comments | ✓ | ✗ does not mention PromQL comments |
| 33.3 | Checks for proper structured logging with slog and context variants | ✓ | ✓ |
| 33.4 | Checks for OpenTelemetry spans on service methods, DB queries, and external calls | ✓ | ✓ |
| 33.5 | Checks for dashboards and alerts being wired up | ✓ | ✓ |
| 33.6 | Checks that errors are either logged OR returned, never both | ✓ | ✗ not part of typical deploy checklist |
| 34. grafana-dashboard-ids — specific Grafana dashboard IDs for Go runtime monitoring | 4/4 | 0/4 | |
| 34.1 | Recommends specific Grafana dashboard IDs (21221, 6671, or 10826) | ✓ | ✗ suggests building custom dashboards |
| 34.2 | Mentions dashboard 21221 for host + runtime combined view | ✓ | ✗ |
| 34.3 | Explains that these dashboards use default Go collector metrics from the Prometheus client library | ✓ | ✗ |
| 34.4 | Shows how to import: Dashboards > New > Import, enter the dashboard ID | ✓ | ✗ |
| 35. slog-with-request-scoped-attrs — slog.With() for request-scoped attributes | 4/4 | 4/4 | |
| 35.1 | Recommends using slog.With() to create a child logger with request-scoped attributes | ✓ | ✓ |
| 35.2 | Shows middleware pattern that creates the enriched logger | ✓ | ✓ |
| 35.3 | Shows storing the enriched logger in context for downstream use | ✓ | ✓ |
| 35.4 | Includes request_id, method, and path as the attributes to inject | ✓ | ✓ |
| 36. slog-ecosystem-handlers — slog handler ecosystem beyond stdlib | 5/5 | 1/5 | |
| 36.1 | Recommends samber/slog-multi for fan-out to multiple handlers | ✓ | ✗ suggests custom io.MultiWriter |
| 36.2 | Mentions samber/slog-sentry for sending errors to Sentry | ✓ | ✗ |
| 36.3 | Mentions samber/slog-datadog for sending logs to Datadog | ✓ | ✗ |
| 36.4 | Explains that slog supports pluggable handlers | ✓ | ✓ |
| 36.5 | References the slog ecosystem (go.dev/wiki/Resources-for-slog or similar) | ✓ | ✗ |
| 37. parallel-observability-audit — parallel sub-agents for observability audits | 5/5 | 0/5 | |
| 37.1 | Recommends using up to 5 parallel sub-agents (via the Agent tool) | ✓ | ✗ suggests linear approach |
| 37.2 | Assigns one sub-agent per signal: metrics, logging, tracing, profiling, RUM | ✓ | ✗ |
| 37.3 | Sub-agent for metrics: verify metric declarations and PromQL comments | ✓ | ✗ |
| 37.4 | Sub-agent for logging: check structured logging, PII in logs, error logging patterns | ✓ | ✗ |
| 37.5 | Sub-agent for tracing: verify span creation in service methods, DB calls, API calls | ✓ | ✗ |
| 38. predict-linear-for-saturation — predict_linear for anticipating resource exhaustion | 4/4 | 4/4 | |
| 38.1 | Recommends using predict_linear() PromQL function to extrapolate trends | ✓ | ✓ |
| 38.2 | Shows expression: predict_linear(db_connections_active[15m], 600) > db_connections_max | ✓ | ✓ |
| 38.3 | Explains that predict_linear extrapolates from recent trend to predict future value | ✓ | ✓ |
| 38.4 | Also suggests a threshold alert (e.g., > 90%) as a complementary alert | ✓ | ✓ |
| 39. self-hosted-rum-gdpr — self-hosted analytics for GDPR compliance | 4/4 | 4/4 | |
| 39.1 | Recommends self-hosted analytics (PostHog or Matomo) for EU data residency | ✓ | ✓ |
| 39.2 | Explains that self-hosting eliminates cross-border data transfer concerns | ✓ | ✓ |
| 39.3 | Compares self-hosted vs SaaS tradeoffs (data residency, cost, maintenance, features) | ✓ | ✓ |
| 39.4 | Mentions that PostHog can be self-hosted to keep data in your own infrastructure | ✓ | ✓ |
| 40. oops-structured-errors-tracing — samber/oops for structured errors in tracing | 4/4 | 0/4 | |
| 40.1 | Recommends samber/oops for structured errors with stack traces | ✓ | ✗ suggests fmt.Errorf or manual attrs |
| 40.2 | Shows using oops with .In(), .Code(), and .With() for structured context | ✓ | ✗ |
| 40.3 | Explains that oops errors carry stack trace, structured context, and work with span.RecordError() | ✓ | ✗ |
| 40.4 | Mentions compatibility with errors.Is/errors.As and slog | ✓ | ✗ |
golang-performance — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 272/272 (100%) | 167/272 (61%) | +39pp |
Full breakdown (272 assertions across 68 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. profile-before-optimizing — Tests whether the model insists on profiling before applying optimizations | 5/5 | 3/5 | |
| 1.1 | Recommends profiling (pprof, fgprof, or tracing) before making code changes | ✓ | ✓ |
| 1.2 | Identifies fetchFromDB as the likely bottleneck (external I/O, not Go code) | ✓ | ✓ |
| 1.3 | Mentions that intuition about bottlenecks is often wrong (~80% of the time) | ✓ | ✗ no specific stat |
| 1.4 | Does NOT primarily focus on micro-optimizing strings.ToUpper or JSON encoding | ✓ | ✗ suggests strings.Builder/JSON fixes |
| 1.5 | Suggests investigating the database query (query tuning, caching, connection pool) | ✓ | ✓ |
| 2. fgprof-off-cpu-bottleneck — Tests whether the model recommends fgprof for off-CPU bottlenecks | 5/5 | 4/5 | |
| 2.1 | Recommends fgprof as the primary tool for capturing off-CPU wait time | ✓ | ✗ suggests goroutine profiles or tracing |
| 2.2 | Explains that standard pprof CPU profile only captures on-CPU time | ✓ | ✓ |
| 2.3 | Suggests the bottleneck is likely I/O wait (network, database, filesystem) | ✓ | ✓ |
| 2.4 | Mentions goroutine profile as a complementary diagnostic | ✓ | ✓ |
| 2.5 | Suggests distributed tracing (OpenTelemetry) for identifying slow upstream services | ✓ | ✓ |
| 3. iterative-benchmark-methodology — Tests iterative benchmark approach (one change at a time, benchstat) | 5/5 | 1/5 | |
| 3.1 | Recommends writing an atomic benchmark for ProcessRecords first | ✓ | ✓ |
| 3.2 | Recommends measuring a baseline with -benchmem and -count=6 for statistical significance | ✓ | ✗ skips -count=6 |
| 3.3 | Recommends applying ONE optimization at a time, not all at once | ✓ | ✗ applies all at once |
| 3.4 | Recommends using benchstat to compare before/after with statistical significance | ✓ | ✗ no benchstat |
| 3.5 | Suggests keeping report files as an audit trail (e.g., /tmp/report-1.txt) | ✓ | ✗ |
| 4. slice-reuse-append-zero — Tests knowledge of append(s[:0], ...) for reusing slice backing arrays | 4/4 | 2/4 | |
| 4.1 | Suggests using append(mode[:0], item) to reuse the backing array | ✓ | ✗ suggests sync.Pool or new slice |
| 4.2 | Explains that reslicing to zero length retains the backing array | ✓ | ✗ |
| 4.3 | Moves the mode variable declaration outside the loop to enable reuse | ✓ | ✓ |
| 4.4 | Does NOT suggest sync.Pool as the primary solution for this simple case | ✓ | ✓ |
| 5. direct-indexing-vs-append — Tests direct indexing over append when output size equals input | 3/3 | 2/3 | |
| 5.1 | Suggests using make([]Result, len(input)) with direct assignment result[i] = convert(...) | ✓ | ✓ |
| 5.2 | Explains that direct assignment avoids per-element append overhead (bounds check, length increment) | ✓ | ✗ |
| 5.3 | Notes that append is better when the result might be smaller (filtering) | ✓ | ✓ |
| 6. map-range-double-lookup — Tests catching redundant map lookups in range loops | 2/2 | 2/2 | |
| 6.1 | Identifies that for k := range in { in[k] } does two map lookups per iteration | ✓ | ✓ |
| 6.2 | Suggests for k, v := range in { result[k] = strconv.Itoa(v) } for single lookup | ✓ | ✓ |
| 7. sentinel-errors-hot-path — Tests preallocated sentinel errors over fmt.Errorf in hot paths | 4/4 | 4/4 | |
| 7.1 | Converts the static error to a preallocated sentinel using errors.New at package level | ✓ | ✓ |
| 7.2 | Keeps fmt.Errorf for the dynamic error that includes %d | ✓ | ✓ |
| 7.3 | Explains that fmt.Errorf allocates on every call, while sentinels allocate once | ✓ | ✓ |
| 7.4 | Does NOT convert the dynamic error to a sentinel | ✓ | ✓ |
| 8. interface-boxing-hot-path — Tests interface boxing allocation cost and generics fix | 4/4 | 3/4 | |
| 8.1 | Identifies interface boxing (any parameter) as the source of allocations | ✓ | ✓ |
| 8.2 | Suggests typed functions or generics to eliminate boxing | ✓ | ✓ |
| 8.3 | Explains that each concrete value passed through any requires a heap allocation for boxing | ✓ | ✗ |
| 8.4 | Does NOT focus only on the type switch as the optimization target | ✓ | ✓ |
| 9. backing-array-leak-slice — Tests backing array retention when returning a small reslice | 4/4 | 4/4 | |
| 9.1 | Identifies that message[:32] retains the entire backing array (1-10MB) | ✓ | ✓ |
| 9.2 | Suggests using copy() to create an independent 32-byte slice | ✓ | ✓ |
| 9.3 | Explains that the original large array cannot be GC'd while the reslice exists | ✓ | ✓ |
| 9.4 | Provides correct copy-based fix: make([]byte, 32) + copy(header, message[:32]) | ✓ | ✓ |
| 10. substring-memory-leak — Tests strings.Clone pattern for substring memory leaks | 3/3 | 2/3 | |
| 10.1 | Identifies that substrings share the backing array of the original string | ✓ | ✓ |
| 10.2 | Suggests strings.Clone(line[12:48]) to create an independent copy | ✓ | ✗ suggests string([]byte(s)) |
| 10.3 | Explains that each 36-char ID retains the entire 500-2000 byte log line | ✓ | ✓ |
| 11. map-never-shrinks — Tests Go maps never release bucket memory and compact pattern | 4/4 | 4/4 | |
| 11.1 | Explains that Go maps never release bucket memory when entries are deleted | ✓ | ✓ |
| 11.2 | Suggests periodically recreating the map by copying entries to a new map | ✓ | ✓ |
| 11.3 | Provides the compact pattern: make(map[K]V, len(old)) + range copy | ✓ | ✓ |
| 11.4 | Does NOT suggest that calling runtime.GC() or tuning GOGC will fix this | ✓ | ✓ |
| 12. sync-pool-rules — Tests proper sync.Pool usage: reset, size limits, pointer pooling | 5/5 | 2/5 | |
| 12.1 | Identifies that HandleLargeUpload puts oversized buffers (100MB+) back into the pool | ✓ | ✓ |
| 12.2 | Mentions the 32KB guideline — don't pool objects larger than ~32KB | ✓ | ✗ vague size advice |
| 12.3 | Identifies that w.Write(buf) may retain the buffer after bufPool.Put(buf) | ✓ | ✗ |
| 12.4 | Suggests pooling pointers (*[]byte) instead of values to avoid allocation on Get | ✓ | ✗ |
| 12.5 | Recommends resetting/clearing state before Put | ✓ | ✓ |
| 13. struct-field-alignment — Tests struct field ordering for optimal memory layout | 5/5 | 4/5 | |
| 13.1 | Identifies that the struct has wasted padding bytes due to alignment | ✓ | ✓ |
| 13.2 | Suggests reordering fields from largest to smallest | ✓ | ✓ |
| 13.3 | Provides a reordered struct that is smaller than the original | ✓ | ✓ |
| 13.4 | Mentions the fieldalignment tool for automated detection | ✓ | ✗ |
| 13.5 | States alignment requirements (bool=1, int32=4, int64/float64=8) | ✓ | ✓ |
| 14. zero-size-field-end-of-struct — Tests struct{} at end of struct adds word-sized padding | 4/4 | 0/4 | |
| 14.1 | Explains that a zero-size field at the end of a struct causes word-sized padding | ✓ | ✗ |
| 14.2 | Explains the reason: preventing pointer overlap with next memory block | ✓ | ✗ |
| 14.3 | Suggests moving struct{} to the beginning of the struct | ✓ | ✗ |
| 14.4 | Shows the fix: type Entry struct { Flag struct{}; Value int64 } = 8 bytes | ✓ | ✗ |
| 15. map-pointer-vs-value-tradeoff — Tests map[K]*V vs map[K]V tradeoff for large structs | 5/5 | 3/5 | |
| 15.1 | Suggests using map[string]*Player to allow direct field modification | ✓ | ✓ |
| 15.2 | Explains that map values are not addressable | ✓ | ✓ |
| 15.3 | Shows players[id].Score += delta with pointer map | ✓ | ✓ |
| 15.4 | Mentions the tradeoff: pointer maps add GC pressure from separate heap allocations | ✓ | ✗ no GC tradeoff |
| 15.5 | Notes that for small, mostly-read structs, map[K]V (value) is better | ✓ | ✗ |
| 16. inlining-log-in-hot-path — Tests that log calls prevent function inlining | 4/4 | 2/4 | |
| 16.1 | Identifies that log calls prevent the function from being inlined | ✓ | ✗ focuses on formatting cost |
| 16.2 | Suggests removing log calls from the hot-path function | ✓ | ✓ |
| 16.3 | Mentions using go build -gcflags="-m" to verify inlining decisions | ✓ | ✗ |
| 16.4 | Explains that function call overhead matters when called millions of times | ✓ | ✓ |
| 17. value-receiver-inlining — Tests value receivers enable inlining for fluent chains | 4/4 | 0/4 | |
| 17.1 | Suggests changing to value receivers instead of pointer receivers | ✓ | ✗ |
| 17.2 | Explains that value receivers allow the compiler to inline the fluent chain | ✓ | ✗ |
| 17.3 | Explains that pointer receivers add indirection that blocks inlining | ✓ | ✗ |
| 17.4 | Shows the value receiver signature: func (c Config) WithTimeout(d time.Duration) Config | ✓ | ✗ |
| 18. cache-locality-matrix-traversal — Tests row-major vs column-major traversal and cache effects | 5/5 | 4/5 | |
| 18.1 | Identifies the column-first traversal as the cause (cache misses) | ✓ | ✓ |
| 18.2 | Explains that Go stores 2D arrays in row-major order | ✓ | ✓ |
| 18.3 | Suggests swapping loop order to row-first | ✓ | ✓ |
| 18.4 | Mentions the performance difference from cache effects (10-50x or similar) | ✓ | ✗ says "significant" without quantifying |
| 18.5 | Does NOT primarily suggest parallelism or SIMD as the first fix | ✓ | ✓ |
| 19. contiguous-2d-allocation — Tests contiguous 2D allocation for cache-friendly matrix access | 4/4 | 4/4 | |
| 19.1 | Identifies that per-row allocation scatters data across the heap | ✓ | ✓ |
| 19.2 | Suggests single contiguous allocation: make([]float64, rows*cols) | ✓ | ✓ |
| 19.3 | Shows slicing the contiguous array into row views: data[i*cols : (i+1)*cols] | ✓ | ✓ |
| 19.4 | Explains that contiguous memory improves cache locality | ✓ | ✓ |
| 20. soa-vs-aos — Tests Struct of Arrays vs Array of Structs for single-field iteration | 4/4 | 1/4 | |
| 20.1 | Identifies that loading entire Particle structs wastes cache space when only X is needed | ✓ | ✓ |
| 20.2 | Suggests Struct of Arrays (SoA) layout with separate slices | ✓ | ✗ suggests parallelism |
| 20.3 | Explains cache utilization improvement (contiguous X values vs scattered) | ✓ | ✗ |
| 20.4 | Notes that AoS is fine when accessing all fields together or for small structs | ✓ | ✗ |
| 21. false-sharing-concurrent-counters — Tests false sharing and cache-line padding | 5/5 | 0/5 | |
| 21.1 | Identifies false sharing as the cause (fields share same cache line) | ✓ | ✗ suggests mutex or sharding |
| 21.2 | Explains that writes to one field invalidate the cache line for other cores | ✓ | ✗ |
| 21.3 | Suggests cache-line padding (56-byte array between fields) | ✓ | ✗ |
| 21.4 | Mentions the 64-byte cache line size | ✓ | ✗ |
| 21.5 | Notes this should only be applied when profiling confirms contention | ✓ | ✗ |
| 22. ilp-multi-accumulator — Tests instruction-level parallelism with multiple accumulators | 5/5 | 0/5 | |
| 22.1 | Identifies the sequential dependency chain as the bottleneck | ✓ | ✗ suggests goroutines or SIMD |
| 22.2 | Suggests using multiple accumulators (e.g., 4) for ILP | ✓ | ✗ |
| 22.3 | Shows code with 4 independent accumulators summing every 4th element | ✓ | ✗ |
| 22.4 | Handles the remainder elements (when len(data) is not divisible by 4) | ✓ | ✗ |
| 22.5 | Mentions expected 2-4x improvement from ILP | ✓ | ✗ |
| 23. index-based-tree-cache-locality — Tests index-based vs pointer-based data structures | 4/4 | 4/4 | |
| 23.1 | Recommends Option B (index-based) for better cache locality | ✓ | ✓ |
| 23.2 | Explains that pointer-based nodes are scattered across the heap | ✓ | ✓ |
| 23.3 | Explains that index-based nodes are stored in a contiguous array | ✓ | ✓ |
| 23.4 | Mentions CPU cache lines or memory prefetching | ✓ | ✓ |
| 24. tight-loop-scheduler-starvation — Tests tight CPU loops starving the Go scheduler | 5/5 | 2/5 | |
| 24.1 | Explains that tight CPU loops with inlined operations can delay scheduler preemption | ✓ | ✓ |
| 24.2 | Suggests breaking work into batches processed by a non-inlined function call | ✓ | ✗ suggests runtime.Gosched() |
| 24.3 | Mentions //go:noinline as an option to force preemption points | ✓ | ✗ |
| 24.4 | Explains the tradeoff: //go:noinline adds overhead but ensures scheduler fairness | ✓ | ✗ |
| 24.5 | Mentions that Go 1.14+ has async preemption but tight loops can still cause issues | ✓ | ✓ |
| 25. reflect-deepequal-performance — Tests reflect.DeepEqual 50-200x slower than typed comparison | 4/4 | 2/4 | |
| 25.1 | Identifies reflect.DeepEqual as 50-200x slower than typed comparison | ✓ | ✓ |
| 25.2 | Suggests using slices.Equal for the Hosts field | ✓ | ✗ uses manual loop |
| 25.3 | Suggests using maps.Equal for the Settings field | ✓ | ✗ |
| 25.4 | Provides a hand-written typed comparison function | ✓ | ✓ |
| 26. type-switch-vs-repeated-assertions — Tests type switch dispatches in one evaluation | 3/3 | 3/3 | |
| 26.1 | Suggests replacing repeated type assertions with a type switch | ✓ | ✓ |
| 26.2 | Explains that a type switch dispatches in a single evaluation | ✓ | ✓ |
| 26.3 | Shows the switch v := v.(type) { case string: ... } pattern | ✓ | ✓ |
| 27. http-transport-maxidleconnsperhost — Tests default MaxIdleConnsPerHost=2 gotcha | 5/5 | 3/5 | |
| 27.1 | Identifies MaxIdleConnsPerHost defaulting to 2 as the root cause | ✓ | ✓ |
| 27.2 | Suggests configuring http.Transport with higher MaxIdleConnsPerHost | ✓ | ✓ |
| 27.3 | Shows complete Transport configuration with MaxIdleConns, MaxIdleConnsPerHost, MaxConnsPerHost | ✓ | ✗ only sets MaxIdleConnsPerHost |
| 27.4 | Mentions draining resp.Body for connection reuse (io.Copy to io.Discard) | ✓ | ✗ |
| 27.5 | Does NOT suggest a third-party connection pool library as the primary solution | ✓ | ✓ |
| 28. response-body-drain — Tests HTTP connections only reused when body fully read | 3/3 | 3/3 | |
| 28.1 | Identifies that the response body is not being fully read/drained | ✓ | ✓ |
| 28.2 | Explains that connections are only reused when the body is fully consumed | ✓ | ✓ |
| 28.3 | Suggests adding io.Copy(io.Discard, resp.Body) | ✓ | ✓ |
| 29. streaming-vs-readall — Tests streaming vs buffering for large payloads | 4/4 | 3/4 | |
| 29.1 | Identifies io.ReadAll as the cause of OOM | ✓ | ✓ |
| 29.2 | Suggests using io.Copy(w, resp.Body) to stream with constant memory | ✓ | ✓ |
| 29.3 | Mentions the 32KB internal buffer of io.Copy | ✓ | ✗ |
| 29.4 | Notes that io.ReadAll is fine for small, bounded payloads (< 1MB) | ✓ | ✓ |
| 30. json-streaming-decoder — Tests json.NewDecoder for streaming large JSON payloads | 3/3 | 2/3 | |
| 30.1 | Suggests using json.NewDecoder with r.Body directly | ✓ | ✓ |
| 30.2 | Shows the dec.More() + dec.Decode(&item) streaming pattern | ✓ | ✗ |
| 30.3 | Explains that this processes one item at a time with O(1) memory per item | ✓ | ✓ |
| 31. cgo-overhead-tight-loop — Tests cgo call overhead (~50-100ns) and batching strategy | 4/4 | 2/4 | |
| 31.1 | Identifies cgo overhead (~50-100ns per call) as the bottleneck | ✓ | ✗ says "overhead" without quantifying |
| 31.2 | Suggests using math.Sqrt (pure Go, inlineable) instead of C.sqrt | ✓ | ✓ |
| 31.3 | For unavoidable C code, suggests batching: pass entire array to C in one call | ✓ | ✓ |
| 31.4 | Mentions that goroutine is pinned to OS thread during cgo calls | ✓ | ✗ |
| 32. gogc-gomemlimit-container — Tests GOMEMLIMIT for containerized applications | 5/5 | 3/5 | |
| 32.1 | Recommends setting GOMEMLIMIT to 80-90% of container memory (400-450MiB) | ✓ | ✗ suggests GOMEMLIMIT without ratio |
| 32.2 | Explains that the GC needs GOMEMLIMIT to know about the container ceiling | ✓ | ✓ |
| 32.3 | Shows the GOMEMLIMIT=450MiB environment variable or debug.SetMemoryLimit | ✓ | ✓ |
| 32.4 | Explains the gap (goroutine stacks, OS buffers, non-heap memory) | ✓ | ✗ |
| 32.5 | Does NOT recommend the ballast pattern (obsolete since Go 1.19) | ✓ | ✓ |
| 33. ballast-pattern-obsolete — Tests GOMEMLIMIT over ballast pattern | 4/4 | 4/4 | |
| 33.1 | Identifies the ballast pattern as obsolete since Go 1.19 | ✓ | ✓ |
| 33.2 | Recommends GOMEMLIMIT as the replacement | ✓ | ✓ |
| 33.3 | Explains that GOMEMLIMIT provides the same benefit without wasting physical memory | ✓ | ✓ |
| 33.4 | Shows the GOMEMLIMIT environment variable or debug.SetMemoryLimit call | ✓ | ✓ |
| 34. gomaxprocs-container-go125 — Tests Go 1.25+ container-aware GOMAXPROCS vs automaxprocs | 4/4 | 1/4 | |
| 34.1 | States that Go 1.25+ automatically detects container CPU limits | ✓ | ✗ |
| 34.2 | Recommends removing the automaxprocs dependency | ✓ | ✗ keeps automaxprocs |
| 34.3 | Mentions that automaxprocs IS needed for Go 1.24 and earlier | ✓ | ✓ |
| 34.4 | Mentions cgroup CPU quota detection as the mechanism | ✓ | ✗ |
| 35. pgo-workflow — Tests PGO workflow and expected gains | 5/5 | 3/5 | |
| 35.1 | Recommends Profile-Guided Optimization (PGO) | ✓ | ✓ |
| 35.2 | Describes the workflow: collect profile, save as default.pgo, rebuild | ✓ | ✓ |
| 35.3 | Mentions expected improvement of 2-7% | ✓ | ✗ no quantification |
| 35.4 | Explains PGO benefits: aggressive inlining and devirtualization | ✓ | ✓ |
| 35.5 | Notes that profiles should be refreshed after significant code changes | ✓ | ✗ |
| 36. slog-logattrs-hot-path — Tests slog.LogAttrs for zero-allocation logging | 4/4 | 1/4 | |
| 36.1 | Explains that log arguments are evaluated/boxed before the level check | ✓ | ✓ |
| 36.2 | Recommends slog.LogAttrs for zero allocations when level is disabled | ✓ | ✗ suggests level check guard |
| 36.3 | Shows typed attributes: slog.Int("id", item.ID), slog.String("name", item.Name) | ✓ | ✗ |
| 36.4 | Notes that slog.Any can still allocate, so typed attributes are preferred | ✓ | ✗ |
| 37. regexp-compile-per-call — Tests compiled pattern caching vs per-call compilation | 4/4 | 3/4 | |
| 37.1 | Identifies that regexp compilation happens on every call (~5,700ns per compile) | ✓ | ✓ |
| 37.2 | Suggests moving regexp.MustCompile to package-level variables | ✓ | ✓ |
| 37.3 | Notes that compiled regexps are safe for concurrent use | ✓ | ✓ |
| 37.4 | Quantifies the waste (10-12x overhead from recompilation vs match-only) | ✓ | ✗ |
| 38. singleflight-cache-stampede — Tests singleflight for cache stampede prevention | 4/4 | 4/4 | |
| 38.1 | Identifies this as a cache stampede problem | ✓ | ✓ |
| 38.2 | Recommends golang.org/x/sync/singleflight.Group | ✓ | ✓ |
| 38.3 | Shows sf.Do(key, func) pattern | ✓ | ✓ |
| 38.4 | Does NOT suggest a global mutex as the primary solution | ✓ | ✓ |
| 39. algorithmic-complexity-slice-contains-loop — Tests O(n*m) from slices.Contains in loop | 4/4 | 3/4 | |
| 39.1 | Identifies O(n*m) complexity from slices.Contains inside a loop | ✓ | ✓ |
| 39.2 | Suggests building a map[string]struct{} from the valid slice first | ✓ | ✓ |
| 39.3 | Shows the O(n+m) solution with map lookup | ✓ | ✓ |
| 39.4 | Uses struct{} (0 bytes) for the map value type, not bool | ✓ | ✗ uses bool |
| 40. early-return-full-scan — Tests early returns to avoid full collection scans | 3/3 | 3/3 | |
| 40.1 | Identifies that the function always scans the full collection | ✓ | ✓ |
| 40.2 | Adds an early return when the first expired session is found | ✓ | ✓ |
| 40.3 | Removes the found variable in favor of direct returns | ✓ | ✓ |
| 41. iterator-chain-vs-direct-loop — Tests iterator chains create closure overhead | 4/4 | 4/4 | |
| 41.1 | Recommends Option B (direct loop) for performance | ✓ | ✓ |
| 41.2 | Explains that iterator chains create closures and intermediate allocations | ✓ | ✓ |
| 41.3 | Notes that Filter processes ALL elements before First can pick one | ✓ | ✓ |
| 41.4 | Acknowledges that iterator chains are fine for non-hot paths | ✓ | ✓ |
| 42. indirect-function-calls-closure — Tests closure indirection prevents inlining | 4/4 | 2/4 | |
| 42.1 | Suggests replacing the Map+closure pattern with a direct loop | ✓ | ✓ |
| 42.2 | Explains that the closure/function call indirection prevents inlining | ✓ | ✗ |
| 42.3 | Shows the direct loop: result[i] = *ptrs[i] | ✓ | ✓ |
| 42.4 | Mentions the expected improvement range (13-17%) | ✓ | ✗ |
| 43. http-server-no-timeouts — Tests zero-value http.Server has NO timeouts | 4/4 | 4/4 | |
| 43.1 | Identifies that the default http.Server has NO timeouts | ✓ | ✓ |
| 43.2 | Mentions Slowloris attack or slow client holding connections indefinitely | ✓ | ✓ |
| 43.3 | Suggests setting ReadTimeout, WriteTimeout, and IdleTimeout | ✓ | ✓ |
| 43.4 | Shows creating an explicit http.Server struct with timeout values | ✓ | ✓ |
| 44. http-keepalive-crawler — Tests disabling keep-alive for crawlers | 3/3 | 1/3 | |
| 44.1 | Identifies that idle connections accumulate across many different hosts | ✓ | ✓ |
| 44.2 | Suggests DisableKeepAlives: true for the crawler client | ✓ | ✗ lowers MaxIdleConnsPerHost instead |
| 44.3 | Explains that keep-alive is counterproductive for many unique hosts | ✓ | ✗ |
| 45. buffered-io-syscall-reduction — Tests bufio for reducing syscall count | 4/4 | 4/4 | |
| 45.1 | Identifies that each WriteString issues a separate syscall | ✓ | ✓ |
| 45.2 | Suggests using bufio.NewWriter(f) to batch writes | ✓ | ✓ |
| 45.3 | Includes w.Flush() at the end | ✓ | ✓ |
| 45.4 | Does NOT primarily suggest strings.Builder | ✓ | ✓ |
| 46. concurrent-pipeline-when-not-to-use — Tests when concurrent pipelines are NOT beneficial | 5/5 | 1/5 | |
| 46.1 | Recommends AGAINST concurrent pipelines for this case | ✓ | ✗ recommends concurrency |
| 46.2 | Explains that all three stages compete for the same resource (CPU) | ✓ | ✗ |
| 46.3 | Notes that concurrency only helps when stages saturate DIFFERENT resources | ✓ | ✗ |
| 46.4 | Mentions context-switching overhead as a cost | ✓ | ✓ |
| 46.5 | Suggests sequential processing or batching as simpler alternative | ✓ | ✗ |
| 47. batch-db-inserts — Tests batch database inserts vs row-by-row | 5/5 | 5/5 | |
| 47.1 | Identifies individual inserts as the problem (50K round-trips) | ✓ | ✓ |
| 47.2 | Suggests batch inserts (multi-row VALUES or COPY protocol) | ✓ | ✓ |
| 47.3 | Shows a batching pattern with configurable batch size | ✓ | ✓ |
| 47.4 | Wraps batches in transactions for atomicity | ✓ | ✓ |
| 47.5 | Does NOT suggest only connection pooling or prepared statements | ✓ | ✓ |
| 48. panic-recover-control-flow — Tests panic/recover not for control flow | 4/4 | 3/4 | |
| 48.1 | Identifies that panic/recover is unnecessary since strconv.Atoi returns errors | ✓ | ✓ |
| 48.2 | Explains that panic allocates a stack trace and unwinds the stack (10-100x overhead) | ✓ | ✗ says "overhead" without quantifying |
| 48.3 | Suggests using simple error checking: v, err := strconv.Atoi(s) | ✓ | ✓ |
| 48.4 | States that panic/recover should only be used for truly unrecoverable situations | ✓ | ✓ |
| 49. monotonic-time-since — Tests time.Since monotonic clock for elapsed time | 3/3 | 3/3 | |
| 49.1 | Suggests using time.Since(start) for elapsed time measurement | ✓ | ✓ |
| 49.2 | Explains that time.Since uses the monotonic clock | ✓ | ✓ |
| 49.3 | Notes that time.Now() already captures monotonic time for Sub() operations | ✓ | ✓ |
| 50. prometheus-gc-pressure-queries — Tests specific PromQL queries for GC pressure | 4/4 | 0/4 | |
| 50.1 | Provides rate(go_gc_duration_seconds_count[5m]) for GC frequency | ✓ | ✗ generic approach |
| 50.2 | Provides go_gc_duration_seconds{quantile="1"} for worst-case GC pause | ✓ | ✗ |
| 50.3 | Mentions >2 cycles/s sustained as a signal of excessive allocation rate | ✓ | ✗ |
| 50.4 | Suggests rate(go_memstats_alloc_bytes_total[5m]) for allocation rate | ✓ | ✗ |
| 51. goroutine-leak-prometheus — Tests PromQL for detecting goroutine leaks | 3/3 | 2/3 | |
| 51.1 | Provides go_goroutines metric for goroutine count monitoring | ✓ | ✓ |
| 51.2 | Suggests delta(go_goroutines[1h]) for detecting net goroutine increase | ✓ | ✗ |
| 51.3 | Notes that goroutine count should correlate with load | ✓ | ✓ |
| 52. continuous-profiling-tools — Tests continuous profiling tools and tradeoffs | 5/5 | 2/5 | |
| 52.1 | Recommends Grafana Pyroscope, Parca, or similar platform | ✓ | ✓ |
| 52.2 | Mentions overhead estimates (1-5% range) | ✓ | ✗ |
| 52.3 | Describes push vs pull collection modes | ✓ | ✗ |
| 52.4 | Mentions historical flamegraph comparison as a key feature | ✓ | ✓ |
| 52.5 | Suggests feeding profiles into PGO for build optimization | ✓ | ✗ |
| 53. gogc-high-vs-low-tradeoff — Tests GOGC tuning tradeoffs (latency vs throughput) | 4/4 | 4/4 | |
| 53.1 | Recommends lower GOGC (e.g., 50) for Service A (latency-sensitive) | ✓ | ✓ |
| 53.2 | Recommends higher GOGC (e.g., 200) for Service B (throughput-oriented) | ✓ | ✓ |
| 53.3 | Explains the tradeoff: lower GOGC = more frequent but shorter pauses | ✓ | ✓ |
| 53.4 | Explains the tradeoff: higher GOGC = less frequent GC but more memory | ✓ | ✓ |
| 54. godebug-gctrace — Tests GODEBUG=gctrace=1 output interpretation | 5/5 | 3/5 | |
| 54.1 | Correctly identifies gc 142 as the 142nd GC cycle | ✓ | ✓ |
| 54.2 | Identifies 12% as total CPU time spent in GC (which is high) | ✓ | ✓ |
| 54.3 | Interprets 180->340->200 MB as heap before, peak during, and after collection | ✓ | ✗ misparses the three values |
| 54.4 | Identifies 400 MB goal as the target heap size based on GOGC/GOMEMLIMIT | ✓ | ✗ |
| 54.5 | Notes that 12% GC CPU is concerning | ✓ | ✓ |
| 55. unsafe-without-benchmark-proof — Tests warning against premature unsafe usage | 4/4 | 3/4 | |
| 55.1 | Recommends against using unsafe here | ✓ | ✓ |
| 55.2 | Notes that 100 req/s is not a hot path | ✓ | ✓ |
| 55.3 | States that unsafe requires benchmark proof showing >10% improvement | ✓ | ✗ no specific threshold |
| 55.4 | Mentions safety risks of unsafe | ✓ | ✓ |
| 56. precomputed-lookup-table — Tests precomputed lookup tables for pure functions | 3/3 | 2/3 | |
| 56.1 | Suggests a precomputed lookup table (var hexDigit = [16]byte{...}) | ✓ | ✓ |
| 56.2 | Shows the table lookup: hexDigit[b>>4], hexDigit[b&0x0f] | ✓ | ✓ |
| 56.3 | Explains that the lookup table fits in L1 cache and is faster than branching | ✓ | ✗ says "faster" without L1 cache explanation |
| 57. json-performance-alternatives — Tests JSON performance alternatives | 5/5 | 4/5 | |
| 57.1 | Mentions custom MarshalJSON/UnmarshalJSON methods | ✓ | ✓ |
| 57.2 | Mentions code-generation libraries (easyjson, ffjson) | ✓ | ✓ |
| 57.3 | Mentions drop-in replacement libraries (goccy/go-json, json-iterator, sonic) | ✓ | ✓ |
| 57.4 | Explains that encoding/json uses reflection | ✓ | ✓ |
| 57.5 | Quantifies expected improvement (2-5x or similar) | ✓ | ✗ |
| 58. channel-batch-processing — Tests batch processing from channels with timeout flush | 5/5 | 4/5 | |
| 58.1 | Uses select with both channel receive and ticker/timer for timeout | ✓ | ✓ |
| 58.2 | Flushes on batch size threshold | ✓ | ✓ |
| 58.3 | Flushes on timeout (ticker) | ✓ | ✓ |
| 58.4 | Handles channel close (flushes remaining items) | ✓ | ✓ |
| 58.5 | Reuses the batch slice (batch[:0]) to reduce allocations | ✓ | ✗ re-allocates |
| 59. allocation-reduction-vs-gc-tuning — Tests that reducing allocations beats GOGC tuning | 4/4 | 1/4 | |
| 59.1 | Recommends reducing allocations as the primary approach over GOGC tuning | ✓ | ✗ suggests GOGC tuning first |
| 59.2 | Explains that GOGC tuning manages the symptom while allocation reduction addresses root cause | ✓ | ✗ |
| 59.3 | Suggests specific allocation reduction strategies | ✓ | ✓ |
| 59.4 | Acknowledges GOGC tuning as a secondary measure after allocation reduction | ✓ | ✗ |
| 60. gctrace-key-fields — Tests what to monitor in GC traces | 4/4 | 2/4 | |
| 60.1 | Mentions high GC frequency as a signal of too many allocations | ✓ | ✓ |
| 60.2 | Mentions high pause times as a signal of large heap or many pointers | ✓ | ✗ |
| 60.3 | Mentions high GC CPU% (>5%) as concerning | ✓ | ✗ |
| 60.4 | Provides the GODEBUG=gctrace=1 command | ✓ | ✓ |
| 61. non-go-memory-leak-detection — Tests detecting non-Go memory leaks via Prometheus | 4/4 | 2/4 | |
| 61.1 | Suggests process_resident_memory_bytes - go_memstats_sys_bytes to isolate non-Go memory | ✓ | ✗ |
| 61.2 | Identifies this as a likely C/cgo memory leak | ✓ | ✓ |
| 61.3 | Explains that growing gap between RSS and Go sys bytes indicates non-Go memory growth | ✓ | ✗ |
| 61.4 | Suggests C-level memory profiling tools (valgrind, AddressSanitizer) | ✓ | ✓ |
| 62. cpu-saturation-prometheus — Tests detecting CPU saturation via Prometheus | 3/3 | 1/3 | |
| 62.1 | Provides rate(process_cpu_seconds_total[5m]) for CPU cores consumed | ✓ | ✓ |
| 62.2 | Divides by GOMAXPROCS to get utilization ratio | ✓ | ✗ |
| 62.3 | States that >0.8 sustained indicates CPU saturation | ✓ | ✗ |
| 63. document-optimizations — Tests recommending documentation of optimizations | 3/3 | 2/3 | |
| 63.1 | Strongly recommends adding comments explaining WHY the optimization was made | ✓ | ✓ |
| 63.2 | Suggests including benchmark numbers in the comments | ✓ | ✗ |
| 63.3 | Explains that future readers may revert optimizations they don't understand | ✓ | ✓ |
| 64. lru-cache-freelru — Tests high-performance LRU cache alternatives | 3/3 | 0/3 | |
| 64.1 | Notes that container/list has poor cache locality (separate heap allocation per node) | ✓ | ✗ |
| 64.2 | Recommends elastic/go-freelru or hashicorp/golang-lru as alternatives | ✓ | ✗ |
| 64.3 | Mentions the performance advantage of contiguous memory layouts for LRU | ✓ | ✗ |
| 65. simd-when-not-worth — Tests when SIMD is NOT worth pursuing | 4/4 | 4/4 | |
| 65.1 | Recommends against SIMD for this case | ✓ | ✓ |
| 65.2 | Explains that the bottleneck is allocations/GC, not CPU-bound computation | ✓ | ✓ |
| 65.3 | States that SIMD only helps CPU-bound numeric inner loops | ✓ | ✓ |
| 65.4 | Suggests reducing allocations as the correct approach | ✓ | ✓ |
| 66. set-map-struct-zero-size — Tests struct{} vs bool for set maps | 2/2 | 2/2 | |
| 66.1 | Recommends struct{} for set maps | ✓ | ✓ |
| 66.2 | Explains that struct{} is 0 bytes vs bool at 1 byte per entry | ✓ | ✓ |
| 67. regression-detection-prometheus — Tests PromQL for deployment regression detection | 3/3 | 2/3 | |
| 67.1 | Suggests comparing rate(go_memstats_alloc_bytes_total[5m]) before/after deploy | ✓ | ✗ |
| 67.2 | Suggests monitoring p99 latency histogram_quantile for increase after deploy | ✓ | ✓ |
| 67.3 | Mentions comparing metrics between old and new deployment versions | ✓ | ✓ |
| 68. statsviz-development-profiling — Tests real-time development visualization tools | 3/3 | 0/3 | |
| 68.1 | Recommends statsviz (github.com/arl/statsviz) for real-time browser visualization | ✓ | ✗ suggests pprof web UI |
| 68.2 | Mentions the /debug/statsviz endpoint or statsviz.Register pattern | ✓ | ✗ |
| 68.3 | Notes that it shows heap, GC pauses, goroutines, and scheduler in real-time | ✓ | ✗ |
golang-troubleshooting — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 186/186 (100%) | 127/186 (68%) | +32pp |
Full breakdown (186 assertions across 37 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. reproduce-before-fix-failing-test-first — Model must write a failing test before proposing any fix | 5/5 | 2/5 | |
| 1.1 | Writes a failing test (or proposes writing one) BEFORE suggesting any code fix | ✓ | ✗ jumps to adding error handling |
| 1.2 | Test covers multiple input variations to identify which inputs trigger the 500 | ✓ | ✗ single fix without test |
| 1.3 | Does NOT propose a fix without first understanding/reproducing the root cause | ✓ | ✗ proposes nil check immediately |
| 1.4 | Asks clarifying questions about the error (logs, request bodies, frequency) or proposes to gather evidence | ✓ | ✓ |
| 1.5 | Mentions running the test to confirm it reproduces the issue before fixing | ✓ | ✓ |
| 2. one-hypothesis-at-a-time — Change one thing and measure, not multiple simultaneous changes | 5/5 | 2/5 | |
| 2.1 | Refuses or strongly advises against making all three changes simultaneously | ✓ | ✗ applies all three as requested |
| 2.2 | Recommends measuring/profiling first to identify which problem is actually causing the symptoms | ✓ | ✓ |
| 2.3 | Suggests testing one hypothesis at a time with measurement between changes | ✓ | ✗ bundles changes |
| 2.4 | Mentions pprof, benchmarks, or race detector as diagnostic tools to identify the real bottleneck | ✓ | ✓ |
| 2.5 | Explains why multiple simultaneous changes are harmful (can't tell what worked, may introduce new bugs) | ✓ | ✗ |
| 3. root-cause-not-symptom-fix — Fix at the source where bad data originates, not where the panic occurs | 5/5 | 3/5 | |
| 3.1 | Identifies that the nil check in the handler is a symptom fix, not the root cause | ✓ | ✓ |
| 3.2 | Traces backward to the constructor/initialization code to find why db is nil | ✓ | ✓ |
| 3.3 | Suggests validating db != nil in the constructor (e.g., NewServer) and failing fast there | ✓ | ✓ |
| 3.4 | Does NOT suggest improving the nil check in the handler as the primary fix | ✓ | ✗ improves the nil check response |
| 3.5 | Explains that fixing at the symptom location masks the real bug | ✓ | ✗ |
| 4. interface-nil-gotcha — Typed nil in interface is not nil | 5/5 | 4/5 | |
| 4.1 | Identifies the interface nil gotcha: a typed nil *ValidationError wrapped in an error interface is NOT a nil interface | ✓ | ✓ |
| 4.2 | Explains that the interface has a non-nil type descriptor even when the pointer value is nil | ✓ | ✓ |
| 4.3 | Recommends returning nil explicitly (return nil) instead of returning the typed nil variable | ✓ | ✓ |
| 4.4 | Does NOT suggest adding an if verr == nil check before return as the primary fix | ✓ | ✗ suggests if verr == nil check |
| 4.5 | Shows or describes the correct fix: check verr != nil before return, and return nil in the else branch | ✓ | ✓ |
| 5. variable-shadowing-err — Inner err shadows outer err | 5/5 | 4/5 | |
| 5.1 | Identifies that := inside the if block creates a NEW err variable that shadows the outer one | ✓ | ✓ |
| 5.2 | Explains that the outer err remains nil because the inner := never assigned to it | ✓ | ✓ |
| 5.3 | Recommends using = (assignment) instead of := to assign to the outer err variable | ✓ | ✓ |
| 5.4 | Shows the fix: declare result separately (var result ResultType) and use result, err = someFunc() | ✓ | ✓ |
| 5.5 | Mentions go vet -shadow or shadow analyzer as a detection tool | ✓ | ✗ |
| 6. defer-in-loop-resource-leak — Deferred calls pile up until function returns | 5/5 | 4/5 | |
| 6.1 | Identifies that defer f.Close() inside a for loop keeps all files open until the function returns | ✓ | ✓ |
| 6.2 | Recommends wrapping the loop body in an anonymous function (closure) so defer runs each iteration | ✓ | ✓ |
| 6.3 | Does NOT suggest increasing ulimit or file descriptor limits as the primary solution | ✓ | ✓ |
| 6.4 | Shows the correct pattern with func() { f, err := os.Open(...); defer f.Close(); ... }() | ✓ | ✓ |
| 6.5 | Alternatively suggests extracting the loop body into a named function | ✓ | ✗ only shows closure |
| 7. break-in-select-inside-for-loop — Bare break only exits select, not the enclosing for loop | 5/5 | 4/5 | |
| 7.1 | Identifies that break inside a select only exits the select statement, not the for loop | ✓ | ✓ |
| 7.2 | Recommends using a labeled break (e.g., break loop) with a label on the for statement | ✓ | ✓ |
| 7.3 | Shows the correct pattern with a label like 'loop:' on the for statement and 'break loop' inside select | ✓ | ✓ |
| 7.4 | Also fixes the ctx.Done() case which has the same break issue | ✓ | ✗ only fixes the quit case |
| 7.5 | Alternatively mentions return as a solution if the function should exit entirely | ✓ | ✓ |
| 8. concurrent-map-fatal-not-panic — Concurrent map access is fatal and unrecoverable | 5/5 | 4/5 | |
| 8.1 | Explains that concurrent map read/write is a FATAL error that cannot be caught by recover() | ✓ | ✓ |
| 8.2 | Distinguishes this from regular panics — the Go runtime kills the process immediately | ✓ | ✗ treats it as a regular panic |
| 8.3 | Recommends protecting the map with sync.RWMutex or using sync.Map | ✓ | ✓ |
| 8.4 | Recommends using go test -race to find the race condition | ✓ | ✓ |
| 8.5 | Does NOT suggest fixing the recover middleware as the solution | ✓ | ✓ |
| 9. waitgroup-add-inside-goroutine — Add inside goroutine races with Wait | 5/5 | 3/5 | |
| 9.1 | Identifies that wg.Add(1) is called inside the goroutine instead of before it | ✓ | ✓ |
| 9.2 | Explains that wg.Wait() may return before all goroutines have called wg.Add(1) | ✓ | ✓ |
| 9.3 | Recommends moving wg.Add(1) before the go func() call | ✓ | ✓ |
| 9.4 | Notes this is a race condition that passes most of the time but fails intermittently | ✓ | ✗ |
| 9.5 | Does NOT focus primarily on the mutex/slice synchronization as the root cause | ✓ | ✗ discusses mutex issues equally |
| 10. missing-return-after-http-error — Missing return after http.Error() | 5/5 | 4/5 | |
| 10.1 | Identifies the missing return statement after http.Error(w, 'Forbidden', http.StatusForbidden) | ✓ | ✓ |
| 10.2 | Explains that http.Error() does NOT stop handler execution — it only writes to the ResponseWriter | ✓ | ✓ |
| 10.3 | Adds return statements after each http.Error() call | ✓ | ✓ |
| 10.4 | Does NOT primarily blame the isAuthorized function | ✓ | ✓ |
| 10.5 | Mentions this is a common Go bug pattern and a security concern | ✓ | ✗ |
| 11. json-numbers-float64-interface — Numbers into interface{} become float64 | 5/5 | 4/5 | |
| 11.1 | Explains that JSON numbers unmarshaled into interface{} become float64, not int or int64 | ✓ | ✓ |
| 11.2 | Notes that large integers (> 2^53) silently lose precision when stored as float64 | ✓ | ✗ |
| 11.3 | Recommends using a typed struct with int64 field as the preferred solution | ✓ | ✓ |
| 11.4 | Alternatively mentions json.NewDecoder with UseNumber() and json.Number for when interface{} is required | ✓ | ✓ |
| 11.5 | Explains the type assertion panics because the actual type is float64, not int64 | ✓ | ✓ |
| 12. strings-trim-vs-trimprefix — Trim treats argument as character set | 5/5 | 4/5 | |
| 12.1 | Explains that strings.Trim treats its second argument as a SET of characters to strip, not as a substring | ✓ | ✓ |
| 12.2 | Shows why 'json' becomes 'js' — the characters j, o, n are in the set {a,p,l,i,c,t,o,n,/} | ✓ | ✓ |
| 12.3 | Recommends strings.TrimPrefix for removing a substring prefix | ✓ | ✓ |
| 12.4 | Mentions strings.TrimSuffix for removing suffixes | ✓ | ✗ |
| 12.5 | Confirms this is NOT a Go bug — it's working as documented | ✓ | ✓ |
| 13. closed-channel-busy-loop-in-select — Closed channel in select causes busy loop | 5/5 | 3/5 | |
| 13.1 | Identifies that a closed channel always returns immediately (zero value) in a select case | ✓ | ✓ |
| 13.2 | Explains this causes the select case to fire continuously — a busy loop burning CPU | ✓ | ✓ |
| 13.3 | Recommends using the comma-ok idiom (job, ok := <-ch) and nil-ing the channel when closed (ch = nil) | ✓ | ✗ suggests only break/return |
| 13.4 | Explains that a nil channel blocks forever in select, effectively disabling that case | ✓ | ✗ |
| 13.5 | Does NOT suggest adding a default case with time.Sleep as the fix | ✓ | ✓ |
| 14. select-default-spin-loop — Select with default inside for loop is a busy-wait | 5/5 | 4/5 | |
| 14.1 | Identifies that select with default inside a for loop is a busy-wait spin loop | ✓ | ✓ |
| 14.2 | Explains that default runs immediately when no channel is ready, creating a tight loop | ✓ | ✓ |
| 14.3 | Recommends removing the default case and using a second channel or context for the stop signal | ✓ | ✓ |
| 14.4 | Alternatively suggests adding a time.Sleep or ticker in the default to yield CPU if non-blocking is truly required | ✓ | ✓ |
| 14.5 | Shows a solution using a ctx.Done() or stop channel in a second select case | ✓ | ✗ |
| 15. enum-zero-value-iota-ambiguity — Iota starting at 0 makes zero value ambiguous | 5/5 | 3/5 | |
| 15.1 | Identifies that iota starting at 0 makes the zero value (default for uninitialized fields) equal to Admin | ✓ | ✓ |
| 15.2 | Recommends reserving 0 for an Unknown/Unspecified sentinel value | ✓ | ✓ |
| 15.3 | Shows the pattern: RoleUnknown Role = iota, then Admin, Editor, Viewer | ✓ | ✓ |
| 15.4 | Explains this applies to any enum — zero value should be the 'unset' state, not a valid value | ✓ | ✗ |
| 15.5 | Does NOT primarily suggest fixing it in the constructor or registration logic | ✓ | ✗ suggests constructor fix |
| 16. recover-only-same-goroutine — recover() goroutine boundary | 5/5 | 4/5 | |
| 16.1 | Explains that recover() can ONLY catch panics in the same goroutine where it is deferred | ✓ | ✓ |
| 16.2 | States that a panic in a child goroutine will crash the entire program regardless of parent recovery | ✓ | ✓ |
| 16.3 | Recommends adding defer/recover inside the child goroutine (processAsync or its wrapper) | ✓ | ✓ |
| 16.4 | Shows the pattern: go func() { defer func() { if r := recover()... }(); processAsync(...) }() | ✓ | ✓ |
| 16.5 | Does NOT suggest reconfiguring the parent middleware as the solution | ✓ | ✗ also discusses middleware changes |
| 17. os-exit-skips-defers — log.Fatal calls os.Exit which skips deferred functions | 5/5 | 4/5 | |
| 17.1 | Identifies that log.Fatal (or log.Fatalf) calls os.Exit(1) internally | ✓ | ✓ |
| 17.2 | Explains that os.Exit skips all deferred functions — cleanup never runs | ✓ | ✓ |
| 17.3 | Recommends restructuring to avoid log.Fatal — use a run() function pattern or return errors | ✓ | ✓ |
| 17.4 | Shows the pattern: move logic into a run() error function, call os.Exit in main only after run returns | ✓ | ✗ suggests cleanup before Fatal instead |
| 17.5 | Does NOT suggest explicitly calling os.Remove before log.Fatal as the primary fix | ✓ | ✓ |
| 18. time-equal-not-double-equals — time.Time == vs .Equal() | 5/5 | 4/5 | |
| 18.1 | Identifies that time.Now() includes a monotonic clock reading that database serialization strips | ✓ | ✓ |
| 18.2 | Explains that == compares all fields including the monotonic component, so it can fail for equal instants | ✓ | ✗ |
| 18.3 | Recommends using .Equal() which ignores the monotonic clock | ✓ | ✓ |
| 18.4 | Alternatively mentions t.Round(0) to strip the monotonic reading before comparison or storage | ✓ | ✓ |
| 18.5 | Does NOT primarily blame database precision or timezone differences | ✓ | ✓ |
| 19. sql-rows-must-be-closed — Missing rows.Close() leaks connections | 5/5 | 4/5 | |
| 19.1 | Identifies the missing defer rows.Close() after the error check | ✓ | ✓ |
| 19.2 | Explains that unclosed sql.Rows holds the database connection until garbage collection | ✓ | ✓ |
| 19.3 | Adds defer rows.Close() immediately after the err check | ✓ | ✓ |
| 19.4 | Also notes the missing rows.Err() check after the loop | ✓ | ✗ |
| 19.5 | Does NOT primarily suggest increasing connection pool size | ✓ | ✓ |
| 20. copying-sync-types-value-receiver — Value receiver copies sync types | 5/5 | 3/5 | |
| 20.1 | Identifies that value receivers (c Counter) copy the entire struct including the Mutex on every call | ✓ | ✓ |
| 20.2 | Explains that each call operates on a copy — increments are lost and the mutex is duplicated | ✓ | ✓ |
| 20.3 | Recommends changing to pointer receivers (c *Counter) | ✓ | ✓ |
| 20.4 | Notes that go vet can detect copied sync types | ✓ | ✗ |
| 20.5 | Explains this applies to ALL sync types (Mutex, RWMutex, WaitGroup, Once, etc.) | ✓ | ✗ only discusses Mutex |
| 21. pprof-production-security — Never expose pprof unauthenticated | 5/5 | 2/5 | |
| 21.1 | Warns that pprof endpoints MUST be protected — never exposed publicly without authentication | ✓ | ✓ |
| 21.2 | Recommends basic auth or similar authentication on pprof endpoints | ✓ | ✗ suggests separate port only |
| 21.3 | Suggests running pprof on a separate port (not the main HTTP port) or localhost only | ✓ | ✓ |
| 21.4 | Recommends toggling pprof via an environment variable (e.g., PPROF_ENABLED) | ✓ | ✗ |
| 21.5 | Explains the risk: pprof leaks goroutine stacks, memory contents, and can be used for DoS | ✓ | ✗ |
| 22. godebug-gc-tracing-interpretation — GODEBUG gctrace output interpretation | 5/5 | 4/5 | |
| 22.1 | Identifies that 18% GC CPU overhead is significantly high (threshold is >10%) | ✓ | ✗ notes 18% but no threshold |
| 22.2 | Explains the heap size breakdown: 1024MB before mark, 900MB after mark, 500MB after sweep | ✓ | ✓ |
| 22.3 | Identifies the large pause times (45ms) as a likely cause of the latency spikes | ✓ | ✓ |
| 22.4 | Suggests the application is over-allocating and recommends investigating allocation patterns | ✓ | ✓ |
| 22.5 | Recommends using pprof heap/alloc profiling to find hot allocation sites | ✓ | ✓ |
| 23. research-codebase-not-just-diff — Trace callers and check upstream validation before flagging | 5/5 | 3/5 | |
| 23.1 | Recommends checking the callers first before flagging the bug | ✓ | ✓ |
| 23.2 | Suggests using Grep or similar to find all call sites of getItem | ✓ | ✗ |
| 23.3 | Notes that upstream code may validate the id (e.g., parsing from uint, bounds checking, positive-only input) | ✓ | ✗ |
| 23.4 | Advises that if callers validate, the severity is reduced but may still warrant a defensive check | ✓ | ✓ |
| 23.5 | Mentions adding an inline comment documenting the assumption if upstream guarantees exist | ✓ | ✓ |
| 24. flaky-test-diagnosis-methodology — Systematic flaky test debugging | 6/6 | 4/6 | |
| 24.1 | Recommends running with -count=100 to reproduce locally | ✓ | ✓ |
| 24.2 | Suggests using -shuffle=on to check for test order dependence | ✓ | ✗ |
| 24.3 | Mentions running with -race to check for data races | ✓ | ✓ |
| 24.4 | Suggests using t.TempDir() instead of shared temp directories to avoid file system pollution | ✓ | ✓ |
| 24.5 | Considers shared mutable state between tests as a potential cause | ✓ | ✓ |
| 24.6 | Does NOT suggest retry logic or skipping the test as a solution | ✓ | ✗ mentions retry as option |
| 25. defense-in-depth-after-fix — Multi-layer defense after fixing a bug | 5/5 | 3/5 | |
| 25.1 | Recommends validating at the entry point (API boundary) — reject paths with .. early | ✓ | ✓ |
| 25.2 | Recommends adding a test that specifically verifies the path traversal is blocked | ✓ | ✓ |
| 25.3 | Suggests adding logging or metrics to detect future traversal attempts (observability) | ✓ | ✗ |
| 25.4 | Considers multiple validation layers — not just one check | ✓ | ✗ focuses on single check |
| 25.5 | Mentions that filepath.Join does not prevent path traversal by itself (common misconception) | ✓ | ✓ |
| 26. escalation-protocol-three-failed-attempts — After 3 failed attempts, question architecture | 5/5 | 1/5 | |
| 26.1 | Recognizes the pattern of cascading failures as a red flag — each fix reveals a new problem | ✓ | ✓ |
| 26.2 | Recommends stepping back to question the overall design/architecture rather than trying another fix | ✓ | ✗ suggests specific fixes |
| 26.3 | Suggests re-reading the code from scratch with fresh eyes | ✓ | ✗ |
| 26.4 | Considers whether the current abstraction is fundamentally sound | ✓ | ✗ |
| 26.5 | Does NOT immediately suggest a 5th specific patch to the existing code | ✓ | ✗ suggests dedup + pool fix |
| 27. git-bisect-for-regression — Using git bisect to find breaking commit | 5/5 | 3/5 | |
| 27.1 | Recommends git bisect to binary-search for the breaking commit | ✓ | ✓ |
| 27.2 | Shows the git bisect start / git bisect bad / git bisect good workflow | ✓ | ✓ |
| 27.3 | Mentions that bisect can be automated with a test command (git bisect run go test -run TestBroken ./...) | ✓ | ✗ |
| 27.4 | Notes this narrows 50 commits to ~6 steps (log2(50)) | ✓ | ✗ |
| 27.5 | Does NOT suggest manually reading all 50 commit diffs | ✓ | ✓ |
| 28. check-external-dependencies-first — Verify external components before assuming code bug | 5/5 | 4/5 | |
| 28.1 | Suggests checking the external payment service health/status first (curl, health endpoint) | ✓ | ✓ |
| 28.2 | Recommends checking DNS resolution (dig or nslookup) | ✓ | ✗ |
| 28.3 | Suggests checking network connectivity (nc, telnet, or similar to the port) | ✓ | ✓ |
| 28.4 | Considers environment-specific causes: expired credentials, DNS changes, firewall rules, certificate rotation | ✓ | ✓ |
| 28.5 | Does NOT start by investigating Go code since nothing changed in the codebase | ✓ | ✓ |
| 29. observability-tools-before-code-dive — Check observability data before diving into code | 5/5 | 3/5 | |
| 29.1 | Recommends checking monitoring/observability tools BEFORE diving into code | ✓ | ✓ |
| 29.2 | Asks what monitoring tools are available (Prometheus, Datadog, Sentry, ELK, etc.) | ✓ | ✗ |
| 29.3 | Suggests checking error rate metrics, latency dashboards, or log aggregation | ✓ | ✗ |
| 29.4 | Mentions specific things to look for: what changed 2 hours ago (deploy, config change, traffic spike) | ✓ | ✓ |
| 29.5 | Does NOT immediately start reading source code files | ✓ | ✓ |
| 30. integer-conversion-silent-truncation — Go integer conversions silently truncate | 5/5 | 4/5 | |
| 30.1 | Explains that Go integer conversions silently truncate without any error or warning | ✓ | ✓ |
| 30.2 | Shows bounds checking before conversion: compare against math.MinInt32 and math.MaxInt32 | ✓ | ✓ |
| 30.3 | Returns an error when the value overflows instead of silently truncating | ✓ | ✓ |
| 30.4 | Notes there is no built-in safe conversion function — you must check bounds manually | ✓ | ✗ |
| 30.5 | Mentions this is especially dangerous for external/user-provided data | ✓ | ✓ |
| 31. init-ordering-fragile — init() ordering across files is fragile | 5/5 | 3/5 | |
| 31.1 | Confirms that init() ordering across files depends on filename alphabetical order and can change when files are added | ✓ | ✓ |
| 31.2 | Explains this makes init() dependencies fragile and hard to debug | ✓ | ✓ |
| 31.3 | Recommends replacing init() with explicit initialization in main() | ✓ | ✓ |
| 31.4 | Shows the pattern: cfg := loadConfig(); db := setupDatabase(cfg); startServer(db) | ✓ | ✗ |
| 31.5 | States init() should only be used for truly self-contained setup (registering drivers, codecs) | ✓ | ✗ |
| 32. goroutine-leak-detection-methodology — Goroutine leaks as cause of slow memory growth | 5/5 | 4/5 | |
| 32.1 | Suggests checking goroutine count (runtime.NumGoroutine or pprof goroutine profile) | ✓ | ✓ |
| 32.2 | Explains that goroutine leaks cause slow memory growth without appearing in heap profiles | ✓ | ✓ |
| 32.3 | Recommends using the pprof goroutine endpoint with ?debug=2 for human-readable stack dumps | ✓ | ✗ |
| 32.4 | Lists common causes: unclosed channels, missing context cancellation, forgotten response body close | ✓ | ✓ |
| 32.5 | Suggests goleak for detection in tests | ✓ | ✓ |
| 33. production-capture-before-restart — Capture profiles BEFORE restarting | 5/5 | 3/5 | |
| 33.1 | Recommends capturing profiles (heap, goroutine, CPU) BEFORE restarting | ✓ | ✓ |
| 33.2 | Explains that restarting destroys the evidence needed to diagnose the root cause | ✓ | ✗ |
| 33.3 | Lists specific profiles to capture: heap, goroutine dump (?debug=2), CPU (30s), mutex | ✓ | ✗ |
| 33.4 | Also suggests capturing system metrics (file descriptors, socket state, process info) | ✓ | ✓ |
| 33.5 | Only after capturing all evidence should the service be restarted if needed | ✓ | ✓ |
| 34. lock-contention-diagnosis — runtime.semacquire indicates lock contention | 5/5 | 4/5 | |
| 34.1 | Identifies runtime.semacquire as a signal of lock contention, not CPU computation | ✓ | ✓ |
| 34.2 | Recommends enabling mutex profiling with runtime.SetMutexProfileFraction(1) | ✓ | ✓ |
| 34.3 | Recommends enabling block profiling with runtime.SetBlockProfileRate(1) | ✓ | ✗ |
| 34.4 | Suggests using pprof mutex and block profiles to find the contended locks | ✓ | ✓ |
| 34.5 | Lists solutions: reduce critical section, sharding, RWMutex, atomic operations | ✓ | ✓ |
| 35. race-detector-not-reasoning — Never reason about concurrency, use the race detector | 5/5 | 4/5 | |
| 35.1 | Does NOT conclude safety based on code reasoning alone | ✓ | ✗ reasons about field layout |
| 35.2 | Recommends running go test -race to verify — never trust visual inspection for concurrency | ✓ | ✓ |
| 35.3 | Identifies that ++ is not atomic — RequestCount++ and ErrorCount++ are read-modify-write operations | ✓ | ✓ |
| 35.4 | Recommends using atomic.Int64 or sync.Mutex to protect the fields | ✓ | ✓ |
| 35.5 | Notes that even different fields on the same struct can race if accessed from different goroutines without synchronization | ✓ | ✓ |
| 36. filepath-join-path-traversal — filepath.Join does not prevent path traversal | 5/5 | 4/5 | |
| 36.1 | Identifies that filepath.Join does NOT prevent path traversal | ✓ | ✓ |
| 36.2 | Shows that input like '../../etc/passwd' resolves to '/etc/passwd' after Join | ✓ | ✓ |
| 36.3 | Recommends verifying the result has the base directory as a prefix after Join | ✓ | ✓ |
| 36.4 | Shows a safe pattern: filepath.Clean + strings.HasPrefix check | ✓ | ✓ |
| 36.5 | Notes the path separator must be appended to the base to prevent partial prefix matches | ✓ | ✗ |
| 37. time-after-in-loop-memory-leak — time.After in loop creates timer leak | 5/5 | 3/5 | |
| 37.1 | Identifies that time.After creates a new timer on every loop iteration that isn't garbage collected until it fires | ✓ | ✓ |
| 37.2 | Explains this causes a memory leak proportional to the message rate | ✓ | ✓ |
| 37.3 | Recommends using time.NewTimer with Reset() or time.NewTicker instead | ✓ | ✓ |
| 37.4 | Shows the correct pattern with a reusable timer and defer timer.Stop() | ✓ | ✗ |
| 37.5 | Does NOT suggest heap profiling as the first diagnostic step for this known pattern | ✓ | ✗ suggests heap profile first |
golang-design-patterns — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 87/87 (100%) | 55/87 (63%) | +37pp |
Full breakdown (87 assertions across 18 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. functional-options-over-builder — constructor API for HTTP server with growing config | 5/5 | 4/5 | |
| 1.1 | Uses functional options pattern (Option type as func that modifies the struct) | ✓ | ✓ |
| 1.2 | Constructor accepts variadic ...Option parameter | ✓ | ✓ |
| 1.3 | Each option is a With* function returning an Option | ✓ | ✓ |
| 1.4 | Sets sensible defaults inside the constructor before applying options | ✓ | ✓ |
| 1.5 | Mentions that functional options should return an error if validation can fail | ✓ | ✗ no error-returning variant mentioned |
| 2. avoid-init-function — database init() review | 5/5 | 3/5 | |
| 2.1 | Explicitly recommends against using init() for database initialization | ✓ | ✓ |
| 2.2 | Mentions that init() makes testing harder or unpredictable | ✓ | ✓ |
| 2.3 | Mentions that init() cannot return errors (must panic or log.Fatal) | ✓ | ✗ not mentioned |
| 2.4 | Suggests explicit constructor or initialization function | ✓ | ✓ |
| 2.5 | Mentions that init() runs before main/tests creating hidden dependencies | ✓ | ✗ not mentioned |
| 3. enum-start-at-one — order status enum with iota | 4/4 | 2/4 | |
| 3.1 | Zero value (iota = 0) is either skipped, named Unknown, Invalid, or Unspecified | ✓ | ✗ Pending at iota 0 |
| 3.2 | First meaningful enum value starts at 1 or higher | ✓ | ✗ starts at 0 |
| 3.3 | Explains WHY: Go's zero value would silently pass as the first enum member | ✓ | ✓ |
| 3.4 | Uses a custom type (not raw int or string) | ✓ | ✓ |
| 4. panic-vs-error-judgment — config parser: invalid format, missing field, nil arg | 5/5 | 3/5 | |
| 4.1 | Invalid config format: return error (caller can handle it) | ✓ | ✓ |
| 4.2 | Missing required field: return error (expected validation failure) | ✓ | ✓ |
| 4.3 | Nil passed to non-nil function: panic is acceptable (violated invariant) | ✓ | ✓ |
| 4.4 | Articulates the principle: panic is for bugs/invariant violations, errors are for expected failures | ✓ | ✗ vague distinction |
| 4.5 | Mentions Must* constructor pattern as a valid panic use case | ✓ | ✗ not mentioned |
| 5. runtime-addcleanup-over-setfinalizer — automatic cleanup for cgo resource (Go 1.24) | 5/5 | 1/5 | |
| 5.1 | Recommends runtime.AddCleanup as the preferred approach | ✓ | ✗ uses SetFinalizer |
| 5.2 | Mentions that AddCleanup supports multiple cleanups on the same object | ✓ | ✗ |
| 5.3 | Mentions that AddCleanup avoids object resurrection risk | ✓ | ✗ |
| 5.4 | Mentions that AddCleanup works even with cyclic references | ✓ | ✗ |
| 5.5 | Either warns against SetFinalizer or explains why AddCleanup is better | ✓ | ✓ |
| 6. resource-pool-bounded-channel — connection pool design | 5/5 | 2/5 | |
| 6.1 | Uses a buffered channel (chan *Conn with fixed capacity) as the pool mechanism | ✓ | ✗ uses slice+mutex |
| 6.2 | Pool has a maximum size / bounded capacity | ✓ | ✓ |
| 6.3 | Get operation uses select with context for timeout/cancellation | ✓ | ✗ no context support |
| 6.4 | Put operation handles pool-full case (discards excess connections) | ✓ | ✗ no overflow handling |
| 6.5 | Does NOT use sync.Pool as the primary pooling mechanism | ✓ | ✓ |
| 7. graceful-shutdown-signal-notifycontext — HTTP server SIGINT/SIGTERM handling | 5/5 | 3/5 | |
| 7.1 | Uses signal.NotifyContext (not raw signal.Notify with a channel) | ✓ | ✗ uses signal.Notify channel |
| 7.2 | Listens for both SIGINT and SIGTERM | ✓ | ✓ |
| 7.3 | Starts the HTTP server in a goroutine | ✓ | ✓ |
| 7.4 | Creates a separate timeout context for the shutdown phase | ✓ | ✗ no separate shutdown timeout |
| 7.5 | Closes other resources (DB, queues) after server shutdown | ✓ | ✓ |
| 8. iterator-streaming-large-data — export 2M rows to JSON HTTP response | 5/5 | 2/5 | |
| 8.1 | Does NOT load all 2M rows into a slice in memory | ✓ | ✗ loads all into []User |
| 8.2 | Streams the JSON response (writes records one at a time to the ResponseWriter) | ✓ | ✗ json.Marshal entire slice |
| 8.3 | Uses rows.Next() loop or iter.Seq2 iterator pattern | ✓ | ✓ |
| 8.4 | Defers rows.Close() immediately after query | ✓ | ✓ |
| 8.5 | Mentions OOM risk or memory concern as motivation for streaming | ✓ | ✗ no OOM discussion |
| 9. regexp-compile-once — email validation regex called thousands of times/sec | 4/4 | 3/4 | |
| 9.1 | Compiles the regexp at package level (var emailRegex = regexp.MustCompile(...)) | ✓ | ✓ |
| 9.2 | Does NOT compile the regexp inside the validation function | ✓ | ✓ |
| 9.3 | Uses regexp.MustCompile (not regexp.Compile) for package-level initialization | ✓ | ✓ |
| 9.4 | Explains WHY: compilation is O(n) and allocates, so doing it per-call is wasteful | ✓ | ✗ mentions performance but not O(n) |
| 10. architecture-right-sizing — 200-line CSV CLI tool architecture | 5/5 | 2/5 | |
| 10.1 | Recommends a flat or minimal structure (no multi-layer architecture) | ✓ | ✗ suggests cmd/internal/pkg layout |
| 10.2 | Does NOT suggest clean architecture, hexagonal, DDD, or ports and adapters | ✓ | ✗ suggests layered packages |
| 10.3 | Does NOT suggest dependency injection frameworks | ✓ | ✓ |
| 10.4 | Structure has at most cmd/ and possibly internal/, not handler/service/repository layers | ✓ | ✗ includes service layer |
| 10.5 | Mentions that architecture complexity should match project scope | ✓ | ✓ |
| 11. hexagonal-vs-clean-architecture — 8K-line order service with HTTP/gRPC/MQ | 5/5 | 3/5 | |
| 11.1 | Correctly explains hexagonal: ports (interfaces) + adapters with primary/secondary distinction | ✓ | ✓ |
| 11.2 | Correctly explains clean architecture: dependency rule (inward) with entities/use-cases/adapters | ✓ | ✓ |
| 11.3 | Recommends hexagonal for this case (multiple entry points: HTTP, gRPC, message consumer) | ✓ | ✗ recommends clean architecture |
| 11.4 | Mentions that both keep domain logic pure and free from infrastructure dependencies | ✓ | ✓ |
| 11.5 | Provides a directory structure with adapter/primary/ and adapter/secondary/ or equivalent | ✓ | ✗ generic layered layout |
| 12. ddd-aggregate-root-mutations — Order aggregate with AddItem, Draft status guard | 5/5 | 3/5 | |
| 12.1 | AddItem is a method on the Order aggregate root (not on OrderItem or a service) | ✓ | ✓ |
| 12.2 | Order fields (items, status) are unexported to prevent external mutation | ✓ | ✗ exported Items/Status fields |
| 12.3 | AddItem validates the status constraint (only Draft orders are editable) | ✓ | ✓ |
| 12.4 | Repository interface is defined in the domain package, not in the infrastructure package | ✓ | ✗ repo interface in infrastructure |
| 12.5 | Domain types have no infrastructure imports (no sql, no http, no framework dependencies) | ✓ | ✓ |
| 13. ddd-bounded-context-communication — Order and Billing context communication | 5/5 | 3/5 | |
| 13.1 | Uses domain events (e.g. OrderPlaced event) for cross-context communication | ✓ | ✓ |
| 13.2 | Billing context does NOT directly import order's internal domain types | ✓ | ✗ imports order types directly |
| 13.3 | Shows or describes an anti-corruption layer that translates order events to billing types | ✓ | ✗ no ACL layer |
| 13.4 | Each bounded context has its own domain, application, and adapter layers | ✓ | ✓ |
| 13.5 | Mentions that direct type imports between contexts create tight coupling | ✓ | ✓ |
| 14. make-illegal-states-unrepresentable — Email type to prevent invalid emails | 5/5 | 3/5 | |
| 14.1 | Creates a dedicated Email type (struct with unexported address field) | ✓ | ✗ validates at function entry |
| 14.2 | Email can only be created via a constructor (NewEmail) that validates the address | ✓ | ✗ no constructor |
| 14.3 | The send function accepts the Email type instead of a raw string | ✓ | ✓ |
| 14.4 | Explains the principle: make illegal states unrepresentable through the type system | ✓ | ✓ |
| 14.5 | The unexported field prevents creating an Email without validation | ✓ | ✓ |
| 15. fail-fast-validate-at-boundaries — three-layer validation strategy | 5/5 | 2/5 | |
| 15.1 | Recommends validating at the HTTP handler layer (the system boundary) | ✓ | ✓ |
| 15.2 | Recommends that the service and repository layers trust the data is already valid | ✓ | ✗ suggests validating in service too |
| 15.3 | Explains WHY: re-validating at every layer clutters code and violates DRY | ✓ | ✗ advocates "defense in depth" |
| 15.4 | Does NOT suggest adding the same validation checks in all three layers | ✓ | ✗ validates at all layers |
| 15.5 | May distinguish between input validation (at boundary) and business rule validation (in domain) | ✓ | ✓ |
16. explicit-over-implicit-defaults — struct tag default:"8080" review |
4/4 | 2/4 | |
| 16.1 | Recommends against using struct tags + reflection for defaults | ✓ | ✗ endorses struct tags as convenient |
| 16.2 | Suggests explicit defaults in a constructor function (e.g. NewConfig()) | ✓ | ✓ |
| 16.3 | Explains WHY: Go favors explicitness, struct tags hide behavior | ✓ | ✗ presents it as valid Go pattern |
| 16.4 | Shows a constructor that returns a Config with default values set explicitly | ✓ | ✓ |
| 17. retry-context-check — production-ready retry with exponential backoff | 5/5 | 3/5 | |
| 17.1 | Function accepts a context.Context parameter | ✓ | ✓ |
| 17.2 | Checks ctx.Err() or ctx.Done() between retry attempts | ✓ | ✓ |
| 17.3 | Uses select with ctx.Done() for the backoff delay (not time.Sleep) | ✓ | ✗ uses time.Sleep |
| 17.4 | Implements exponential backoff | ✓ | ✓ |
| 17.5 | Returns the context error if the context is cancelled during retry | ✓ | ✗ returns last operation error |
| 18. ddd-value-object-money — monetary amounts with addition and comparison | 5/5 | 2/5 | |
| 18.1 | Uses int64 (cents) not float64 for the amount | ✓ | ✗ uses float64 |
| 18.2 | Includes a currency field | ✓ | ✓ |
| 18.3 | Fields are unexported (immutable value object, only created via constructor) | ✓ | ✗ exported Amount/Currency |
| 18.4 | Add method validates currency match before addition | ✓ | ✓ |
| 18.5 | Constructor validates input (e.g. currency is required) | ✓ | ✗ no validation in constructor |
golang-cli — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 55/58 (95%) | 30/58 (52%) | +43pp |
Full breakdown (58 assertions across 12 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. minimal-main-and-execute — main.go minimal, os.Exit only in main | 5/5 | 2/5 | |
| 1.1 | main.go only calls Execute() — no config, flag setup, or business logic | ✓ | ✓ |
| 1.2 | Root command sets SilenceUsage: true | ✓ | ✗ not set |
| 1.3 | Root command sets SilenceErrors: true | ✓ | ✗ not set |
| 1.4 | Subcommands do NOT call os.Exit() inside RunE | ✓ | ✗ os.Exit(1) in RunE |
| 1.5 | Push subcommand registered via rootCmd.AddCommand() in init() | ✓ | ✓ |
| 2. viper-config-layering — flags > env > config file > defaults | 5/5 | 2/5 | |
| 2.1 | Calls viper.BindPFlag to bind the port flag to Viper | ✓ | ✗ flag and Viper disconnected |
| 2.2 | Sets env prefix with viper.SetEnvPrefix | ✓ | ✗ no prefix, bare PORT |
| 2.3 | Calls viper.AutomaticEnv() | ✓ | ✓ |
| 2.4 | Handles viper.ConfigFileNotFoundError gracefully | ✓ | ✗ log.Fatal on missing config |
| 2.5 | Precedence order correct: flags > env > config > defaults | ✓ | ✓ |
| 3. persistent-pre-run-config-init — config init in PersistentPreRunE | 4/5 | 2/5 | |
| 3.1 | Config init in PersistentPreRunE on root command | ✓ | ✗ duplicated in each RunE |
| 3.2 | Config init NOT duplicated in each subcommand | ✓ | ✗ three copies of init logic |
| 3.3 | --config is a persistent flag on root command | ✓ | ✓ |
| 3.4 | Uses SetEnvKeyReplacer for hyphens-to-underscores | ✗ | ✗ |
| 3.5 | Logging configured to write to stderr | ✓ | ✓ |
| 4. stdout-vs-stderr-separation — output to stdout, diagnostics to stderr | 5/5 | 2/5 | |
| 4.1 | Program output uses cmd.OutOrStdout(), not os.Stdout | ✓ | ✗ fmt.Println to os.Stdout |
| 4.2 | Logs/progress go to stderr | ✓ | ✗ log.Println to stdout |
| 4.3 | Errors go to stderr via cmd.ErrOrStderr() | ✓ | ✗ errors mixed with stdout |
| 4.4 | Uses cmd.OutOrStdout() instead of os.Stdout directly | ✓ | ✓ |
| 4.5 | Returns error from RunE, not os.Exit() or log.Fatal() | ✓ | ✓ |
| 5. version-ldflags-injection — version via ldflags, not hardcoded | 5/5 | 4/5 | |
| 5.1 | Version/commit/date are var (not const) with placeholder defaults | ✓ | ✓ |
| 5.2 | Shows -ldflags '-X ...' build command | ✓ | ✓ |
| 5.3 | Version command uses cmd.OutOrStdout() | ✓ | ✗ fmt.Println |
| 5.4 | Version NOT hardcoded as const | ✓ | ✓ |
| 5.5 | Includes runtime/debug.ReadBuildInfo() fallback | ✓ | ✓ |
| 6. exit-code-conventions — 0 success, 1 general, 2 usage | 4/5 | 2/5 | |
| 6.1 | Exit codes 0/1/2 follow Unix conventions | ✓ | ✗ all errors exit 1 |
| 6.2 | os.Exit() only in main() | ✓ | ✗ os.Exit inside handlers |
| 6.3 | Typed error or ExitError with Code field to propagate exit codes | ✗ | ✗ |
| 6.4 | Errors returned from commands, not swallowed by os.Exit() | ✓ | ✓ |
| 6.5 | Different error categories map to different exit codes | ✓ | ✓ |
| 7. signal-handling-with-context — signal.NotifyContext for cancellation | 5/5 | 4/5 | |
| 7.1 | Uses signal.NotifyContext, NOT raw channel with signal.Notify | ✓ | ✗ raw channel + select |
| 7.2 | Handles both os.Interrupt and syscall.SIGTERM | ✓ | ✓ |
| 7.3 | Creates shutdown timeout context (10-30s) | ✓ | ✓ |
| 7.4 | Calls srv.Shutdown(ctx), not srv.Close() | ✓ | ✓ |
| 7.5 | Distinguishes http.ErrServerClosed from unexpected errors | ✓ | ✓ |
| 8. flag-binding-and-constraints — persistent vs local, mutual exclusion, Viper binding | 5/5 | 2/5 | |
| 8.1 | --verbose is a persistent flag on root, not local to deploy | ✓ | ✗ local flag on deploy |
| 8.2 | Uses MarkFlagRequired for --env and --tag | ✓ | ✓ |
| 8.3 | Uses MarkFlagsMutuallyExclusive for --dry-run and --force | ✓ | ✗ manual check in RunE |
| 8.4 | Uses RegisterFlagCompletionFunc for --env values | ✓ | ✗ no completion |
| 8.5 | Binds flags to Viper with viper.BindPFlag | ✓ | ✓ |
| 9. argument-validation — Cobra built-in validators, not manual len(args) | 4/4 | 2/4 | |
| 9.1 | Uses cobra.NoArgs for status command | ✓ | ✗ manual len(args) == 0 |
| 9.2 | Uses cobra.ExactArgs(1) for deploy command | ✓ | ✓ |
| 9.3 | Uses cobra.RangeArgs(2, 3) for scale command | ✓ | ✗ manual range check in RunE |
| 9.4 | Validators set on Args field, not inside RunE | ✓ | ✓ |
| 10. cli-testing-pattern — execute programmatically with captured output | 5/5 | 3/5 | |
| 10.1 | Creates executeCommand helper with cmd.SetOut(buf) and cmd.SetErr(buf) | ✓ | ✗ no helper, tests run binary |
| 10.2 | Table-driven tests with multiple cases | ✓ | ✓ |
| 10.3 | Uses cmd.SetArgs() — not os/exec.Command | ✓ | ✗ exec.Command("./greet") |
| 10.4 | Captures output via bytes.Buffer — not redirecting os.Stdout | ✓ | ✓ |
| 10.5 | Tests check both output and error return value | ✓ | ✓ |
| 11. machine-readable-output-format — --output flag with json/table/plain | 4/4 | 2/4 | |
| 11.1 | Supports --output flag with json and table formats | ✓ | ✗ --json boolean toggle only |
| 11.2 | JSON output uses encoding/json to cmd.OutOrStdout() | ✓ | ✓ |
| 11.3 | Table output uses text/tabwriter for aligned columns | ✓ | ✗ fmt.Printf with manual spacing |
| 11.4 | Default format is human-readable table | ✓ | ✓ |
| 12. shell-completion-setup — Cobra built-in generators + custom completions | 4/5 | 3/5 | |
| 12.1 | Completion subcommand supports bash, zsh, fish, powershell | ✓ | ✓ |
| 12.2 | Uses Cobra's built-in GenBashCompletionV2, GenZshCompletion, etc. | ✓ | ✓ |
| 12.3 | Uses RegisterFlagCompletionFunc for --env flag | ✓ | ✗ no custom completions |
| 12.4 | Uses ValidArgs or cobra.ExactValidArgs for completion command | ✓ | ✓ |
| 12.5 | Returns cobra.ShellCompDirectiveNoFileComp for non-file flags | ✗ | ✗ |
golang-concurrency — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 62/62 (100%) | 38/62 (61%) | +39pp |
Full breakdown (62 assertions across 13 evals)
Model: Claude Opus 4.6 | Runs: 13 evals x 2 configs = 26 subagents | Grading: Human-as-a-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. time-after-in-select-loop — avoid time.After in loop, use time.NewTimer+Reset | 5/5 | 2/5 | |
| 1.1 | Does NOT use time.After inside the loop body | ✓ | ✗ uses time.After(5*time.Second) in select |
| 1.2 | Creates time.NewTimer (or time.NewTicker) outside the loop | ✓ | ✗ no timer created outside loop |
| 1.3 | Calls timer.Reset() after handling a message or timeout | ✓ | ✗ no Reset call |
| 1.4 | Calls timer.Stop() (or defers it) to clean up the timer | ✓ | ✓ |
| 1.5 | Drains the timer channel before Reset when appropriate (if !timer.Stop() { <-timer.C }) | ✓ | ✓ |
| 2. channel-closing-ownership — only the sender closes a channel | 5/5 | 3/5 | |
| 2.1 | Does NOT close the channel from the consumer/receiver side | ✓ | ✗ closes from consumer |
| 2.2 | Uses a separate signaling mechanism (done channel, context cancellation) for the consumer to tell the producer to stop | ✓ | ✓ |
| 2.3 | The producer is the one that closes the data channel | ✓ | ✗ consumer closes data channel |
| 2.4 | Explains the panic risk of closing a channel from the receiver side | ✓ | ✓ |
| 2.5 | The producer selects on the stop signal alongside its send operation | ✓ | ✓ |
| 3. waitgroup-add-placement — wg.Add before go statement | 4/4 | 4/4 | |
| 3.1 | Calls wg.Add(1) BEFORE the go statement, not inside the goroutine | ✓ | ✓ |
| 3.2 | Calls defer wg.Done() inside the goroutine | ✓ | ✓ |
| 3.3 | Calls wg.Wait() after the loop to wait for all goroutines | ✓ | ✓ |
| 3.4 | Does not place wg.Add inside the goroutine function body | ✓ | ✓ |
| 4. channel-direction-in-signatures — use chan<- and <-chan in function signatures | 5/5 | 4/5 | |
| 4.1 | The generator function returns <-chan int (receive-only for callers) | ✓ | ✓ |
| 4.2 | The doubler function accepts <-chan int as input parameter | ✓ | ✗ uses bidirectional chan int |
| 4.3 | The doubler function returns <-chan int (receive-only for callers) | ✓ | ✓ |
| 4.4 | Internally, channels are created as bidirectional but exposed as directional through return types | ✓ | ✓ |
| 4.5 | The producer (generator) closes its output channel with defer close(out) | ✓ | ✓ |
| 5. unbuffered-channel-default — default to unbuffered, justify buffers | 5/5 | 2/5 | |
| 5.1 | Recommends starting with unbuffered (or very small buffer like 0 or 1) as the default | ✓ | ✗ suggests buffer of 100 |
| 5.2 | Explains that large buffers mask backpressure problems | ✓ | ✗ no backpressure discussion |
| 5.3 | States that buffer size should be based on measured need, not arbitrary choice | ✓ | ✗ picks arbitrary size |
| 5.4 | Does NOT suggest a large arbitrary buffer (e.g., 100, 1000) without explaining the tradeoffs | ✓ | ✓ |
| 5.5 | Mentions that buffered channels hide the problem of slow consumers | ✓ | ✓ |
| 6. send-copies-not-pointers — send values through channels, not pointers | 4/4 | 1/4 | |
| 6.1 | Sends Task values (not *Task pointers) through the channel, OR explicitly documents why pointers are safe | ✓ | ✗ sends *Task pointers |
| 6.2 | The channel type is chan Task (value type) rather than chan *Task | ✓ | ✗ uses chan *Task |
| 6.3 | Does not mutate the Task struct after sending it on the channel (or sends a copy) | ✓ | ✗ no copy/immutability discussion |
| 6.4 | If pointers are used, explicitly acknowledges the shared-memory risk and explains the mitigation | ✓ | ✓ |
| 7. errgroup-vs-waitgroup-decision — errgroup.SetLimit for bounded concurrent work with errors | 5/5 | 3/5 | |
| 7.1 | Uses errgroup (golang.org/x/sync/errgroup) instead of sync.WaitGroup for error propagation | ✓ | ✓ |
| 7.2 | Uses errgroup.WithContext to cancel siblings on first error | ✓ | ✓ |
| 7.3 | Uses g.SetLimit(10) for bounded concurrency instead of a hand-rolled semaphore | ✓ | ✗ uses semaphore channel |
| 7.4 | Does NOT build a manual worker pool with channels and WaitGroup when errgroup suffices | ✓ | ✗ builds manual pool with channels |
| 7.5 | Each goroutine checks ctx.Done() or uses the context from errgroup.WithContext | ✓ | ✓ |
| 8. ctx-done-in-select — always include ctx.Done() in select to prevent goroutine leaks | 5/5 | 3/5 | |
| 8.1 | The function accepts a context.Context parameter | ✓ | ✓ |
| 8.2 | Every select statement includes a case <-ctx.Done(): return branch | ✓ | ✗ missing ctx.Done in output select |
| 8.3 | Both the read from input channel AND the write to output channel are wrapped in select with ctx.Done() | ✓ | ✗ only read has select |
| 8.4 | The goroutine exits cleanly when context is cancelled | ✓ | ✓ |
| 8.5 | The output channel is closed when the goroutine exits (defer close) | ✓ | ✓ |
| 9. sync-map-vs-rwmutex-decision — RWMutex+map for write-heavy overlapping keys | 5/5 | 2/5 | |
| 9.1 | Recommends sync.RWMutex + plain map over sync.Map for this write-heavy, overlapping-key pattern | ✓ | ✗ recommends sync.Map |
| 9.2 | Explains that sync.Map is optimized for write-once/read-many or disjoint key sets | ✓ | ✗ no access-pattern discussion |
| 9.3 | Explains that for frequent writes with overlapping keys, RWMutex+map is faster | ✓ | ✗ no performance comparison |
| 9.4 | Does NOT unconditionally recommend sync.Map for any concurrent map scenario | ✓ | ✓ |
| 9.5 | Mentions that concurrent map read/write without synchronization causes a hard crash | ✓ | ✓ |
| 10. sync-pool-reset-before-put — Reset() before Put(), not after Get() | 5/5 | 3/5 | |
| 10.1 | Calls buf.Reset() BEFORE bufPool.Put(buf), not after Get() | ✓ | ✗ resets after Get() |
| 10.2 | Uses defer to ensure the buffer is returned to the pool even on error | ✓ | ✓ |
| 10.3 | Does not assume the object from Get() is clean/zeroed | ✓ | ✗ assumes clean after Get |
| 10.4 | The pool's New function creates a new buffer | ✓ | ✓ |
| 10.5 | Does not store persistent state in pooled objects | ✓ | ✓ |
| 11. goroutine-panic-recovery — recover at goroutine boundaries | 4/4 | 3/4 | |
| 11.1 | Adds defer func() { recover() }() or equivalent panic recovery inside the goroutine | ✓ | ✓ |
| 11.2 | Logs or handles the recovered panic (not just silently swallowed) | ✓ | ✗ swallows panic silently |
| 11.3 | The goroutine has a shutdown mechanism (context, done channel, or similar) | ✓ | ✓ |
| 11.4 | Mentions that a panic in a goroutine crashes the entire process | ✓ | ✓ |
| 12. singleflight-cache-stampede — use singleflight to deduplicate concurrent lookups | 5/5 | 4/5 | |
| 12.1 | Recommends golang.org/x/sync/singleflight as the primary solution | ✓ | ✓ |
| 12.2 | Shows usage of group.Do(key, func) where key identifies the deduplicated resource | ✓ | ✓ |
| 12.3 | Explains that only one goroutine executes the function; others wait and share the result | ✓ | ✓ |
| 12.4 | May combine singleflight with a cache layer for TTL-based caching | ✓ | ✓ |
| 12.5 | Does NOT suggest only a plain mutex or only a TTL cache as the solution to thundering herd | ✓ | ✗ also suggests mutex as primary approach |
| 13. iterator-vs-goroutine-pipeline — iterators for sequential CPU-bound transforms | 5/5 | 4/5 | |
| 13.1 | Recommends against goroutine+channel pipeline for this purely sequential, in-memory transform | ✓ | ✓ |
| 13.2 | Suggests Go 1.23+ iterators (iter.Seq) or simple slice operations instead | ✓ | ✗ suggests only plain loop |
| 13.3 | Explains that goroutine+channel pipelines add overhead without benefit for sequential CPU-bound work | ✓ | ✓ |
| 13.4 | Mentions that goroutine pipelines are appropriate when stages involve I/O or need true parallelism | ✓ | ✓ |
| 13.5 | Does NOT build a multi-goroutine pipeline for this use case | ✓ | ✓ |
golang-context — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 48/50 (96%) | 31/50 (62%) | +34pp |
Full breakdown (50 assertions across 10 evals)
Model: Claude Opus 4.6 | Runs: 10 evals x 2 configs = 20 subagents | Grading: Human-as-a-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. context-background-in-handler — use r.Context(), not context.Background() in handlers | 5/5 | 4/5 | |
| 1.1 | Uses r.Context() to obtain the request context, NOT context.Background() inside the handler | ✓ | ✓ |
| 1.2 | Passes the same context (or a derived child) to the database query (*Context variant) | ✓ | ✓ |
| 1.3 | Passes the same context to the external HTTP call via http.NewRequestWithContext | ✓ | ✓ |
| 1.4 | Does NOT use http.NewRequest (without context) for the external service call | ✓ | ✓ |
| 1.5 | Checks ctx.Err() or handles context cancellation when the client disconnects | ✓ | ✗ no ctx.Err() check on disconnect |
| 2. cancel-leak-timeout — defer cancel() for every WithTimeout, per-attempt timeouts | 5/5 | 4/5 | |
| 2.1 | Creates a new context.WithTimeout for each retry attempt (not one timeout for all retries) | ✓ | ✓ |
| 2.2 | Calls defer cancel() (or cancel() before next iteration) for every WithTimeout call | ✓ | ✓ |
| 2.3 | Does NOT discard the cancel function with _ (e.g., ctx, _ = context.WithTimeout(...)) | ✓ | ✗ discards cancel with _ |
| 2.4 | Uses http.NewRequestWithContext to attach the per-attempt timeout context | ✓ | ✓ |
| 2.5 | Accepts a parent context parameter and derives timeouts from it | ✓ | ✓ |
| 3. context-value-key-type — unexported key types for context values | 5/5 | 3/5 | |
| 3.1 | Uses an unexported type for the context key (e.g., type contextKey string or type tenantKey struct{}) | ✓ | ✗ uses plain string key |
| 3.2 | Does NOT use a plain string as the context key | ✓ | ✗ context.WithValue(ctx, "tenant_id", ...) |
| 3.3 | Provides a typed getter function (e.g., TenantIDFromContext) that returns the value with proper type assertion | ✓ | ✓ |
| 3.4 | Provides a setter function or the middleware injects the value using the unexported key | ✓ | ✓ |
| 3.5 | The getter handles the case where the value is missing from the context (returns zero value + bool or error) | ✓ | ✓ |
| 4. context-in-struct-trap — do not store context.Context in a struct field | 5/5 | 3/5 | |
| 4.1 | Does NOT store context.Context as a field in the Worker struct | ✓ | ✗ stores ctx context.Context in struct |
| 4.2 | Passes context as a parameter to Start() or Run() method | ✓ | ✗ uses stored struct field context |
| 4.3 | Uses context cancellation or a done channel for graceful shutdown signaling | ✓ | ✓ |
| 4.4 | Listens to ctx.Done() in a select statement to detect shutdown | ✓ | ✓ |
| 4.5 | ctx is the first parameter where it appears, named ctx context.Context | ✓ | ✓ |
| 5. without-cancel-background-work — use context.WithoutCancel for background work (Go 1.21+) | 5/5 | 2/5 | |
| 5.1 | Uses context.WithoutCancel to create a context for the audit goroutine | ✓ | ✗ uses context.Background() |
| 5.2 | Does NOT pass r.Context() directly to the background audit goroutine | ✓ | ✓ |
| 5.3 | Does NOT use context.Background() for the audit goroutine (that would lose trace_id) | ✓ | ✗ uses context.Background() |
| 5.4 | The audit goroutine preserves request-scoped values (trace_id) from the original context | ✓ | ✗ context.Background() loses all values |
| 5.5 | Launches the audit as a separate goroutine (go keyword) so the handler can return immediately | ✓ | ✓ |
| 6. nested-timeout-shorter-wins — nested timeouts take the shorter deadline | 4/5 | 3/5 | |
| 6.1 | Creates the overall 10-second timeout from the parent context | ✓ | ✓ |
| 6.2 | Creates the cache timeout as a child of the overall context | ✓ | ✓ |
| 6.3 | Acknowledges that the database timeout is bounded by whatever time remains on the parent | ✗ creates 8s child without noting parent constraint | ✗ creates independent 8s timeout |
| 6.4 | Defers cancel() for every WithTimeout call | ✓ | ✓ |
| 6.5 | Does NOT create independent context.Background() timeouts that bypass the overall deadline | ✓ | ✗ creates context.Background() timeout for DB |
| 7. context-todo-vs-background — use context.TODO() as placeholder, not context.Background() | 5/5 | 3/5 | |
| 7.1 | Uses context.TODO() as the temporary placeholder in functions not yet fully migrated | ✓ | ✗ uses context.Background() everywhere |
| 7.2 | Uses context.Background() only at the true top level (main function or test setup) | ✓ | ✗ uses Background() at all levels |
| 7.3 | Shows a migration path where context.TODO() is gradually replaced as callers are updated | ✓ | ✓ |
| 7.4 | Context is always the first parameter, named ctx context.Context | ✓ | ✓ |
| 7.5 | Does NOT pass nil as a context value at any point in the migration | ✓ | ✓ |
| 8. db-context-variants — use *Context database method variants | 5/5 | 5/5 | |
| 8.1 | Uses db.QueryRowContext (not db.QueryRow) for GetByID | ✓ | ✓ |
| 8.2 | Uses db.ExecContext (not db.Exec) for Create, Update, and Delete | ✓ | ✓ |
| 8.3 | Passes the ctx parameter to every database call | ✓ | ✓ |
| 8.4 | Does NOT ignore the ctx parameter by calling non-context database methods | ✓ | ✓ |
| 8.5 | Each method accepts ctx context.Context as its first parameter | ✓ | ✓ |
| 9. context-values-abuse — do not stuff infrastructure deps into context values | 5/5 | 4/5 | |
| 9.1 | Database connection is passed as a struct field or explicit parameter, NOT via context value | ✓ | ✓ |
| 9.2 | Logger is passed as a struct field or explicit parameter, NOT via context value | ✓ | ✓ |
| 9.3 | Order details are passed as an explicit function parameter, NOT via context value | ✓ | ✓ |
| 9.4 | User ID and/or trace ID are stored in context values (these are request-scoped metadata) | ✓ | ✓ |
| 9.5 | Distinguishes between infrastructure dependencies (explicit) and request-scoped metadata (context values) | ✓ | ✗ no explicit distinction made |
| 10. afterfunc-cleanup — use context.AfterFunc for non-blocking cleanup (Go 1.21+) | 4/5 | 2/5 | |
| 10.1 | Uses context.AfterFunc to register the cleanup callback | ✓ | ✗ uses manual goroutine with <-ctx.Done() |
| 10.2 | Does NOT block the main flow waiting for context cancellation | ✓ | ✓ |
| 10.3 | The cleanup function removes the temporary file | ✓ | ✓ |
| 10.4 | Captures the stop function returned by AfterFunc for potential cancellation of the callback | ✓ | ✗ no stop function (uses goroutine) |
| 10.5 | Also includes a defer-based cleanup as a safety net (AfterFunc + defer for belt-and-suspenders) | ✗ only AfterFunc, no defer fallback | ✗ only goroutine, no defer fallback |
golang-continuous-integration — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 66/66 (100%) | 27/66 (41%) | +59pp |
Full breakdown (66 assertions across 13 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. test-workflow-flags — CI test workflow with required flags and matrix | 5/5 | 3/5 | |
| 1.1 | Workflow includes -race flag in the go test command |
✓ | ✓ |
| 1.2 | Workflow includes -shuffle=on flag in the go test command |
✓ | ✗ omitted shuffle |
| 1.3 | Workflow includes -coverprofile flag in the go test command |
✓ | ✓ |
| 1.4 | Strategy uses fail-fast: false |
✓ | ✗ default fail-fast: true |
| 1.5 | Go version matrix includes at least 'stable' and one explicit version | ✓ | ✓ |
| 2. go-mod-tidy-check — enforce go mod tidy via git diff --exit-code | 4/4 | 2/4 | |
| 2.1 | Suggests running go mod tidy as a CI step |
✓ | ✓ |
| 2.2 | Includes git diff --exit-code after go mod tidy to detect uncommitted changes |
✓ | ✗ no git diff check |
| 2.3 | The git diff checks go.mod and/or go.sum specifically, or uses a general git diff --exit-code | ✓ | ✗ only runs tidy |
| 2.4 | Also includes go mod verify or go mod download step |
✓ | ✓ |
| 3. integration-test-caching — integration tests with -count=1 and service containers | 5/5 | 2/5 | |
| 3.1 | Uses -count=1 flag to disable test result caching |
✓ | ✗ no -count=1 |
| 3.2 | Includes -race flag for integration tests |
✓ | ✓ |
| 3.3 | Uses build tags (e.g., -tags=integration) to separate integration tests |
✓ | ✗ no build tags |
| 3.4 | Uses GitHub Actions services block for PostgreSQL and/or Redis |
✓ | ✓ |
| 3.5 | Includes health check options for service containers | ✓ | ✗ no health checks |
| 4. security-scanning-pipeline — full security stack: govulncheck, gosec, CodeQL, Bearer | 6/6 | 2/6 | |
| 4.1 | Recommends govulncheck and explains it only reports actually-called code paths | ✓ | ✗ mentions govulncheck without call-path |
| 4.2 | Recommends gosec for Go security scanning | ✓ | ✓ |
| 4.3 | Recommends CodeQL and mentions the security-extended or security-and-quality query suite | ✓ | ✗ no extended suite |
| 4.4 | Recommends Bearer for sensitive data flow issues | ✓ | ✗ not mentioned |
| 4.5 | Workflow includes security-events: write permission for SARIF upload |
✓ | ✗ no permission block |
| 4.6 | Suggests creating a CodeQL config file to use an extended query suite | ✓ | ✓ |
| 5. dependabot-grouping-strategy — group minor/patch, separate majors | 6/6 | 3/6 | |
| 5.1 | Configures Dependabot for gomod package ecosystem | ✓ | ✓ |
| 5.2 | Configures Dependabot for github-actions package ecosystem | ✓ | ✓ |
| 5.3 | Configures Dependabot for docker package ecosystem | ✓ | ✗ omitted docker |
| 5.4 | Groups minor and patch Go module updates into a single PR | ✓ | ✗ no grouping |
| 5.5 | Major updates are NOT grouped (individual PRs for breaking changes) | ✓ | ✗ all updates grouped |
| 5.6 | Sets a weekly schedule | ✓ | ✓ |
| 6. dependabot-auto-merge-security — auto-merge with actor guard and branch protection | 5/5 | 0/5 | |
| 6.1 | Workflow has if: github.actor == dependabot[bot] guard |
✓ | ✗ no actor guard |
| 6.2 | Workflow checks metadata to exclude major updates from auto-merge | ✓ | ✗ merges all versions |
| 6.3 | Warns about contents: write and pull-requests: write being elevated permissions |
✓ | ✗ no security warning |
| 6.4 | Mentions branch protection rules as the real safety net | ✓ | ✗ not mentioned |
| 6.5 | Notes that github.actor checks are not fully spoof-proof |
✓ | ✗ not mentioned |
| 7. renovate-vs-dependabot — Renovate advantages for monorepos | 5/5 | 2/5 | |
| 7.1 | Recommends Renovate as an alternative to Dependabot | ✓ | ✓ |
| 7.2 | Mentions Renovate's gomodTidy feature | ✓ | ✗ not mentioned |
| 7.3 | Mentions Renovate's native automerge without needing a separate workflow | ✓ | ✗ not mentioned |
| 7.4 | Mentions Renovate's monorepo/workspace support | ✓ | ✓ |
| 7.5 | Mentions Renovate's better grouping rules | ✓ | ✗ not mentioned |
| 8. goreleaser-library-vs-cli — GoReleaser config for libraries vs CLIs | 5/5 | 2/5 | |
| 8.1 | Uses skip: true in the builds section since libraries don't produce binaries |
✓ | ✗ includes full build config |
| 8.2 | Keeps the config minimal (mainly changelog generation) | ✓ | ✗ full CLI-style config |
| 8.3 | Mentions that for libraries, a simple gh release create may be sufficient |
✓ | ✗ not mentioned |
| 8.4 | Does NOT include cross-compilation (goos/goarch) in the library config | ✓ | ✓ |
| 8.5 | Includes changelog configuration | ✓ | ✓ |
| 9. docker-workflow-security — push: false on PRs, per-job permissions, provenance | 6/6 | 2/6 | |
| 9.1 | Sets push to false on pull requests | ✓ | ✗ pushes on all events |
| 9.2 | Uses per-job permissions scoping | ✓ | ✗ top-level permissions only |
| 9.3 | Includes QEMU and Buildx setup for multi-platform builds | ✓ | ✓ |
| 9.4 | Includes provenance and/or SBOM attestation configuration | ✓ | ✗ no attestations |
| 9.5 | Includes packages: write permission for GHCR push |
✓ | ✓ |
| 9.6 | Login step is conditional on non-PR events | ✓ | ✗ unconditional login |
| 10. permissions-least-privilege — read-only GITHUB_TOKEN, branch protection, fork PRs | 7/7 | 3/7 | |
| 10.1 | Recommends setting default GITHUB_TOKEN to read-only | ✓ | ✗ not mentioned |
| 10.2 | Recommends branch protection with required status checks | ✓ | ✓ |
| 10.3 | Recommends requiring PR approvals (at least 1) | ✓ | ✓ |
| 10.4 | Recommends dismissing stale approvals when new commits are pushed | ✓ | ✗ not mentioned |
| 10.5 | Recommends restricting fork PR workflows for outside collaborators | ✓ | ✗ not mentioned |
| 10.6 | Warns against pull_request_target with untrusted code |
✓ | ✗ not mentioned |
| 10.7 | Recommends creating a release environment with required reviewers | ✓ | ✓ |
| 11. release-workflow-fetch-depth — fetch-depth: 0 for changelog generation | 4/4 | 2/4 | |
| 11.1 | Checkout step uses fetch-depth: 0 for full git history |
✓ | ✗ default shallow clone |
| 11.2 | Workflow triggers on tag push with a v* pattern |
✓ | ✓ |
| 11.3 | Uses contents: write permission for creating releases |
✓ | ✗ no permissions block |
| 11.4 | Passes GITHUB_TOKEN to GoReleaser |
✓ | ✓ |
| 12. action-version-pinning — pin to major versions, not branches | 3/3 | 2/3 | |
| 12.1 | Identifies that using @master and @main is wrong and insecure |
✓ | ✓ |
| 12.2 | Recommends pinning to major versions like @v4, @v6 |
✓ | ✓ |
| 12.3 | Explains the risk: branch references can change unexpectedly or be compromised | ✓ | ✗ identifies problem but no risk reasoning |
| 13. coverage-threshold-configuration — codecov.yml with project and patch targets | 5/5 | 2/5 | |
| 13.1 | Configures codecov.yml (not just CLI flags) for coverage thresholds |
✓ | ✗ uses CLI flags only |
| 13.2 | Sets project target to 80% | ✓ | ✓ |
| 13.3 | Sets a threshold value (e.g., 2%) to allow small drops | ✓ | ✗ no threshold |
| 13.4 | Configures patch coverage target for new code in PRs | ✓ | ✗ no patch target |
| 13.5 | Coverage upload is conditional on a single matrix entry (e.g., only on stable) | ✓ | ✓ |
golang-dependency-injection — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 54/55 (98%) | 28/55 (51%) | +47pp |
Full breakdown (55 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. constructor-injection-not-globals — inject via constructors, not globals or init() | 5/5 | 3/5 | |
| 1.1 | Uses constructor injection (NewUserService taking dependencies as parameters) | ✓ | ✓ |
| 1.2 | Explicitly advises against package-level variables for service dependencies | ✓ | ✗ accepts global var approach |
| 1.3 | Explains why globals are problematic (untestable, hidden dependencies, or coupling) | ✓ | ✗ no explanation of problems |
| 1.4 | Does NOT use init() for service initialization |
✓ | ✓ |
| 1.5 | Returns a concrete struct pointer from the constructor, not an interface | ✓ | ✓ |
| 2. interface-defined-at-consumer — interfaces defined where consumed, not implemented | 5/5 | 2/5 | |
| 2.1 | Defines the interface in the consuming package, not the implementation package | ✓ | ✗ defines interface next to implementation |
| 2.2 | Explains the principle: accept interfaces, return structs | ✓ | ✗ not stated |
| 2.3 | The implementation package returns a concrete struct pointer | ✓ | ✓ |
| 2.4 | The consumer depends on its own locally-defined interface | ✓ | ✗ imports provider interface |
| 2.5 | Does NOT have the implementation package import the consumer's interface | ✓ | ✓ |
| 3. container-not-passed-as-dependency — service locator anti-pattern | 5/5 | 2/5 | |
| 3.1 | Advises against passing the injector/container as a dependency | ✓ | ✗ accepts passing injector |
| 3.2 | Identifies this as the service locator anti-pattern | ✓ | ✗ not identified |
| 3.3 | Shows that the Injector should only exist at the composition root | ✓ | ✗ injector used throughout |
| 3.4 | Shows UserService receiving Database and Mailer directly as constructor parameters | ✓ | ✓ |
| 3.5 | Shows the provider function using do.MustInvoke inside the provider, not inside methods |
✓ | ✓ |
| 4. manual-di-for-small-projects — small projects use manual DI, not a library | 5/5 | 2/5 | |
| 4.1 | Recommends manual constructor injection for a project with only 5 services | ✓ | ✗ recommends Wire or Fx |
| 4.2 | Does NOT recommend a DI library as the primary approach | ✓ | ✗ leads with DI library |
| 4.3 | Shows wiring in main() with explicit constructor calls in dependency order |
✓ | ✓ |
| 4.4 | Initializes infrastructure first, then repositories, then services, then transport | ✓ | ✗ no layered ordering |
| 4.5 | Mentions that a DI library becomes worthwhile at 10-20+ services | ✓ | ✓ |
| 5. di-library-selection-judgment — correct library recommendation based on criteria | 5/5 | 3/5 | |
| 5.1 | Recommends samber/do as a strong fit (generics, lifecycle, compile-time safety) | ✓ | ✗ recommends uber-go/fx |
| 5.2 | Explains why uber-go/fx uses reflection (runtime errors, not compile-time) | ✓ | ✓ |
| 5.3 | Explains why google/wire lacks built-in lifecycle management | ✓ | ✓ |
| 5.4 | Mentions that samber/do requires Go 1.18+ for generics | ✓ | ✗ samber/do not discussed |
| 5.5 | Discusses at least 3 DI library options from the decision table | ✓ | ✓ |
| 6. wire-build-constraint-and-codegen — proper google/wire setup with wireinject | 5/5 | 3/5 | |
| 6.1 | Includes //go:build wireinject build constraint in the wire.go file |
✓ | ✗ omitted build constraint |
| 6.2 | Uses wire.Build with all provider functions listed |
✓ | ✓ |
| 6.3 | Shows wire.Bind for binding interface to implementation |
✓ | ✗ no wire.Bind shown |
| 6.4 | Explains that wire generates wire_gen.go with plain constructor calls | ✓ | ✓ |
| 6.5 | Mentions that wire_gen.go must not be edited manually | ✓ | ✓ |
| 7. fx-lifecycle-hooks-pattern — uber-go/fx OnStart/OnStop hooks for DB connection | 5/5 | 4/5 | |
| 7.1 | Uses fx.Lifecycle parameter in the provider function |
✓ | ✓ |
| 7.2 | Registers OnStart hook for establishing the database connection | ✓ | ✓ |
| 7.3 | Registers OnStop hook for closing the database connection | ✓ | ✓ |
| 7.4 | Uses lc.Append(fx.Hook{...}) pattern |
✓ | ✓ |
| 7.5 | OnStart and OnStop take context.Context as parameter |
✓ | ✗ omitted context parameter |
| 8. testing-with-di-mock-injection — inject mocks at the interface boundary | 5/5 | 4/5 | |
| 8.1 | Creates a mock implementation of the UserStore interface | ✓ | ✓ |
| 8.2 | Injects the mock into UserService via the constructor (NewUserService) | ✓ | ✓ |
| 8.3 | Tests both the success path (user found) and the error path (not found) | ✓ | ✓ |
| 8.4 | Does NOT use a real database connection in the test | ✓ | ✓ |
| 8.5 | The mock is defined in the test file, not as a package-level or global variable | ✓ | ✗ mock as package-level var |
| 9. shallow-dependency-graph — deep chains are a design problem | 5/5 | 2/5 | |
| 9.1 | Identifies the deep dependency chain as a design problem | ✓ | ✗ accepts as normal layered architecture |
| 9.2 | Recommends flattening the dependency graph | ✓ | ✗ no flattening suggested |
| 9.3 | Suggests that most services should depend on repositories and config directly | ✓ | ✗ keeps transitive chain |
| 9.4 | Explains negative consequences of deep chains (fragility, hard to test) | ✓ | ✓ |
| 9.5 | Proposes a concrete restructuring | ✓ | ✓ |
| 10. one-container-per-app-not-per-request — container created once, not per request | 5/5 | 2/5 | |
| 10.1 | Identifies creating a new container per request as a mistake | ✓ | ✗ accepts per-request as reasonable |
| 10.2 | Recommends one container per application created at startup | ✓ | ✗ no single-container guidance |
| 10.3 | Explains the performance/correctness problem (recreating singletons, no connection reuse) | ✓ | ✗ no performance concern |
| 10.4 | Suggests using scopes for request-level isolation if needed | ✓ | ✓ |
| 10.5 | Shows the container being created once in main() and services injected into handlers |
✓ | ✓ |
| 11. lazy-vs-eager-initialization — lazy init preference, singleton vs transient | 4/5 | 1/5 | |
| 11.1 | Recommends lazy initialization (services created on first use, not all at startup) | ✓ | ✗ recommends eager for everything |
| 11.2 | Recommends singletons for stateful services like database connections | ✓ | ✓ |
| 11.3 | Recommends transients (or factories) for stateless request processing services | ✓ | ✗ all singletons |
| 11.4 | Explains why lazy loading is beneficial (unused services never created, faster startup) | ✗ | ✗ no lazy benefits explained |
| 11.5 | Notes which DI libraries support lazy loading (samber/do, fx) vs which don't (wire) | ✓ | ✗ no library comparison for lazy |
golang-dependency-management — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 52/52 (100%) | 24/52 (46%) | +54pp |
Full breakdown (52 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
1. ask-before-adding-dependency — AI agent must ask before go get |
5/5 | 1/5 | |
| 1.1 | Asks the user for confirmation before running go get |
✓ | ✗ runs go get immediately |
| 1.2 | Presents the package name and import path | ✓ | ✗ installs without showing path |
| 1.3 | Mentions whether the standard library covers the use case | ✓ | ✗ no stdlib mention |
| 1.4 | Lists known alternatives | ✓ | ✗ suggests one library directly |
| 1.5 | Does NOT silently run go get without asking first |
✓ | ✓ |
| 2. go-sum-must-be-committed — go.sum is critical for supply-chain security | 5/5 | 3/5 | |
| 2.1 | Strongly advises against gitignoring go.sum | ✓ | ✓ |
| 2.2 | Explains that go.sum contains cryptographic checksums for dependency verification | ✓ | ✓ |
| 2.3 | Explains the supply-chain security risk: compromised proxy could substitute malicious code | ✓ | ✗ no supply-chain mention |
| 2.4 | Mentions go mod verify as the mechanism that uses go.sum |
✓ | ✗ not mentioned |
| 2.5 | Recommends removing go.sum from .gitignore | ✓ | ✓ |
3. patch-only-upgrade-preference — go get -u=patch safer than go get -u |
5/5 | 1/5 | |
| 3.1 | Recommends go get -u=patch ./... as the safer default |
✓ | ✗ recommends go get -u ./... |
| 3.2 | Explains that -u=patch only upgrades patch versions (no API changes per semver) |
✓ | ✗ no distinction made |
| 3.3 | Explains that -u upgrades minor versions too, which can change behavior |
✓ | ✗ no warning about minor |
| 3.4 | Mentions running go mod tidy after upgrading |
✓ | ✓ |
| 3.5 | Does NOT recommend go get -u ./... without warning about risk |
✓ | ✗ recommends it as primary |
| 4. mvs-algorithm-understanding — Minimal Version Selection, not latest | 5/5 | 3/5 | |
| 4.1 | Correctly states that Go selects v1.3.0 (not the latest available) | ✓ | ✓ |
| 4.2 | Explains MVS: Go picks the highest minimum required, not the latest available | ✓ | ✓ |
| 4.3 | Distinguishes MVS from other package managers (npm, pip, cargo) | ✓ | ✗ no comparison to other managers |
| 4.4 | Mentions that MVS provides deterministic builds without a lock file | ✓ | ✗ not mentioned |
| 4.5 | Explains that go.sum is integrity verification, not version locking | ✓ | ✓ |
| 5. major-version-suffix-rule — /v2 suffix for v2+ modules | 5/5 | 4/5 | |
| 5.1 | States that the module path must include /v2 suffix |
✓ | ✓ |
| 5.2 | States that all import paths must be updated to include /v2 |
✓ | ✓ |
| 5.3 | Explains this is Go's import compatibility rule | ✓ | ✓ |
| 5.4 | Mentions that v0 and v1 do NOT have a suffix | ✓ | ✗ not mentioned |
| 5.5 | Notes that v1 and v2 can coexist in the same build | ✓ | ✓ |
| 6. replace-directive-library-warning — replace ignored when consumed as dependency | 4/4 | 2/4 | |
| 6.1 | States that replace directives only take effect in the main module's go.mod | ✓ | ✓ |
| 6.2 | States that consumers will NOT use the fork | ✓ | ✓ |
| 6.3 | Recommends removing replace directives before publishing a library | ✓ | ✗ no publish warning |
| 6.4 | Suggests alternative solutions (upstream the fix, publish fork as separate module) | ✓ | ✗ no alternatives offered |
| 7. tools-go-pattern — pin CLI tool versions in go.mod | 5/5 | 1/5 | |
| 7.1 | Recommends the tools.go pattern (file with //go:build tools constraint) |
✓ | ✗ suggests go install @latest in CI |
| 7.2 | Uses blank imports (_ imports) to keep tools in go.mod |
✓ | ✗ no blank imports |
| 7.3 | The build constraint ensures the file is never compiled into production code | ✓ | ✗ no build constraint |
| 7.4 | Mentions running go mod tidy after creating the tools.go file |
✓ | ✗ no tidy mention |
| 7.5 | Explains that go install then uses the pinned version from go.mod |
✓ | ✓ |
| 8. govulncheck-call-path-analysis — static analysis traces call paths to vulnerable funcs | 5/5 | 3/5 | |
| 8.1 | Recommends govulncheck to check if vulnerability is reachable | ✓ | ✓ |
| 8.2 | Explains that govulncheck uses static analysis to trace call paths | ✓ | ✗ no call-path explanation |
| 8.3 | Explains that if code never calls the affected function, govulncheck will NOT flag it | ✓ | ✗ not explained |
| 8.4 | Shows the govulncheck ./... command |
✓ | ✓ |
| 8.5 | Distinguishes govulncheck from generic CVE scanners | ✓ | ✓ |
| 9. go-work-sum-gitignore — go.work.sum should NOT be committed | 5/5 | 2/5 | |
| 9.1 | States that go.work.sum should NOT be committed | ✓ | ✗ suggests committing both |
| 9.2 | Recommends adding go.work.sum to .gitignore | ✓ | ✗ no gitignore mention |
| 9.3 | Explains that go.work is for development only | ✓ | ✓ |
| 9.4 | Distinguishes this from go.sum which MUST be committed | ✓ | ✗ treats both the same |
| 9.5 | May mention that go.work itself can optionally be committed | ✓ | ✓ |
| 10. exclude-vs-retract-distinction — exclude is consumer-side, retract is author-side | 5/5 | 3/5 | |
| 10.1 | Uses retract for the published library (author-side) |
✓ | ✓ |
| 10.2 | Uses exclude for the buggy dependency (consumer-side) |
✓ | ✗ uses replace instead |
| 10.3 | Explains that retract goes in the library's own go.mod | ✓ | ✓ |
| 10.4 | Explains that exclude redirects to the next higher available version | ✓ | ✗ no redirect behavior explained |
| 10.5 | Notes that retracted versions are still downloadable but not selected by default | ✓ | ✓ |
11. test-dependency-upgrade-flag — -t flag for including test deps in upgrades |
3/3 | 1/3 | |
| 11.1 | Explains that go get -u ./... excludes test-only dependencies by default |
✓ | ✗ not explained |
| 11.2 | Recommends go get -u -t ./... to include test dependencies |
✓ | ✗ suggests upgrading individually |
| 11.3 | Explains the difference between -u and -u -t |
✓ | ✓ |
golang-structs-interfaces — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 52/52 (100%) | 34/52 (65%) | +35pp |
Full breakdown (52 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. interface-at-consumer-not-implementor — notification service using email client | 5/5 | 2/5 | |
| 1.1 | Interface is defined in the notification package (the consumer), NOT in the email package | ✓ | ✗ defines in email package |
| 1.2 | Interface has only the methods the notification package needs (not the full email.Client API) | ✓ | ✗ mirrors full Client API |
| 1.3 | Email package exports a concrete Client struct, not an interface | ✓ | ✓ |
| 1.4 | Explains WHY: keeps the consumer in control of the contract, avoids importing for interface | ✓ | ✗ no rationale given |
| 1.5 | Notification service depends on its own interface, not on email package types | ✓ | ✓ |
| 2. return-structs-not-interfaces — NewUserService constructor signature | 5/5 | 3/5 | |
| 2.1 | Constructor returns *UserService (concrete type), NOT an interface | ✓ | ✓ |
| 2.2 | Constructor accepts UserStore as an interface parameter (accept interfaces) | ✓ | ✓ |
| 2.3 | Explains WHY: callers get full access to concrete type; consumers can assign to interface | ✓ | ✗ no rationale |
| 2.4 | Explicitly states that returning interfaces from constructors is bad practice | ✓ | ✗ presents both as valid |
| 2.5 | The accept-interfaces-return-structs principle is stated or demonstrated | ✓ | ✓ |
| 3. premature-interface-trap — single PostgreSQL implementation repository | 5/5 | 1/5 | |
| 3.1 | Recommends starting with a concrete struct (not an interface) when only one implementation | ✓ | ✗ creates interface immediately |
| 3.2 | Mentions the principle: don't design with interfaces, discover them | ✓ | ✗ not mentioned |
| 3.3 | Suggests extracting an interface LATER when a second consumer or test mock demands it | ✓ | ✗ creates interface "for testability" |
| 3.4 | Acknowledges that testability IS a valid reason but it should be a deliberate choice | ✓ | ✗ reflexively adds interface |
| 3.5 | Does NOT reflexively recommend creating an interface just because it is a repository | ✓ | ✓ |
| 4. zero-value-useful-design — Registry panics without NewRegistry | 5/5 | 2/5 | |
| 4.1 | Recommends lazy initialization in the Register method (if r.items == nil) | ✓ | ✗ says "always call constructor" |
| 4.2 | Mentions the Go principle: make the zero value useful | ✓ | ✗ not mentioned |
| 4.3 | The fix allows using var r Registry without calling a constructor | ✓ | ✗ requires constructor |
| 4.4 | Does NOT just say "always use the constructor" as the primary fix | ✓ | ✓ |
| 4.5 | References bytes.Buffer or sync.Mutex as stdlib examples of useful zero values | ✓ | ✓ |
| 5. embedding-vs-named-field — Server with http.Handler and DataStore | 5/5 | 3/5 | |
| 5.1 | Embeds http.Handler (to promote ServeHTTP to the Server) | ✓ | ✓ |
| 5.2 | Uses a named field for DataStore (not embedded, methods should not be exposed) | ✓ | ✗ embeds both |
| 5.3 | Explains the embed vs named field rule: embed for "is a", named field for "has a" | ✓ | ✗ no rule articulated |
| 5.4 | Mentions that embedding promotes ALL methods of the inner type, which can be undesirable | ✓ | ✓ |
| 5.5 | Notes that the receiver of promoted methods is the inner type, not the outer type | ✓ | ✓ |
| 6. compile-time-interface-check — MyBuffer implements io.ReadWriter | 4/4 | 3/4 | |
| 6.1 | Uses var _ io.ReadWriter = (*MyBuffer)(nil) pattern | ✓ | ✓ |
| 6.2 | Places the check near the type definition | ✓ | ✗ places in test file |
| 6.3 | Explains that this costs nothing at runtime | ✓ | ✓ |
| 6.4 | Explains that the build fails immediately if MyBuffer stops satisfying the interface | ✓ | ✓ |
| 7. type-assertion-comma-ok — check if interface{} value is a string | 4/4 | 4/4 | |
| 7.1 | Uses the comma-ok form: s, ok := val.(string) | ✓ | ✓ |
| 7.2 | Checks the ok value before using s | ✓ | ✓ |
| 7.3 | Warns against bare type assertion (s := val.(string)) because it panics | ✓ | ✓ |
| 7.4 | Handles the !ok case explicitly | ✓ | ✓ |
| 8. optional-behavior-type-assertion — flush only if writer supports it | 5/5 | 3/5 | |
| 8.1 | Defines a separate Flusher interface with just the Flush method | ✓ | ✓ |
| 8.2 | Function parameter is io.Writer (not a combined interface) | ✓ | ✗ requires WriteFlusher interface |
| 8.3 | Uses type assertion (f, ok := w.(Flusher)) to check for flush capability | ✓ | ✓ |
| 8.4 | Only calls Flush if the type assertion succeeds | ✓ | ✓ |
| 8.5 | Mentions this pattern is used in the standard library (e.g. http.Flusher, io.ReaderFrom) | ✓ | ✗ no stdlib reference |
| 9. nocopy-sentinel-struct — prevent ConnPool from being passed by value | 5/5 | 1/5 | |
| 9.1 | Recommends embedding a noCopy sentinel struct | ✓ | ✗ says "always use pointers" |
| 9.2 | noCopy implements Lock() and Unlock() methods (empty bodies) | ✓ | ✗ not mentioned |
| 9.3 | Explains that go vet will flag copies of structs containing noCopy | ✓ | ✗ not mentioned |
| 9.4 | Mentions this is the same technique used by sync.WaitGroup, sync.Mutex, or strings.Builder | ✓ | ✗ not mentioned |
| 9.5 | Shows that the struct should be passed by pointer after adding noCopy | ✓ | ✓ |
| 10. generics-over-any-interface — generic Contains function for any comparable type | 5/5 | 4/5 | |
| 10.1 | Uses generics with a type parameter: func Contains[T comparable] | ✓ | ✓ |
| 10.2 | Does NOT use []any or interface{} parameters | ✓ | ✓ |
| 10.3 | Uses the comparable constraint for the type parameter | ✓ | ✓ |
| 10.4 | Explains WHY: generics preserve type safety, while any loses it | ✓ | ✓ |
| 10.5 | Mentions that any should only be used at true boundaries where type is genuinely unknown | ✓ | ✗ no boundary guidance |
| 11. receiver-consistency-rule — mixed pointer/value receivers on same type | 4/4 | 2/4 | |
| 11.1 | Says mixing pointer and value receivers on the same type is wrong or not recommended | ✓ | ✗ says mixing is fine per-method |
| 11.2 | Recommends making ALL methods use pointer receivers since one needs to mutate | ✓ | ✓ |
| 11.3 | Explains WHY: consistency rule -- if any method uses a pointer receiver, all should | ✓ | ✗ no consistency principle articulated |
| 11.4 | Mentions that method sets differ for T and *T which affects interface satisfaction | ✓ | ✓ |
golang-lint — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 49/51 (96%) | 28/51 (55%) | +41pp |
Full breakdown (51 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. nolint-directive-specificity — linter name + justification on every //nolint | 5/5 | 3/5 | |
| 1.1 | Every //nolint specifies the linter name — no bare //nolint | ✓ | ✗ bare //nolint on 3 lines |
| 1.2 | Every //nolint includes a justification comment | ✓ | ✗ no justifications |
| 1.3 | Type assertion uses //nolint:forcetypeassert with safety explanation | ✓ | ✓ |
| 1.4 | Long test uses //nolint:funlen with table-driven justification | ✓ | ✓ |
| 1.5 | Cyclomatic complexity uses //nolint:gocyclo with orchestration justification | ✓ | ✓ |
| 2. nolint-fix-vs-suppress-judgment — fix correctness bugs, suppress style issues | 6/6 | 4/6 | |
| 2.1 | Recommends FIXING bodyclose — unclosed response bodies leak connections | ✓ | ✓ |
| 2.2 | Recommends SUPPRESSING funlen on table-driven test | ✓ | ✓ |
| 2.3 | Recommends FIXING errcheck on database query in request handler | ✓ | ✗ suppressed as "not critical" |
| 2.4 | Recommends SUPPRESSING dupl on intentional parallel structure | ✓ | ✓ |
| 2.5 | Recommends FIXING sqlclosecheck — unclosed sql.Rows leak connections | ✓ | ✗ suppressed with defer comment |
| 2.6 | Recommends SUPPRESSING goconst in tests | ✓ | ✓ |
| 3. golangci-yml-version-2-structure — version: "2", formatters section, linters.enable | 4/5 | 1/5 | |
| 3.1 | Config has version: "2" at the top | ✓ | ✗ no version field |
| 3.2 | Linters listed under linters.enable (not enable-all) | ✓ | ✗ uses enable-all with disable list |
| 3.3 | gofumpt under formatters.enable, NOT linters.enable | ✓ | ✗ gofumpt in linters.enable |
| 3.4 | errcheck has check-type-assertions: true in linters.settings | ✓ | ✓ |
| 3.5 | Timeout set under run.timeout: 5m | ✗ | ✗ |
| 4. linter-categories-correctness-vs-style — prioritize bug-finding over style | 5/5 | 3/5 | |
| 4.1 | Includes govet and staticcheck as highest-value correctness linters | ✓ | ✓ |
| 4.2 | Includes errcheck for unchecked errors | ✓ | ✓ |
| 4.3 | Prioritizes correctness/safety over style linters | ✓ | ✗ revive and godot ranked above gosec |
| 4.4 | Includes at least one security linter (bodyclose, gosec, sqlclosecheck) | ✓ | ✗ no security linters in top 10 |
| 4.5 | Does NOT include redundant complexity checkers (gocyclo + cyclop) | ✓ | ✓ |
| 5. legacy-codebase-incremental-adoption — new-from-rev, not mass //nolint | 5/5 | 3/5 | |
| 5.1 | Recommends issues.new-from-rev to only lint new/changed code | ✓ | ✓ |
| 5.2 | Does NOT suggest adding //nolint to all 2000+ warnings | ✓ | ✗ suggests mass //nolint annotations |
| 5.3 | Suggests gradual cleanup of old code over time | ✓ | ✓ |
| 5.4 | Suggests golangci-lint run --fix for auto-fixable issues | ✓ | ✓ |
| 5.5 | Mentions parallel sub-agents or batching fixes by category | ✓ | ✗ sequential manual cleanup |
| 6. interpreting-lint-output-format — use linter name to triage fix vs suppress | 5/5 | 3/5 | |
| 6.1 | Identifies errcheck on DB.Close as real issue to fix | ✓ | ✓ |
| 6.2 | Identifies bodyclose as critical resource leak to fix | ✓ | ✗ treats all warnings equally |
| 6.3 | Identifies unused validateToken as dead code to remove | ✓ | ✓ |
| 6.4 | Evaluates gocyclo based on function nature (orchestration vs complex) | ✓ | ✗ blanket "refactor" recommendation |
| 6.5 | Identifies revive comment warning as lower-priority style issue | ✓ | ✓ |
| 7. disabled-linters-with-rationale — exhaustruct, gochecknoglobals, wrapcheck, mnd, varnamelen | 5/5 | 3/5 | |
| 7.1 | Recommends AGAINST exhaustruct — breaks zero-value idiom | ✓ | ✗ enables it for "completeness" |
| 7.2 | Recommends AGAINST gochecknoglobals — valid global uses exist | ✓ | ✓ |
| 7.3 | Recommends AGAINST wrapcheck as default — too noisy | ✓ | ✗ enables it unconditionally |
| 7.4 | Recommends AGAINST mnd — flags obvious constants like HTTP status codes | ✓ | ✓ |
| 7.5 | Recommends AGAINST varnamelen — conflicts with Go's short name idiom | ✓ | ✓ |
| 8. nolintlint-meta-linter — enforce proper //nolint hygiene automatically | 4/4 | 2/4 | |
| 8.1 | Recommends enabling nolintlint linter | ✓ | ✓ |
| 8.2 | Configures require-specific: true | ✓ | ✗ no nolintlint config |
| 8.3 | Configures require-explanation: true | ✓ | ✗ no nolintlint config |
| 8.4 | Shows correct location: linters.settings.nolintlint | ✓ | ✓ |
| 9. multiple-nolint-comma-syntax — comma-separated linters in single directive | 3/3 | 2/3 | |
| 9.1 | Uses comma-separated: //nolint:errcheck,gosec — not two directives | ✓ | ✗ two separate //nolint lines |
| 9.2 | Includes justification explaining both false positives | ✓ | ✓ |
| 9.3 | Directive on same line as flagged code or line above | ✓ | ✓ |
| 10. common-config-issues — timeout, v1-to-v2 migration, linter-not-found | 3/4 | 2/4 | |
| 10.1 | Recommends increasing run.timeout for deadline exceeded | ✓ | ✓ |
| 10.2 | Recommends golangci-lint migrate for v1 config errors | ✓ | ✗ suggests manual rewrite |
| 10.3 | For linter not found: check version — modernize requires v2.6.0+ | ✓ | ✗ suggests reinstalling |
| 10.4 | Mentions golangci-lint linters to check available linters | ✗ | ✓ |
| 11. formatter-vs-linter-distinction — formatters section, fmt subcommand | 4/4 | 2/4 | |
| 11.1 | gofumpt under formatters.enable, NOT linters.enable | ✓ | ✗ linters.enable |
| 11.2 | Sets gofumpt extra-rules: true under formatters.settings | ✓ | ✓ |
| 11.3 | Mentions golangci-lint fmt ./... command for formatting | ✓ | ✗ only golangci-lint run |
| 11.4 | Notes gci and goimports are redundant with gofumpt | ✓ | ✓ |
golang-popular-libraries — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 54/54 (100%) | 38/54 (70%) | +30pp |
Full breakdown (54 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. stdlib-first-json — encoding/json before third-party | 5/5 | 5/5 | |
| 1.1 | Recommends encoding/json as the first option | ✓ | ✓ |
| 1.2 | Third-party alternatives only for specific performance needs | ✓ | ✓ |
| 1.3 | Explains stdlib is sufficient for most JSON use cases | ✓ | ✓ |
| 1.4 | Alternatives presented for measured performance requirements | ✓ | ✓ |
| 1.5 | Does NOT recommend third-party without first considering stdlib | ✓ | ✓ |
| 2. pgx-over-lib-pq — pgx as primary PostgreSQL driver | 5/5 | 3/5 | |
| 2.1 | Recommends pgx as the primary recommendation | ✓ | ✗ lib/pq as primary |
| 2.2 | Mentions pgx is faster than lib/pq | ✓ | ✓ |
| 2.3 | Notes pgx supports all PostgreSQL types and advanced features | ✓ | ✓ |
| 2.4 | May mention lib/pq but positions pgx as preferred | ✓ | ✓ |
| 2.5 | Does NOT recommend lib/pq as primary without mentioning pgx | ✓ | ✗ lib/pq recommended first |
| 3. chi-for-minimal-router — chi when staying close to net/http | 5/5 | 3/5 | |
| 3.1 | Recommends chi as a strong match for stated requirements | ✓ | ✗ recommends Gin |
| 3.2 | Explains chi is lightweight and composes with net/http | ✓ | ✓ |
| 3.3 | Notes chi has minimal dependencies | ✓ | ✓ |
| 3.4 | May mention Gin/Echo but positions chi as better fit for net/http | ✓ | ✓ |
| 3.5 | Does NOT recommend a full framework as primary when user wants to stay close to net/http | ✓ | ✗ Gin as primary recommendation |
| 4. slog-over-external-loggers — consider log/slog before zap/zerolog | 5/5 | 4/5 | |
| 4.1 | Mentions log/slog as standard library option (Go 1.21+) | ✓ | ✓ |
| 4.2 | Presents slog as a viable option, not afterthought | ✓ | ✗ brief mention at end |
| 4.3 | If recommending external, explains specific value over slog | ✓ | ✓ |
| 4.4 | Does NOT skip standard library consideration entirely | ✓ | ✓ |
| 4.5 | May mention zap/zerolog for specific use cases | ✓ | ✓ |
| 5. sqlc-vs-orm-decision — sqlc for compile-time SQL safety | 5/5 | 2/5 | |
| 5.1 | Recommends sqlc as a primary option for compile-time SQL safety | ✓ | ✗ recommends GORM only |
| 5.2 | Explains sqlc generates type-safe code with no runtime reflection | ✓ | ✗ no mention of sqlc |
| 5.3 | Mentions GORM uses runtime reflection — doesn't catch SQL errors at compile time | ✓ | ✗ GORM presented without caveat |
| 5.4 | May also mention ent as a code-generated alternative | ✓ | ✓ |
| 5.5 | Does NOT recommend only GORM when user asks for compile-time safety | ✓ | ✓ |
| 6. rate-limiter-stdlib-first — golang.org/x/time/rate before third-party | 5/5 | 4/5 | |
| 6.1 | Recommends golang.org/x/time/rate as official option | ✓ | ✓ |
| 6.2 | Explains token bucket algorithm | ✓ | ✓ |
| 6.3 | May mention Tollbooth/limiter for HTTP middleware integration | ✓ | ✓ |
| 6.4 | Does NOT skip the official x/time/rate package | ✓ | ✗ no mention of x/time/rate |
| 6.5 | Explains when third-party might be preferred (per-IP, distributed) | ✓ | ✓ |
| 7. franz-go-for-kafka — franz-go over legacy sarama | 4/4 | 2/4 | |
| 7.1 | Recommends franz-go as primary recommendation | ✓ | ✗ sarama as primary |
| 7.2 | Describes franz-go as modern, high-performance, feature-complete | ✓ | ✓ |
| 7.3 | Does NOT recommend only sarama without mentioning franz-go | ✓ | ✗ only sarama mentioned |
| 7.4 | May mention sarama but positions franz-go as preferred modern choice | ✓ | ✓ |
| 8. check-maintenance-before-recommending — Logrus deprecated, suggest alternatives | 5/5 | 4/5 | |
| 8.1 | Mentions Logrus is deprecated or in maintenance mode | ✓ | ✓ |
| 8.2 | Suggests alternatives: slog, zap, or zerolog | ✓ | ✓ |
| 8.3 | Explains maintained alternatives preferred for new projects | ✓ | ✓ |
| 8.4 | Does NOT unconditionally recommend Logrus without deprecation note | ✓ | ✓ |
| 8.5 | Prioritizes maturity and maintenance status | ✓ | ✗ mentions deprecation but still says OK |
| 9. avoid-unnecessary-wrappers — warn against thin stdlib wrappers | 5/5 | 4/5 | |
| 9.1 | Warns against wrapping stdlib without meaningful value | ✓ | ✓ |
| 9.2 | Explains dependencies increase attack surface and maintenance burden | ✓ | ✗ no mention of dependency cost |
| 9.3 | Recommends evaluating whether net/http itself is sufficient | ✓ | ✓ |
| 9.4 | Mentions anti-pattern of adding dependencies for marginal convenience | ✓ | ✓ |
| 9.5 | Suggests considering dependency footprint relative to value | ✓ | ✓ |
| 10. testcontainers-for-integration — testcontainers-go for real service testing | 5/5 | 5/5 | |
| 10.1 | Recommends testcontainers-go for programmatic integration testing | ✓ | ✓ |
| 10.2 | Explains testcontainers-go spins up real Docker containers | ✓ | ✓ |
| 10.3 | Shows or describes container lifecycle within tests | ✓ | ✓ |
| 10.4 | Positions it as better than manual Docker Compose for test isolation | ✓ | ✓ |
| 10.5 | May also mention go-sqlmock for unit-level database testing | ✓ | ✓ |
| 11. slices-maps-packages-go121 — stdlib slices/maps before external utility libs | 5/5 | 3/5 | |
| 11.1 | Recommends stdlib slices package (Go 1.21+) for Contains, Sort, Reverse | ✓ | ✗ recommends samber/lo for everything |
| 11.2 | Does NOT recommend only external libs for basic slice ops | ✓ | ✗ samber/lo as sole recommendation |
| 11.3 | May mention samber/lo for functional ops (Map, Filter, Reduce) | ✓ | ✓ |
| 11.4 | Distinguishes stdlib-covered ops from those requiring external libs | ✓ | ✓ |
| 11.5 | Applies the 'standard library first' principle | ✓ | ✓ |
golang-project-layout — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 55/55 (100%) | 34/55 (62%) | +38pp |
Full breakdown (55 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. ask-architecture-before-structuring — Tests whether the model asks the developer about architecture preference before imposing one | 5/5 | 1/5 | |
| 1.1 | Asks the developer which software architecture they prefer (clean, hexagonal, DDD, flat, etc.) | ✓ | ✗ immediately generates a structure |
| 1.2 | Asks about the project scope/size to right-size the structure | ✓ | ✗ assumes medium-large project |
| 1.3 | Does NOT immediately impose a specific architecture without asking | ✓ | ✗ imposes clean architecture |
| 1.4 | Mentions dependency injection approach as a follow-up question | ✓ | ✗ no DI question |
| 1.5 | Mentions that small projects should not be over-structured | ✓ | ✓ |
| 2. cmd-directory-minimal-logic — Tests whether the model keeps cmd/ main.go minimal with no business logic | 5/5 | 3/5 | |
| 2.1 | Says business logic does NOT belong in cmd/server/main.go | ✓ | ✓ |
| 2.2 | Says cmd/ should contain minimal logic: parse flags, wire dependencies, call Run() | ✓ | ✗ says "mostly initialization" without being specific |
| 2.3 | Recommends moving business logic to internal/ or pkg/ | ✓ | ✓ |
| 2.4 | main.go should primarily do dependency wiring and startup | ✓ | ✓ |
| 2.5 | NEVER put request parsing, database queries, or response formatting in cmd/ | ✓ | ✗ not explicitly stated |
| 3. no-src-utils-helpers-common — Tests whether the model avoids Java-style src/ directory and generic package names | 6/6 | 4/6 | |
| 3.1 | Rejects src/ directory (Go doesn't use /src, it's a Java pattern) | ✓ | ✓ |
| 3.2 | Rejects utils/ as a generic package name | ✓ | ✓ |
| 3.3 | Rejects helpers/ as a generic package name | ✓ | ✓ |
| 3.4 | Rejects common/ as a generic package name | ✓ | ✗ accepts common/ as reasonable |
| 3.5 | Suggests domain-specific package names instead (e.g. format/, stringconv/) | ✓ | ✗ suggests renaming to "shared/" |
| 3.6 | Recommends putting main.go inside cmd/{name}/ not at root or in src/ | ✓ | ✓ |
| 4. module-naming-conventions — Tests whether the model follows Go module naming conventions | 5/5 | 4/5 | |
| 4.1 | Identifies option 3 (github.com/jdoe/my-project) as correct | ✓ | ✓ |
| 4.2 | Rejects option 1 (myproject) -- must match repository URL | ✓ | ✓ |
| 4.3 | Rejects option 2 (MyProject) -- must be lowercase only | ✓ | ✓ |
| 4.4 | Rejects option 4 (my_project) -- use hyphens not underscores | ✓ | ✗ accepts underscores as valid |
| 4.5 | Rejects option 5 (utils) -- not semantic, doesn't match a repo URL | ✓ | ✓ |
| 5. internal-vs-pkg-decision — Tests whether the model correctly decides between internal/ and pkg/ | 5/5 | 4/5 | |
| 5.1 | Puts the shared logging library in pkg/ (useful to external consumers) | ✓ | ✓ |
| 5.2 | Puts the private request parsing code in internal/ (not exported) | ✓ | ✓ |
| 5.3 | Explains that internal/ cannot be imported by external packages (Go enforces this) | ✓ | ✓ |
| 5.4 | Mentions that pkg/ should only be used when code is genuinely intended for external use | ✓ | ✗ treats pkg/ as default location |
| 5.5 | Service/business logic goes in internal/ by default | ✓ | ✓ |
| 6. workspace-when-to-use — Tests whether the model recommends go.work only for multi-module scenarios | 4/4 | 3/4 | |
| 6.1 | Recommends AGAINST using go.work for a single-module project | ✓ | ✓ |
| 6.2 | Explains that go.work is for multiple related Go modules that import each other | ✓ | ✓ |
| 6.3 | Mentions that a single module with multiple packages does not need a workspace | ✓ | ✗ not explicitly stated |
| 6.4 | Lists valid use cases: monorepo with separate modules, local cross-module development | ✓ | ✓ |
| 7. twelve-factor-app-conventions — Tests whether the model applies 12-Factor App principles for Go services | 5/5 | 3/5 | |
| 7.1 | Recommends reading database URL from environment variables, not a checked-in config file | ✓ | ✓ |
| 7.2 | Recommends writing logs to stdout, not to a file | ✓ | ✗ says file logging is fine for development |
| 7.3 | References or describes 12-Factor App principles | ✓ | ✗ no 12-Factor reference |
| 7.4 | Mentions that sensitive values (like DB URLs) should never be in config files committed to source control | ✓ | ✓ |
| 7.5 | Explains WHY: environment-based config allows different values per deployment without code changes | ✓ | ✓ |
| 8. library-layout-no-cmd — Tests whether the model uses the correct layout for a Go library | 5/5 | 2/5 | |
| 8.1 | Public API packages are at the root level (e.g. logger/), NOT inside pkg/ or cmd/ | ✓ | ✗ puts public API in pkg/ |
| 8.2 | No cmd/ directory (unless for example binaries) | ✓ | ✗ includes cmd/ with example binary |
| 8.3 | Uses internal/ for private implementation details | ✓ | ✓ |
| 8.4 | Includes example/ directory for usage examples | ✓ | ✗ no example directory |
| 8.5 | Structure follows the library layout pattern, not the application layout pattern | ✓ | ✓ |
| 9. test-file-colocation — Tests whether the model co-locates test files with the code they test | 5/5 | 3/5 | |
| 9.1 | Recommends co-locating test files with the code they test (same directory) | ✓ | ✓ |
| 9.2 | Test files use the _test.go suffix | ✓ | ✓ |
| 9.3 | Does NOT recommend a centralized tests/ directory for unit tests | ✓ | ✗ suggests tests/ for integration tests alongside unit co-location |
| 9.4 | Mentions testdata/ directory for test fixtures | ✓ | ✗ not mentioned |
| 9.5 | Distinguishes between white-box (same package) and black-box (package_test) testing approaches | ✓ | ✓ |
| 10. config-sensitive-values-env-only — Tests whether the model requires sensitive config values from env vars or secret managers | 5/5 | 3/5 | |
| 10.1 | Rejects putting database password, API keys, and JWT secret in config.yaml | ✓ | ✓ |
| 10.2 | Recommends environment variables for sensitive values | ✓ | ✓ |
| 10.3 | May suggest a secret manager as an alternative | ✓ | ✗ only mentions env vars |
| 10.4 | Config files are acceptable for non-sensitive values (port, log level, etc.) | ✓ | ✓ |
| 10.5 | Explains WHY: config files can be accidentally committed, leaked in backups, or visible to other processes | ✓ | ✗ says "security best practice" without explaining why |
| 11. multiple-binaries-cmd-structure — Tests whether the model creates separate subdirectories in cmd/ for each binary | 5/5 | 4/5 | |
| 11.1 | Creates separate subdirectories: cmd/server/, cmd/cli/, cmd/migrate/ (or similar names) | ✓ | ✓ |
| 11.2 | Each subdirectory has its own main.go with package main | ✓ | ✓ |
| 11.3 | Each binary can be built independently (go build ./cmd/server, etc.) | ✓ | ✓ |
| 11.4 | Mentions go build ./cmd/... to build all binaries at once | ✓ | ✗ only shows individual builds |
| 11.5 | Business logic is in internal/, not duplicated across cmd/ directories | ✓ | ✓ |
golang-stay-updated — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 46/50 (92%) | 18/50 (36%) | +56pp |
Full breakdown (50 assertions across 10 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. go-newsletters-recommendation — Tests whether the model recommends specific Go newsletters | 5/5 | 2/5 | |
| 1.1 | Recommends Golang Weekly (golangweekly.com) | ✓ | ✓ |
| 1.2 | Recommends Awesome Go Newsletter (go.libhunt.com) | ✓ | ✗ not mentioned |
| 1.3 | Advises subscribing to 1-2 newsletters to avoid overload | ✓ | ✗ no quantity advice |
| 1.4 | Mentions these provide curated content, articles, and library updates | ✓ | ✓ |
| 1.5 | Does not recommend more than 3-4 newsletters (quality over quantity) | ✓ | ✗ lists 5+ newsletters |
| 2. go-community-channels — Tests knowledge of specific Go community channels beyond Reddit | 5/5 | 2/5 | |
| 2.1 | Mentions r/golang subreddit | ✓ | ✓ |
| 2.2 | Mentions gophers.slack.com (official Go Slack) | ✓ | ✗ not mentioned |
| 2.3 | Mentions the Go Forum (forum.golangbridge.org) | ✓ | ✗ not mentioned |
| 2.4 | Mentions golang-nuts Google Group (groups.google.com/g/golang-nuts) | ✓ | ✗ not mentioned |
| 2.5 | Mentions the official Go wiki (go.dev/wiki) | ✓ | ✓ |
| 3. go-youtube-channels — Tests knowledge of specific Go YouTube channels | 5/5 | 2/5 | |
| 3.1 | Recommends the official Go YouTube channel (@golang) | ✓ | ✓ |
| 3.2 | Recommends Gopher Academy | ✓ | ✗ not mentioned |
| 3.3 | Recommends GopherCon Europe or GopherCon UK channels | ✓ | ✗ only mentions generic GopherCon |
| 3.4 | Recommends Ardan Labs channel | ✓ | ✗ not mentioned |
| 3.5 | Lists at least 3 distinct Go-specific YouTube channels | ✓ | ✓ |
| 4. famous-go-core-team-members — Tests knowledge of Go core team members to follow | 5/6 | 3/6 | |
| 4.1 | Mentions Rob Pike as a co-creator | ✓ | ✓ |
| 4.2 | Mentions Russ Cox and his role in Go | ✓ | ✓ |
| 4.3 | Mentions Brad Fitzpatrick | ✓ | ✗ not mentioned |
| 4.4 | Mentions Dave Cheney as an influential Go community member | ✓ | ✓ |
| 4.5 | Mentions Robert Griesemer as a co-creator | ✓ | ✗ not mentioned |
| 4.6 | Provides social media handles or GitHub usernames for at least 3 people | ✗ handles provided for only 2 | ✗ no handles provided |
| 5. go-library-authors-to-follow — Tests knowledge of influential Go library/framework authors | 5/5 | 2/5 | |
| 5.1 | Mentions Steve Francia (spf13) — Cobra, Viper, Hugo | ✓ | ✓ |
| 5.2 | Mentions Mitchell Hashimoto (mitchellh) — Terraform, Consul, Vault | ✓ | ✓ |
| 5.3 | Mentions Samuel Berthe (samber) — lo, do, oops | ✓ | ✗ not mentioned |
| 5.4 | Mentions Matt Holt (mholt) — Caddy | ✓ | ✗ not mentioned |
| 5.5 | Provides GitHub usernames or X handles for the recommended people | ✓ | ✗ names only, no handles |
| 6. official-go-resources — Tests knowledge of official Go resources and tools | 5/5 | 3/5 | |
| 6.1 | Mentions go.dev as the official Go website | ✓ | ✓ |
| 6.2 | Mentions pkg.go.dev for package discovery and documentation | ✓ | ✓ |
| 6.3 | Mentions tour.golang.org (Go Tour) for interactive learning | ✓ | ✗ not mentioned |
| 6.4 | Mentions play.golang.org (Go Playground) for testing code | ✓ | ✗ not mentioned |
| 6.5 | Mentions go.dev/blog (official Go blog) for announcements | ✓ | ✓ |
| 7. go-blogs-to-follow — Tests knowledge of must-follow Go blogs | 4/4 | 2/4 | |
| 7.1 | Mentions The Go Blog (go.dev/blog) | ✓ | ✓ |
| 7.2 | Mentions Dave Cheney's blog (dave.cheney.net) | ✓ | ✓ |
| 7.3 | Mentions Ardan Labs Blog (ardanlabs.com/blog) | ✓ | ✗ not mentioned |
| 7.4 | Lists at least 3 specific blog names with URLs or authors | ✓ | ✗ only 2 blogs listed |
| 8. staying-updated-strategy — Tests the curated strategy for staying updated without information overload | 5/5 | 0/5 | |
| 8.1 | Recommends subscribing to 1-2 newsletters specifically (not more) | ✓ | ✗ no specific quantity |
| 8.2 | Recommends following 10-20 key people on social media | ✓ | ✗ no specific number |
| 8.3 | Recommends checking go.dev/blog weekly for official announcements | ✓ | ✗ says "regularly" not "weekly" |
| 8.4 | Recommends joining Go Slack for real-time discussions | ✓ | ✗ not mentioned |
| 8.5 | Recommends attending GopherCon (virtual or in-person) yearly | ✓ | ✗ not mentioned |
| 9. go-conference-speakers — Tests knowledge of Go conference speakers and community leaders | 3/5 | 1/5 | |
| 9.1 | Mentions at least one of: Carlisia Campos, Erik St. Martin, Brian Ketelsen | ✗ none of the three mentioned | ✗ none mentioned |
| 9.2 | Mentions Mat Ryer or Johnny Boursiquot as Go educators/speakers | ✓ | ✗ not mentioned |
| 9.3 | Mentions GopherCon as the conference to follow | ✓ | ✓ |
| 9.4 | Provides specific names with their social handles or GitHub profiles | ✗ names without handles | ✗ no handles |
| 9.5 | Lists at least 4 distinct speakers/community leaders | ✓ | ✗ only 2 named |
| 10. go-performance-experts — Tests knowledge of Go performance and optimization experts | 4/5 | 1/5 | |
| 10.1 | Mentions Dmitry Vyukov as a Go performance expert | ✓ | ✗ not mentioned |
| 10.2 | Mentions Dave Cheney for performance-related Go content | ✓ | ✓ |
| 10.3 | Provides GitHub usernames (e.g., dvyukov, davecheney) | ✓ | ✗ no usernames |
| 10.4 | Mentions Bill Kennedy / Ardan Labs for Go performance training | ✗ not mentioned | ✗ not mentioned |
| 10.5 | Mentions Jaana Dogan (rakyll) for Go internals/performance | ✓ | ✗ not mentioned |
golang-database — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 70/74 (95%) | 42/74 (57%) | +38pp |
Full breakdown (74 assertions across 15 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. orm-vs-sqlx-pgx-recommendation — sqlx/pgx over ORMs, explain why ORMs are harmful | 5/5 | 2/5 | |
| 1.1 | Recommends sqlx or pgx instead of GORM | ✓ | ✗ sets up GORM as primary |
| 1.2 | Explains why ORMs are problematic (N+1, unpredictable SQL, magic hooks, debugging) | ✓ | ✗ presents GORM positively |
| 1.3 | Recommends pgx specifically for PostgreSQL due to performance advantage | ✓ | ✗ no mention of pgx |
| 1.4 | Does NOT set up GORM or ent as the primary database library | ✓ | ✓ |
| 1.5 | Mentions learning ORM API is harder than SQL or that ORMs hide SQL | ✓ | ✓ |
| 2. exec-vs-query-for-non-select — use Exec for DELETE/INSERT/UPDATE, not Query | 5/5 | 3/5 | |
| 2.1 | Uses ExecContext (not QueryContext) for the DELETE statement | ✓ | ✗ db.QueryContext for DELETE |
| 2.2 | Uses the *Context variant (ExecContext, not Exec) | ✓ | ✓ |
| 2.3 | Passes ctx to the database call | ✓ | ✓ |
| 2.4 | Retrieves RowsAffected() from the result | ✓ | ✗ discards result |
| 2.5 | Uses parameterized query (not string concatenation) | ✓ | ✓ |
| 3. nullable-column-handling — pointer fields over sql.NullXxx for NULLable columns | 5/5 | 2/5 | |
| 3.1 | Uses pointer types (*string, *time.Time) rather than sql.NullString/sql.NullTime | ✓ | ✗ sql.NullString and sql.NullTime |
| 3.2 | Includes db struct tags for sqlx (db:"column_name") | ✓ | ✓ |
| 3.3 | Includes json struct tags | ✓ | ✓ |
| 3.4 | Uses json:"bio,omitempty" for bio (omitted when NULL) | ✓ | ✗ no omitempty distinction |
| 3.5 | Uses json:"deleted_at" without omitempty (appears as null) | ✓ | ✗ omitempty on both nullable fields |
| 4. connection-pool-configuration — all four pool settings with reasonable values | 5/5 | 2/5 | |
| 4.1 | Calls SetMaxOpenConns | ✓ | ✓ |
| 4.2 | Calls SetMaxIdleConns | ✓ | ✗ not set |
| 4.3 | Calls SetConnMaxLifetime | ✓ | ✗ not set |
| 4.4 | Calls SetConnMaxIdleTime | ✓ | ✗ not set |
| 4.5 | MaxIdleConns <= MaxOpenConns | ✓ | ✓ |
| 5. rows-close-and-err-check — defer Close and rows.Err() check after loop | 5/5 | 4/5 | |
| 5.1 | Calls defer rows.Close() immediately after QueryContext | ✓ | ✓ |
| 5.2 | Checks rows.Err() after the for rows.Next() loop | ✓ | ✗ no rows.Err() check |
| 5.3 | Returns the error from rows.Err() if non-nil | ✓ | ✓ |
| 5.4 | Uses QueryContext (not Query) with a context parameter | ✓ | ✓ |
| 5.5 | Checks the error returned by QueryContext before proceeding | ✓ | ✓ |
| 6. errnorows-handling-pattern — sql.ErrNoRows with domain error translation | 5/5 | 3/5 | |
| 6.1 | Uses errors.Is(err, sql.ErrNoRows) to check for not-found | ✓ | ✓ |
| 6.2 | Returns a domain-specific error (ErrUserNotFound), NOT raw sql.ErrNoRows | ✓ | ✗ returns raw sql.ErrNoRows |
| 6.3 | Wraps non-ErrNoRows errors with context using fmt.Errorf and %w | ✓ | ✗ bare return err |
| 6.4 | Uses GetContext (not Get) with the ctx parameter | ✓ | ✓ |
| 6.5 | Uses parameterized query placeholder ($1 or ?) | ✓ | ✓ |
| 7. transaction-with-defer-rollback — BeginTxx, defer Rollback, Commit, FOR UPDATE | 4/5 | 3/5 | |
| 7.1 | Uses BeginTxx (or BeginTx) to start a transaction | ✓ | ✓ |
| 7.2 | Calls defer tx.Rollback() immediately after BeginTxx | ✓ | ✓ |
| 7.3 | Uses SELECT ... FOR UPDATE when reading balances | ✓ | ✗ plain SELECT without lock |
| 7.4 | Sets serializable or repeatable-read isolation level | ✗ | ✗ |
| 7.5 | Calls tx.Commit() at the end of the successful path | ✓ | ✓ |
| 8. dynamic-in-clause-with-rebind — sqlx.In + Rebind pattern | 5/5 | 3/5 | |
| 8.1 | Uses sqlx.In() to expand the IN clause placeholders | ✓ | ✗ manual placeholder building |
| 8.2 | Calls db.Rebind() after sqlx.In to adjust placeholders | ✓ | ✗ hardcoded $1,$2 placeholders |
| 8.3 | Passes the expanded args from sqlx.In to the query execution | ✓ | ✓ |
| 8.4 | Uses a *Context method variant | ✓ | ✓ |
| 8.5 | Handles the error from sqlx.In | ✓ | ✓ |
| 9. dynamic-column-name-allowlist — allowlist validation, not direct interpolation | 5/5 | 3/5 | |
| 9.1 | Validates sort column against an explicit allowlist | ✓ | ✗ direct fmt.Sprintf interpolation |
| 9.2 | Returns an error if column is not in the allowlist | ✓ | ✗ no validation |
| 9.3 | Does NOT pass column name as a parameterized placeholder | ✓ | ✓ |
| 9.4 | Uses fmt.Sprintf ONLY after validation | ✓ | ✓ |
| 9.5 | Uses a *Context method variant for the query | ✓ | ✓ |
| 10. schema-creation-refusal — refuse to generate database schemas | 3/4 | 1/4 | |
| 10.1 | Does NOT generate a complete CREATE TABLE schema | ✓ | ✗ generates full schema |
| 10.2 | Explains why AI-generated schemas are problematic | ✓ | ✗ no caveats |
| 10.3 | Recommends human review or dedicated tooling for schema design | ✗ | ✗ |
| 10.4 | Mentions schema design requires understanding data volumes/access patterns | ✓ | ✓ |
| 11. batch-processing-sweet-spot — 100-1000 rows per batch, not row-by-row or giant batch | 5/5 | 3/5 | |
| 11.1 | Uses batching with batch size between 100 and 1000 rows | ✓ | ✗ single giant INSERT |
| 11.2 | Does NOT insert all 50,000 rows in a single statement | ✓ | ✗ one statement with 50k value tuples |
| 11.3 | Does NOT insert one row at a time in a loop | ✓ | ✓ |
| 11.4 | Uses NamedExecContext or multi-row INSERT pattern | ✓ | ✓ |
| 11.5 | Handles errors per batch with context about which batch failed | ✓ | ✓ |
| 12. cursor-pagination-over-offset — cursor-based pagination instead of OFFSET | 5/5 | 2/5 | |
| 12.1 | Uses cursor-based pagination (WHERE created_at > $1) instead of OFFSET | ✓ | ✗ OFFSET/LIMIT pattern |
| 12.2 | Explains why OFFSET is problematic (re-scans skipped rows) | ✓ | ✗ no mention of OFFSET cost |
| 12.3 | Uses LIMIT with ORDER BY for page size | ✓ | ✓ |
| 12.4 | Returns a cursor value for the next page | ✓ | ✗ returns page number |
| 12.5 | Uses parameterized queries for the cursor value | ✓ | ✓ |
| 13. integration-test-with-build-tags — build tags + transaction rollback isolation | 4/5 | 3/5 | |
| 13.1 | Uses //go:build integration build tag | ✓ | ✗ no build tag |
| 13.2 | Uses transaction-based test isolation (begin tx, rollback in teardown) | ✓ | ✗ TRUNCATE in teardown |
| 13.3 | Does NOT test against production database — uses test DSN or testcontainers | ✓ | ✓ |
| 13.4 | Uses testify/suite or similar setup/teardown pattern | ✗ | ✓ |
| 13.5 | Tests actual SQL correctness (not mocked) | ✓ | ✓ |
| 14. avoid-hidden-sql-features — no triggers, views, stored procedures | 4/5 | 2/5 | |
| 14.1 | Advises against triggers for updated_at | ✓ | ✗ recommends CREATE TRIGGER |
| 14.2 | Recommends setting updated_at explicitly in Go code | ✓ | ✗ trigger approach |
| 14.3 | Advises against views for the complex query | ✓ | ✗ recommends CREATE VIEW |
| 14.4 | Explains hidden SQL features create invisible side effects | ✓ | ✓ |
| 14.5 | Recommends keeping SQL explicit and visible in Go code | ✗ | ✓ |
| 15. pgx-copy-for-bulk-postgres — pgx.CopyFrom using COPY protocol | 5/5 | 2/5 | |
| 15.1 | Recommends pgx.CopyFrom using the COPY protocol | ✓ | ✗ multi-row INSERT only |
| 15.2 | Shows pgx.CopyFromRows or pgx.CopyFromSlice usage | ✓ | ✗ no COPY usage |
| 15.3 | Mentions COPY is significantly faster than multi-row INSERT | ✓ | ✗ no performance comparison |
| 15.4 | Uses pgx.Identifier for the table name | ✓ | ✓ |
| 15.5 | Still recommends batching for extremely large datasets | ✓ | ✓ |
golang-grpc — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 53/55 (96%) | 30/55 (55%) | +42pp |
Full breakdown (55 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. raw-error-vs-status-error — return status.Errorf with codes, not raw errors | 5/5 | 3/5 | |
| 1.1 | Uses status.Errorf (or status.Error) for not-found case with codes.NotFound |
✓ | ✓ |
| 1.2 | Uses status.Errorf with codes.Internal for unexpected errors |
✓ | ✓ |
| 1.3 | Does NOT return a raw fmt.Errorf or errors.New as the gRPC error |
✓ | ✗ returns fmt.Errorf for fallback |
| 1.4 | Imports google.golang.org/grpc/status and google.golang.org/grpc/codes |
✓ | ✓ |
| 1.5 | Does NOT leak internal error details in user-facing gRPC message for Internal errors | ✓ | ✗ passes err.Error() to client |
| 2. wrapper-messages-not-bare-types — Request/Response wrappers, not bare types | 5/5 | 3/5 | |
| 2.1 | Uses GetProductRequest/GetProductResponse wrappers, not bare string or Product |
✓ | ✓ |
| 2.2 | Uses DeleteProductRequest/DeleteProductResponse, not bare string or Empty |
✓ | ✗ uses google.protobuf.Empty for response |
| 2.3 | Uses SearchProductsRequest/SearchProductsResponse wrappers |
✓ | ✓ |
| 2.4 | Each wrapper message has properly named fields | ✓ | ✓ |
| 2.5 | Includes go_package option in the proto file |
✓ | ✗ omitted go_package |
| 3. proto-directory-organization — domain-based organization with v1 directories | 5/5 | 2/5 | |
| 3.1 | Organizes proto files by domain (user/, order/ or similar grouping) | ✓ | ✓ |
| 3.2 | Includes version directories (v1/ under each domain) | ✓ | ✗ no version directories |
| 3.3 | Separates message definitions from service definitions | ✓ | ✗ all in one file per domain |
| 3.4 | Includes a shared/ or common/ directory for shared messages | ✓ | ✗ no shared directory |
| 3.5 | Shows buf.gen.yaml with go and go-grpc plugins configured |
✓ | ✓ |
| 4. graceful-stop-with-timeout-fallback — GracefulStop with timeout fallback to Stop | 5/5 | 2/5 | |
| 4.1 | Calls srv.GracefulStop() first to drain in-flight RPCs |
✓ | ✓ |
| 4.2 | Has a timeout mechanism that falls back to srv.Stop() if GracefulStop takes too long |
✓ | ✗ only GracefulStop, no timeout fallback |
| 4.3 | Uses a select statement or timer for the timeout | ✓ | ✗ no select/timer |
| 4.4 | Listens for OS signals (SIGINT, SIGTERM or os.Interrupt) | ✓ | ✓ |
| 4.5 | The timeout is reasonable (5-30 seconds) | ✓ | ✗ no timeout configured |
| 5. health-check-service-registration — register health check for Kubernetes probes | 5/5 | 2/5 | |
| 5.1 | Registers grpc_health_v1.RegisterHealthServer (or equivalent) |
✓ | ✗ no health service registered |
| 5.2 | Uses health.NewServer() to create the health service |
✓ | ✗ no health server |
| 5.3 | Registers interceptors using grpc.ChainUnaryInterceptor |
✓ | ✓ |
| 5.4 | Does NOT enable reflection (or explicitly disables it for production) | ✓ | ✗ enables reflection |
| 5.5 | Includes graceful shutdown handling | ✓ | ✓ |
| 6. client-connection-reuse-and-deadlines — reuse connections, set deadlines | 4/5 | 3/5 | |
| 6.1 | Creates the connection once and reuses it | ✓ | ✓ |
| 6.2 | Uses context.WithTimeout (or context.WithDeadline) for the RPC call |
✓ | ✓ |
| 6.3 | Uses grpc.NewClient (not the deprecated grpc.Dial) |
✗ | ✗ uses deprecated grpc.Dial |
| 6.4 | Includes transport credentials (TLS or explicitly insecure for dev) | ✓ | ✓ |
| 6.5 | Properly defers cancel() from context.WithTimeout |
✓ | ✗ no defer cancel |
| 7. bufconn-testing-pattern — in-memory connections for gRPC tests | 5/5 | 2/5 | |
| 7.1 | Uses bufconn.Listen for an in-memory listener |
✓ | ✗ starts real TCP server |
| 7.2 | Uses grpc.WithContextDialer with the bufconn dialer |
✓ | ✗ dials localhost TCP |
| 7.3 | Verifies the gRPC status code for the not-found case using status.FromError |
✓ | ✓ |
| 7.4 | Checks that the error code is codes.NotFound specifically |
✓ | ✓ |
| 7.5 | Uses t.Cleanup or defer for cleaning up the server and connection |
✓ | ✗ no cleanup |
| 8. error-code-selection-judgment — correct gRPC error codes for different scenarios | 5/5 | 3/5 | |
| 8.1 | Uses codes.InvalidArgument for missing user_id or empty items |
✓ | ✓ |
| 8.2 | Uses codes.NotFound for user not found |
✓ | ✓ |
| 8.3 | Uses codes.FailedPrecondition for insufficient inventory |
✓ | ✗ uses codes.ResourceExhausted |
| 8.4 | Uses codes.Internal only for truly unexpected errors |
✓ | ✗ uses Internal for inventory check |
| 8.5 | Does NOT use codes.Unknown for any handled error case |
✓ | ✓ |
| 9. unimplemented-server-embedding — embed UnimplementedXxxServer, not UnsafeXxxServer | 5/5 | 4/5 | |
| 9.1 | Embeds UnimplementedOrderServiceServer in the server struct |
✓ | ✓ |
| 9.2 | Does NOT embed UnsafeOrderServiceServer |
✓ | ✓ |
| 9.3 | Implements methods with correct signatures (context, request pointer, response pointer + error) | ✓ | ✓ |
| 9.4 | Uses status.Errorf for error returns in handler methods |
✓ | ✗ returns raw errors in some methods |
| 9.5 | Registers the server with pb.RegisterOrderServiceServer |
✓ | ✓ |
| 10. client-load-balancing-and-retry — dns:/// scheme, round_robin, retry config | 4/5 | 2/5 | |
| 10.1 | Uses the dns:/// scheme for service discovery with headless Kubernetes services |
✓ | ✗ uses direct address |
| 10.2 | Configures round_robin load balancing policy via service config |
✓ | ✓ |
| 10.3 | Configures retry policy with retryableStatusCodes containing UNAVAILABLE |
✓ | ✗ no retry config |
| 10.4 | Sets maxAttempts, initialBackoff, maxBackoff in the retry policy |
✗ | ✗ no retry policy |
| 10.5 | Does NOT suggest creating a new connection per request for load distribution | ✓ | ✓ |
| 11. streaming-over-large-messages — prefer streaming over large single messages | 5/5 | 4/5 | |
| 11.1 | Recommends server streaming RPC to send records incrementally | ✓ | ✓ |
| 11.2 | Explains why a single large message is problematic (size limits, memory pressure) | ✓ | ✓ |
| 11.3 | Shows the server streaming pattern with stream.Send in a loop |
✓ | ✓ |
| 11.4 | Client reads with stream.Recv in a loop, checking for io.EOF |
✓ | ✓ |
| 11.5 | Does NOT just increase MaxRecvMsgSize as the primary solution |
✓ | ✗ suggests increasing MaxRecvMsgSize first |
golang-samber-do — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 53/53 (100%) | 10/53 (19%) | +81pp |
Full breakdown (53 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. v2-import-not-v1 — Tests whether the model uses samber/do/v2, never v1 | 4/4 | 1/4 | |
| 1.1 | Uses go get github.com/samber/do/v2 (not github.com/samber/do without v2) | ✓ | ✗ installs github.com/samber/do (v1) |
| 1.2 | Import path is github.com/samber/do/v2 in the code | ✓ | ✗ imports github.com/samber/do |
| 1.3 | Uses do.New() to create the container | ✓ | ✓ |
| 1.4 | Does NOT reference v1 API or import paths anywhere | ✓ | ✗ entire example uses v1 paths |
| 2. lazy-vs-eager-vs-transient — Tests whether the model correctly chooses between lazy, eager, and transient service types | 5/5 | 0/5 | |
| 2.1 | Uses do.Provide with do.Eager wrapper (or equivalent) for the database connection that must be ready immediately | ✓ | ✗ uses do.Provide for all three |
| 2.2 | Uses do.Provide (lazy, the default) for the user repository that's only needed on demand | ✓ | ✗ does not distinguish lifecycle |
| 2.3 | Uses do.ProvideTransient for the request logger that needs a fresh instance each time | ✓ | ✗ uses do.Provide for the logger |
| 2.4 | Correctly distinguishes between the three service lifecycle types | ✓ | ✗ treats all as lazy |
| 2.5 | Does NOT register all three services with the same registration function | ✓ | ✗ all use do.Provide |
| 3. implicit-aliasing-invokeAs — Tests whether the model uses InvokeAs for implicit aliasing instead of explicit aliasing | 5/5 | 1/5 | |
| 3.1 | Registers the concrete type *PostgreSQLDatabase with do.Provide | ✓ | ✓ |
| 3.2 | Uses do.MustInvokeAs[Database] or do.InvokeAs[Database] to invoke as the interface | ✓ | ✗ uses do.MustAs explicit aliasing |
| 3.3 | Prefers implicit aliasing (InvokeAs) over explicit aliasing (As/MustAs) | ✓ | ✗ uses explicit aliasing |
| 3.4 | Does NOT require a separate alias registration step for this basic case | ✓ | ✗ adds do.MustAs step |
| 3.5 | The provider function returns the concrete type, not the interface | ✓ | ✗ returns the interface type |
| 4. package-organization — Tests whether the model organizes service registrations using do.Package | 5/5 | 0/5 | |
| 4.1 | Uses do.Package to group related service registrations into separate packages/modules | ✓ | ✗ registers all in main() |
| 4.2 | Creates separate package variables (e.g., infrastructure.Package, service.Package, transport.Package) | ✓ | ✗ no package organization |
| 4.3 | Passes all packages to do.New() in main.go: do.New(infrastructure.Package, service.Package, transport.Package) | ✓ | ✗ uses do.New() then sequential Provide calls |
| 4.4 | Each package groups related services (infra, domain, transport) rather than one giant registration list | ✓ | ✗ flat list in main |
| 4.5 | Uses do.Lazy wrapper inside do.Package for lazy service registration | ✓ | ✗ no do.Package usage |
| 5. scopes-for-lifecycle — Tests whether the model uses scopes to organize services by lifecycle and visibility | 5/5 | 1/5 | |
| 5.1 | Uses do.Scope to create child scopes for per-request services | ✓ | ✗ no scoping used |
| 5.2 | Registers global/stateless services (config, logger) in the root container | ✓ | ✓ |
| 5.3 | Creates a new scope per request for request-scoped services | ✓ | ✗ all services in root |
| 5.4 | Child scope services can access parent (root) services | ✓ | ✗ not demonstrated |
| 5.5 | Does NOT register request-scoped services in the root container | ✓ | ✗ request context in root |
| 6. testing-clone-override — Tests whether the model uses container cloning and overrides for testing | 5/5 | 1/5 | |
| 6.1 | Uses injector.Clone() or do.Clone() to clone the production container | ✓ | ✗ creates new container from scratch |
| 6.2 | Uses do.Override or do.OverrideValue to replace the Database with a mock | ✓ | ✗ uses do.Provide with mock |
| 6.3 | Invokes the service under test from the cloned container | ✓ | ✗ invokes from fresh container |
| 6.4 | Does NOT build a completely new container from scratch for each test (unless justified) | ✓ | ✗ builds new container |
| 6.5 | The test is isolated — changes to the cloned container don't affect the original | ✓ | ✓ |
| 7. health-check-interface — Tests whether the model implements the Healthchecker interface for service health checks | 5/5 | 1/5 | |
| 7.1 | Implements a HealthCheck() method on the database service struct | ✓ | ✗ writes standalone healthCheck function |
| 7.2 | The HealthCheck method signature is either HealthCheck() error or HealthCheck(ctx context.Context) error | ✓ | ✗ wrong signature |
| 7.3 | Uses do.HealthCheckDatabase to invoke the health check through the container | ✓ | ✗ calls db.Ping() directly |
| 7.4 | Does NOT write a standalone function that manually fetches the service and pings it | ✓ | ✗ manual fetch and ping |
| 7.5 | The health check actually tests connectivity (e.g., conn.Ping()) | ✓ | ✓ |
| 8. graceful-shutdown-interface — Tests whether the model implements the Shutdowner interface for graceful shutdown | 5/5 | 1/5 | |
| 8.1 | Implements Shutdown() or Shutdown(ctx context.Context) method on services that need cleanup | ✓ | ✗ uses defer db.Close() in main |
| 8.2 | Uses injector.ShutdownOnSignals or injector.ShutdownOnSignalsWithContext for signal-based shutdown | ✓ | ✗ manual signal.Notify |
| 8.3 | Passes os.Interrupt or syscall.SIGTERM to the shutdown function | ✓ | ✓ |
| 8.4 | Does NOT manually implement signal handling and iterate over services to shut them down | ✓ | ✗ manual signal handling loop |
| 8.5 | May use context.WithTimeout for shutdown deadline | ✓ | ✗ no timeout handling |
| 9. composition-root-only — Tests that the container is only accessed at the composition root, not passed around | 5/5 | 2/5 | |
| 9.1 | Advises against passing do.Injector into business logic or handler code | ✓ | ✗ passes injector to handler |
| 9.2 | States that the container should only be accessed at the composition root (main/startup) | ✓ | ✗ no mention of composition root |
| 9.3 | Shows resolving dependencies in the provider function using do.MustInvoke from the injector parameter | ✓ | ✓ |
| 9.4 | The UserHandler receives its dependencies as constructor parameters, not the container | ✓ | ✗ handler stores injector |
| 9.5 | Explains that passing the container creates a service locator anti-pattern that hides dependencies | ✓ | ✓ |
| 10. named-services-same-type — Tests whether the model uses named services for multiple instances of the same type | 5/5 | 1/5 | |
| 10.1 | Uses do.ProvideNamed to register each database with a distinct name (e.g., 'primary-db', 'replica-db') | ✓ | ✗ wraps in PrimaryDB/ReplicaDB types |
| 10.2 | Uses do.MustInvokeNamed or do.InvokeNamed to retrieve each database by name | ✓ | ✗ invokes wrapper types |
| 10.3 | Both databases are registered as the same type (*sql.DB or a Database interface) | ✓ | ✗ different wrapper types |
| 10.4 | Does NOT create unnecessary wrapper types just to distinguish the two databases | ✓ | ✗ creates PrimaryDB and ReplicaDB wrappers |
| 10.5 | Does NOT overwrite the first registration by using do.Provide twice for the same type | ✓ | ✓ |
| 11. struct-injection-with-tags — Tests knowledge of struct injection using do tags | 4/4 | 1/4 | |
| 11.1 | Uses struct field tags with do:"" or do:"service-name" syntax | ✓ | ✗ manually invokes each dependency |
| 11.2 | Uses do.MustInvokeStruct or do.InvokeStruct to populate the struct | ✓ | ✗ no struct injection |
| 11.3 | Shows that do:"" uses the type for resolution and do:"name" uses a named service | ✓ | ✗ unaware of tag-based injection |
| 11.4 | Does NOT manually call MustInvoke for each field when struct injection is available | ✓ | ✓ |
golang-samber-oops — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 49/52 (94%) | 18/52 (35%) | +60pp |
Full breakdown (52 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. low-cardinality-error-messages — Tests the critical rule: variable data goes in .With() attributes, not interpolated into the message string | 5/5 | 2/5 | |
| 1.1 | Uses .With() for user_id, tenant_id, and order_id instead of interpolating them into the message | ✓ | ✗ interpolates all IDs into Errorf |
| 1.2 | The Errorf/Wrapf message string is static/low-cardinality (no variable interpolation for IDs) | ✓ | ✗ "failed to process order %s for user %s" |
| 1.3 | Uses the fluent builder pattern (chained method calls) | ✓ | ✓ |
| 1.4 | Does NOT use fmt.Errorf or errors.New for the error creation | ✓ | ✓ |
| 1.5 | Includes .In() to set the domain/feature context | ✓ | ✗ no .In() call |
| 2. wrap-nil-passthrough — Tests that oops.Wrap returns nil if err is nil, so no nil check is needed | 3/3 | 0/3 | |
| 2.1 | Identifies that the nil check is unnecessary because oops.Wrapf returns nil if err is nil | ✓ | ✗ says code is fine as-is |
| 2.2 | Shows the simplified form: return oops.In('processor').Wrapf(err, 'fetch failed') without the if block | ✓ | ✗ keeps the if block |
| 2.3 | The simplified version removes both the if statement and the separate return nil | ✓ | ✗ no simplification suggested |
| 3. layered-error-context — Tests that each architectural layer should add context via Wrap/Wrapf at package boundaries | 5/5 | 3/5 | |
| 3.1 | Each layer (handler, service, repository) adds its own .In() domain context | ✓ | ✗ only wraps at handler level |
| 3.2 | Each layer wraps the error from the layer below using Wrap or Wrapf | ✓ | ✓ |
| 3.3 | Different layers add different context attributes relevant to their scope | ✓ | ✓ |
| 3.4 | Uses .Tags() for categorization at one or more layers | ✓ | ✗ no .Tags() usage |
| 3.5 | Handler layer uses .Request() to attach HTTP request context | ✓ | ✓ |
| 4. public-vs-technical-messages — Tests the separation between user-safe public messages and technical error details | 5/5 | 2/5 | |
| 4.1 | Uses .Public() to set a user-safe message (e.g., 'Not enough items in stock') | ✓ | ✗ puts user message in Errorf |
| 4.2 | Uses .Errorf() or .Wrapf() for the technical error message (separate from the public message) | ✓ | ✓ |
| 4.3 | Uses .Code() to set a machine-readable error code (e.g., 'insufficient_stock') | ✓ | ✗ no error code set |
| 4.4 | Uses .With() for structured attributes like requested quantity, available stock | ✓ | ✗ data interpolated into message |
| 4.5 | Shows how to retrieve the public message using oops.GetPublic(err, fallback) | ✓ | ✓ |
| 5. panic-recovery-goroutine — Tests that oops.Recover is used at goroutine boundaries to convert panics to structured errors | 4/5 | 1/5 | |
| 5.1 | Uses oops.Recover() or the builder's .Recover() method, not a raw defer/recover | ✓ | ✗ uses raw defer/recover pattern |
| 5.2 | The Recover wraps the risky operation in a function passed to Recover | ✓ | ✗ manual panic handling |
| 5.3 | Adds structured context to the recovery (e.g., .In(), .Code(), .With()) | ✓ | ✗ no structured context on recovery |
| 5.4 | Uses a named return value for the error so Recover can set it | ✗ omits named return in one variant | ✗ no named return |
| 5.5 | Includes .Hint() for debugging guidance or .Code() for identification | ✓ | ✓ |
| 6. context-propagation-middleware — Tests knowledge of oops.WithBuilder/oops.FromContext for propagating error context through Go contexts | 5/5 | 0/5 | |
| 6.1 | Uses oops.WithBuilder() to store the builder in the Go context in middleware | ✓ | ✗ passes builder as function parameter |
| 6.2 | Uses oops.FromContext(ctx) in downstream functions to retrieve the pre-configured builder | ✓ | ✗ no context-based propagation |
| 6.3 | Middleware sets trace ID, request info, and user context on the builder | ✓ | ✗ no middleware pattern |
| 6.4 | Shows the middleware pattern with http.Handler wrapping | ✓ | ✗ no http.Handler wrapper |
| 6.5 | Downstream handlers/services can add more context (e.g., .Tags()) on top of the base builder | ✓ | ✗ each handler builds from scratch |
| 7. reusable-builder-pattern — Tests the pattern of creating a reusable builder at the top of a function and reusing it for multiple error paths | 5/5 | 2/5 | |
| 7.1 | Creates a single base builder variable at the top of the function with shared context (user, tenant, domain) | ✓ | ✗ duplicates builder chain at each error site |
| 7.2 | Each error return path extends the base builder with error-specific attributes using .With() or .Code() | ✓ | ✗ full chain repeated |
| 7.3 | The base builder is NOT terminated (no .Errorf/.Wrap call) — it's reused | ✓ | ✗ no shared builder |
| 7.4 | Uses .In() on the shared builder for the domain/feature | ✓ | ✓ |
| 7.5 | Uses .User() and/or .Tenant() on the shared builder | ✓ | ✓ |
| 8. accessing-oops-error-info — Tests knowledge of the OopsError type assertion to access structured fields | 6/6 | 3/6 | |
| 8.1 | Type-asserts the error to oops.OopsError | ✓ | ✓ |
| 8.2 | Uses .Code() method to get the error code | ✓ | ✓ |
| 8.3 | Uses .Domain() method to get the domain | ✓ | ✗ tries .In() on the error |
| 8.4 | Uses .Tags() method to get the tags | ✓ | ✗ not accessed |
| 8.5 | Uses .Context() method to get the key-value attributes map | ✓ | ✗ tries .Attributes() |
| 8.6 | Uses .Stacktrace() method to get the stack trace | ✓ | ✓ |
| 9. user-and-tenant-context — Tests the .User() and .Tenant() methods with their key-value attribute support | 5/5 | 1/5 | |
| 9.1 | Uses .User(id, key, value) method with the user ID and additional attributes like email | ✓ | ✗ uses .With("user_id", id) |
| 9.2 | Uses .Tenant(id, key, value) method with the tenant ID and additional attributes like plan | ✓ | ✗ uses .With("tenant_id", id) |
| 9.3 | Does NOT just use .With() for user/tenant info when .User()/.Tenant() are available | ✓ | ✗ only uses .With() |
| 9.4 | Includes a .Code() for the permission error | ✓ | ✗ no error code |
| 9.5 | Uses .Public() for a user-facing permission denied message | ✓ | ✓ |
| 10. oops-assertions — Tests knowledge of oops.Assert/oops.Assertf for invariant checks wrapped in Recover | 3/5 | 1/5 | |
| 10.1 | Uses oops.Assertf or oops.Assert to check the invariant (amount > 0) | ✓ | ✗ uses if/return error pattern |
| 10.2 | Wraps the assertion in an oops.Recover() call to convert the panic to a structured error | ✗ Recover missing in one code path | ✗ no Recover wrapping |
| 10.3 | Uses a named error return value so Recover can set it | ✗ omits named return | ✗ no named return |
| 10.4 | Notes that assertions should be rare in Go and used only for truly impossible/bug states | ✓ | ✗ no caveat about rare usage |
| 10.5 | Adds structured context (.In(), .Code(), etc.) to the Recover builder | ✓ | ✓ |
| 11. oops-configuration — Tests knowledge of oops global configuration options | 3/3 | 3/3 | |
| 11.1 | Uses oops.StackTraceMaxDepth to control stack trace depth | ✓ | ✓ |
| 11.2 | Uses oops.Local with time.LoadLocation for timezone configuration | ✓ | ✓ |
| 11.3 | Mentions oops.SourceFragmentsHidden as another available configuration option | ✓ | ✓ |
golang-stretchr-testify — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 47/47 (100%) | 25/47 (53%) | +47pp |
Full breakdown (47 assertions across 11 evals)
Model: Claude Opus 4.6 | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. assert-vs-require-precondition — Tests whether the model uses require for preconditions and assert for verifications | 5/5 | 3/5 | |
| 1.1 | Uses require (not assert) for the NoError check on parsing | ✓ | ✓ |
| 1.2 | Uses require (not assert) for the NotNil check on config | ✓ | ✗ uses assert.NotNil |
| 1.3 | Uses assert for the subsequent value checks (Port, Host, Debug) | ✓ | ✓ |
| 1.4 | Does NOT use require for all assertions indiscriminately | ✓ | ✓ |
| 1.5 | Argument order is (expected, actual) not (actual, expected) for Equal calls | ✓ | ✗ swaps to (actual, expected) |
| 2. assert-new-naming-convention — Tests the skill's specific naming convention: 'is' for assert.New(t) and 'must' for require.New(t) | 5/5 | 0/5 | |
| 2.1 | Uses assert.New(t) to create a reusable assertion object | ✓ | ✗ uses assert.Equal(t, ...) directly |
| 2.2 | Uses require.New(t) to create a reusable require object | ✓ | ✗ uses require.NoError(t, ...) directly |
| 2.3 | Names the assert.New(t) variable 'is' | ✓ | ✗ no New() usage |
| 2.4 | Names the require.New(t) variable 'must' | ✓ | ✗ no New() usage |
| 2.5 | Shows the 'is' and 'must' variables being used for different purposes (preconditions vs verifications) | ✓ | ✗ no differentiated naming |
| 3. error-chain-assertion — Tests knowledge that is.Equal(ErrNotFound, err) fails on wrapped errors and ErrorIs should be used | 4/4 | 2/4 | |
| 3.1 | Uses ErrorIs (not Equal) to check the error against ErrNotFound | ✓ | ✓ |
| 3.2 | Does NOT use assert.Equal or is.Equal to compare errors directly | ✓ | ✓ |
| 3.3 | Uses require for the initial error existence check if subsequent assertions depend on it | ✓ | ✗ uses assert for error check |
| 3.4 | Argument order for ErrorIs is (err, target) not (target, err) | ✓ | ✗ swaps arguments |
| 4. mock-assert-expectations — Tests whether AssertExpectations is called | 5/5 | 3/5 | |
| 4.1 | Mock embeds mock.Mock | ✓ | ✓ |
| 4.2 | Mock method uses m.Called() to forward arguments | ✓ | ✓ |
| 4.3 | Test calls m.AssertExpectations(t) to verify all expectations were met | ✓ | ✗ missing AssertExpectations |
| 4.4 | Uses .Once() or equivalent call modifier to enforce exactly one call | ✓ | ✗ no call modifier |
| 4.5 | Uses mock.Anything for arguments that don't need specific matching (e.g., context) | ✓ | ✓ |
| 5. mock-matched-by-predicate — Tests knowledge of mock.MatchedBy for custom argument matching | 5/5 | 3/5 | |
| 5.1 | Uses mock.MatchedBy with a predicate function for the LogEntry argument | ✓ | ✓ |
| 5.2 | The predicate checks Level == 'error' | ✓ | ✓ |
| 5.3 | The predicate checks that Message contains 'timeout' (using strings.Contains or similar) | ✓ | ✓ |
| 5.4 | Uses mock.Anything for the context argument | ✓ | ✗ passes context.Background() literally |
| 5.5 | Calls AssertExpectations at the end | ✓ | ✗ missing AssertExpectations |
| 6. mock-retry-different-returns — Tests knowledge of chaining .Once() calls for retry testing | 4/4 | 2/4 | |
| 6.1 | Sets up first On().Return() with an error and .Once() | ✓ | ✓ |
| 6.2 | Sets up second On().Return() with success data and .Once() | ✓ | ✓ |
| 6.3 | The two expectations are on the same method with the same arguments | ✓ | ✗ uses different argument patterns |
| 6.4 | Calls AssertExpectations to verify both calls happened | ✓ | ✗ missing AssertExpectations |
| 7. suite-lifecycle-and-launcher — Tests that suite requires a launcher function and understands lifecycle order | 5/5 | 3/5 | |
| 7.1 | Creates a suite struct embedding suite.Suite | ✓ | ✓ |
| 7.2 | Uses SetupTest (not SetupSuite) for per-test mock store initialization | ✓ | ✗ uses SetupSuite for mock init |
| 7.3 | Includes a launcher function: func TestXxxSuite(t *testing.T) with suite.Run() | ✓ | ✗ missing launcher function |
| 7.4 | Test methods are named TestXxx (starting with Test) on the suite receiver | ✓ | ✓ |
| 7.5 | Uses SetupSuite or TearDownSuite for the shared database connection (one-time setup) | ✓ | ✓ |
| 8. suite-require-syntax — Tests that suite methods use s.Require().NotNil() for require behavior | 3/3 | 1/3 | |
| 8.1 | Uses s.Require().NotNil() (not just s.NotNil()) for fail-fast behavior | ✓ | ✗ uses s.NotNil() thinking it's require |
| 8.2 | Explains that s.NotNil() and similar suite methods behave like assert (continue on failure) | ✓ | ✗ no explanation of default behavior |
| 8.3 | Shows that s.Require() returns a require-style assertion object | ✓ | ✓ |
| 9. pointer-comparison-trap — Tests awareness that is.Equal(ptr1, ptr2) compares addresses, not values | 3/3 | 2/3 | |
| 9.1 | Identifies that assert.Equal on pointers compares memory addresses, not struct values | ✓ | ✓ |
| 9.2 | Recommends dereferencing the pointers or using EqualExportedValues | ✓ | ✓ |
| 9.3 | Mentions EqualExportedValues as an alternative for comparing only exported fields | ✓ | ✗ only suggests dereferencing |
| 10. eventually-with-rich-assertions — Tests knowledge of EventuallyWithT for async polling with multiple rich assertions | 5/5 | 3/5 | |
| 10.1 | Uses EventuallyWithT (not just Eventually) for rich assertions | ✓ | ✗ uses Eventually with bool |
| 10.2 | The callback receives *assert.CollectT (or similar collect parameter) | ✓ | ✗ callback returns bool |
| 10.3 | Multiple assertions are made inside the callback (status check AND result count check) | ✓ | ✓ |
| 10.4 | Uses assert.NoError/assert.Equal with the CollectT parameter inside the callback, not with t | ✓ | ✓ |
| 10.5 | Specifies timeout (10s) and polling interval as separate parameters | ✓ | ✓ |
| 11. testifylint-recommendation — Tests whether the model recommends testifylint for catching common testify mistakes | 3/3 | 3/3 | |
| 11.1 | Recommends testifylint as a linter for testify-specific issues | ✓ | ✓ |
| 11.2 | Mentions that testifylint catches wrong argument order | ✓ | ✓ |
| 11.3 | Mentions that testifylint catches assert/require misuse | ✓ | ✓ |
golang-samber-mo — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 95/108 (88%) | 43/108 (40%) | +48pp |
Full breakdown (108 assertions)
Model: Claude Opus 4.6 | Runs: 23 evals × 2 configs = 46 subagents | Grading: Human-as-Judge + LLM-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. option-vs-pointer-for-nullable-db-field — Option[T] for nullable DB columns | 5/5 | 5/5 | |
| 1.1 | Uses mo.Option[string] for nullable fields | ✓ | ✓ |
| 1.2 | Mentions sql.Scanner and driver.Valuer | ✓ | ✓ |
| 1.3 | Mentions json.Marshaler/Unmarshaler | ✓ | ✓ |
| 1.4 | Shows row.Scan with Option type | ✓ | ✓ |
| 1.5 | Does NOT recommend sql.NullString | ✓ | ✓ |
| 2. result-vs-tuple-error-boundary — Result at API boundary vs internal | 5/5 | 2/5 | |
| 2.1 | Returns (Config, error) at public boundary | ✓ | ✗ returns Result publicly |
| 2.2 | Uses Result internally for chaining | ✓ | ✗ uses Result everywhere |
| 2.3 | Shows TupleToResult at boundary | ✓ | ✓ |
| 2.4 | Shows .Get() at end for conversion back | ✓ | ✗ returns Result directly |
| 2.5 | Explains Result for internal composition | ✓ | ✓ |
| 3. either-vs-result-two-valid-types — Either for non-error alternatives | 5/5 | 2/5 | |
| 3.1 | Uses Either[CachedUser, FreshUser] | ✓ | ✗ uses interface or Result |
| 3.2 | Does NOT use Result | ✓ | ✗ suggests common interface |
| 3.3 | Explains Either vs Result | ✓ | ✗ no distinction made |
| 3.4 | Shows Left/Right constructors | ✓ | ✓ |
| 3.5 | Shows Match or IsLeft/IsRight | ✓ | ✓ |
| 4. sub-package-for-type-changing-map — Sub-package for type-changing Map | 5/5 | 1/5 | |
| 4.1 | Uses option.Map from sub-package | ✓ | ✗ uses nonexistent mo.Map(opt, fn) |
| 4.2 | Does NOT use .Map for type change | ✓ | ✓ |
| 4.3 | Imports mo/option | ✓ | ✗ no sub-package import |
| 4.4 | Shows curried form | ✓ | ✗ uses mo.Map(opt, fn) pattern |
| 4.5 | Explains why sub-packages exist | ✓ | ✗ claims mo.Map is package-level |
| 5. do-notation-for-imperative-monadic — mo.Do for imperative monadic code | 5/5 | 0/5 | |
| 5.1 | Suggests mo.Do | ✓ | ✗ suggests unwrap-early pattern |
| 5.2 | Shows MustGet inside Do | ✓ | ✗ manual .Get() checks |
| 5.3 | Explains panic catching | ✓ | ✗ no mention of Do |
| 5.4 | Do returns Result[T] | ✓ | ✗ no mention of Do |
| 5.5 | Cleaner than FlatMap chains | ✓ | ✗ no Do alternative |
| 6. pipe-composition-multi-step — Pipe for multi-step type-changing pipelines | 5/5 | 0/5 | |
| 6.1 | Uses option.Pipe3 | ✓ | ✗ nests function calls |
| 6.2 | Each step uses option.Map/FlatMap | ✓ | ✗ manual intermediate vars |
| 6.3 | Pipeline reads top-to-bottom | ✓ | ✗ nested calls |
| 6.4 | Imports mo/option sub-package | ✓ | ✗ no sub-package awareness |
| 6.5 | FlatMap for validation step | ✓ | ✗ manual if/else |
| 7. future-vs-task-eager-vs-lazy — Future (eager) vs Task (lazy) | 5/5 | 4/5 | |
| 7.1 | Recommends Task | ✓ | ✓ |
| 7.2 | Future starts immediately | ✓ | ✓ |
| 7.3 | Task runs on .Run() | ✓ | ✓ |
| 7.4 | Run returns *Future[T] | ✓ | ✓ |
| 7.5 | Deferred execution pattern | ✓ | ✗ wrong constructor signature |
| 8. tuple-to-result-wrapping-stdlib — TupleToResult for stdlib wrapping | 5/5 | 3/5 | |
| 8.1 | Uses TupleToResult(os.ReadFile(path)) | ✓ | ✓ |
| 8.2 | Uses TupleToResult or Try for Atoi | ✓ | ✓ |
| 8.3 | Does NOT manually check err | ✓ | ✗ manually constructs Ok/Err |
| 8.4 | Chains results | ✓ | ✓ |
| 8.5 | Explains TupleToResult | ✓ | ✗ no explanation |
| 9. when-not-to-use-monads — Advises against monads for simple cases | 4/4 | 3/4 | |
| 9.1 | Advises against Result here | ✓ | ✓ |
| 9.2 | Recommends if err != nil | ✓ | ✓ |
| 9.3 | Result shines with chains | ✓ | ✓ |
| 9.4 | Does NOT over-apply Result | ✓ | ✗ still wraps in Result |
| 10. option-json-serialization — Option JSON marshaling behavior | 5/5 | 3/5 | |
| 10.1 | Uses Option[string] | ✓ | ✓ |
| 10.2 | Some->value, None->null | ✓ | ✓ |
| 10.3 | Mentions omitzero (Go 1.24+) | ✓ | ✗ no omitzero awareness |
| 10.4 | Shows struct with json tag | ✓ | ✓ |
| 10.5 | No custom MarshalJSON needed | ✓ | ✗ suggests custom marshaler |
| 11. emptyable-to-option-zero-value — EmptyableToOption for zero detection | 4/4 | 0/4 | |
| 11.1 | Uses EmptyableToOption | ✓ | ✗ manual if/else |
| 11.2 | None for zero, Some for non-zero | ✓ | ✗ manual construction |
| 11.3 | No manual check | ✓ | ✗ uses if/else |
| 11.4 | Works for any comparable type | ✓ | ✗ only handles string |
| 12. pointer-to-option-nil-handling — PointerToOption for nil conversion | 3/3 | 1/3 | |
| 12.1 | Uses PointerToOption(ptr) | ✓ | ✗ manual nil check |
| 12.2 | nil->None, non-nil->Some | ✓ | ✓ |
| 12.3 | No manual nil check | ✓ | ✗ manual check |
| 13. result-map-vs-flatmap-choice — Map vs FlatMap on Result | 5/5 | 2/5 | |
| 13.1 | MapValue for infallible uppercase | ✓ | ✗ uses Map for both |
| 13.2 | FlatMap for fallible parse | ✓ | ✓ |
| 13.3 | No FlatMap for uppercase | ✓ | ✗ uses FlatMap for both |
| 13.4 | Shows chain | ✓ | ✓ |
| 13.5 | Explains distinction | ✓ | ✗ no distinction explained |
| 14. option-map-bool-semantics — Option.Map's (T, bool) return type | 4/4 | 0/4 | |
| 14.1 | Uses Map with (int, bool) return | ✓ | ✗ claims Map can't filter |
| 14.2 | false converts to None | ✓ | ✗ doesn't know bool return |
| 14.3 | Shows (T, bool) signature | ✓ | ✗ claims func(T) T |
| 14.4 | No FlatMap for filtering | ✓ | ✗ uses FlatMap |
| 15. foldable-interface-uniform-matching — Fold across Option/Result/Either | 0/5 | 0/5 | |
| 15.1 | Uses mo.Fold | ✗ not in skill | ✗ separate switch blocks |
| 15.2 | Shows Foldable interface | ✗ not in skill | ✗ no Foldable awareness |
| 15.3 | successFn/failureFn callbacks | ✗ not in skill | ✗ manual matching |
| 15.4 | Single Fold for all types | ✗ not in skill | ✗ separate logic |
| 15.5 | No separate handling per type | ✗ not in skill | ✗ separate blocks |
| 16. io-for-testable-side-effects — IO for deferring side effects | 4/4 | 0/4 | |
| 16.1 | Uses IO or IOEither | ✓ | ✗ uses interfaces/DI |
| 16.2 | IO is lazy, runs on Run() | ✓ | ✗ no IO awareness |
| 16.3 | IO1/IO2 parameterize | ✓ | ✗ no IO awareness |
| 16.4 | Composability of IO | ✓ | ✗ no IO awareness |
| 17. result-to-either-conversion — ToEither() conversion | 3/3 | 1/3 | |
| 17.1 | Uses result.ToEither() | ✓ | ✗ manual IsOk + Left/Right |
| 17.2 | Ok->Right, Err->Left | ✓ | ✓ |
| 17.3 | No manual IsOk check | ✓ | ✗ manual check |
| 18. either3-for-multi-type-union — Either3+ for n-ary unions | 5/5 | 2/5 | |
| 18.1 | Uses Either3 | ✓ | ✗ uses interface or type switch |
| 18.2 | Shows constructors | ✓ | ✗ no constructor awareness |
| 18.3 | Shows Match | ✓ | ✗ type switch |
| 18.4 | No interface{}/any | ✓ | ✓ |
| 18.5 | Mentions Either4/Either5 | ✓ | ✓ |
| 19. option-vs-zero-value-distinction — When Option adds value vs zero values | 4/4 | 3/4 | |
| 19.1 | Plain int for count | ✓ | ✓ |
| 19.2 | Option[string] for nickname | ✓ | ✓ |
| 19.3 | Explains absence vs zero | ✓ | ✓ |
| 19.4 | No Option[int] for count | ✓ | ✗ uses Option for both |
| 20. map-lookup-to-option — TupleToOption for map lookups | 4/4 | 1/4 | |
| 20.1 | Uses TupleToOption(m[key]) | ✓ | ✗ manual ok check |
| 20.2 | Chains Map/FlatMap | ✓ | ✓ |
| 20.3 | No manual ok check | ✓ | ✗ manual check |
| 20.4 | Complete pattern | ✓ | ✗ manual construction |
| 21. result-try-catch-panics — Try for wrapping panicky functions | 2/4 | 2/4 | |
| 21.1 | Uses mo.Try | ✓ | ✓ |
| 21.2 | Try catches panics | ✗ skill says Do catches panics, not Try | ✗ claims Try only wraps (T, error) |
| 21.3 | Shows signature | ✓ | ✓ |
| 21.4 | No separate defer/recover | ✗ skill suggests Do for panics | ✗ adds defer/recover |
| 22. state-monad-for-accumulation — State monad for threading state | 4/4 | 0/4 | |
| 22.1 | Uses State/NewState | ✓ | ✗ uses mutable struct |
| 22.2 | State params = state + result | ✓ | ✗ mutable position field |
| 22.3 | Run(initialState) | ✓ | ✗ no State awareness |
| 22.4 | State threaded, not mutated | ✓ | ✗ mutates struct |
| 23. result-map-signature — Result.Map's (T, error) return type | 4/4 | 2/4 | |
| 23.1 | Map callback returns (int, error) | ✓ | ✓ |
| 23.2 | Does NOT use func(int) int | ✓ | ✓ |
| 23.3 | MapValue alternative | ✓ | ✗ no MapValue mention |
| 23.4 | Map vs MapValue signatures | ✓ | ✗ no distinction |
golang-samber-hot — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 61/65 (94%) | 26/65 (40%) | +54pp |
Full breakdown (65 assertions)
Model: Claude Opus 4.6 | Runs: 15 evals × 2 configs = 30 subagents | Grading: LLM-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. algorithm-selection-default — mixed workload, should pick W-TinyLFU over LRU | 6/6 | 0/6 | |
| 1.1 | Uses hot.WTinyLFU as the eviction algorithm (not hot.LRU) | ✓ | ✗ used hot.NewLRU |
| 1.2 | Uses hot.NewHotCache constructor with generic type parameters | ✓ | ✗ fabricated API |
| 1.3 | Chains .WithTTL(10 * time.Minute) or equivalent | ✓ | ✗ positional arg |
| 1.4 | Chains .WithJanitor() in the builder | ✓ | ✗ missing |
| 1.5 | Calls defer cache.StopJanitor() after Build() | ✓ | ✗ missing |
| 1.6 | Calls .Build() to finalize the cache | ✓ | ✗ no builder pattern |
| 2. algorithm-selection-frequency — stable power-law DNS, should pick LFU | 4/4 | 4/4 | |
| 2.1 | Recommends hot.LFU for stable-frequency workload | ✓ | ✓ |
| 2.2 | Explains why LFU fits: keeps frequently accessed items | ✓ | ✓ |
| 2.3 | Mentions LFU weakness doesn't apply because rankings are stable | ✓ | ✓ |
| 2.4 | Does NOT default to LRU | ✓ | ✓ |
| 3. janitor-required — simple cache with TTL must include WithJanitor() | 4/4 | 1/4 | |
| 3.1 | Includes .WithJanitor() in the builder chain | ✓ | ✗ fabricated config struct |
| 3.2 | Includes defer cache.StopJanitor() for cleanup | ✓ | ✗ missing |
| 3.3 | Uses .WithTTL() for expiration | ✓ | ✗ config struct field |
| 3.4 | Does NOT call cache.Janitor() separately after Build() | ✓ | ✓ |
| 4. missing-cache-panic-prevention — SetMissing requires config | 4/4 | 0/4 | |
| 4.1 | Configures WithMissingCache() or WithMissingSharedCache() in builder | ✓ | ✗ stores nil in main cache |
| 4.2 | Uses cache.SetMissing() or cache.SetMissingWithTTL() | ✓ | ✗ uses cache.Set(key, nil) |
| 4.3 | Missing cache configured BEFORE Build() | ✓ | ✗ no config |
| 4.4 | Explains SetMissing without config panics | ✓ | ✗ unaware of SetMissing |
| 5. loader-pattern-singleflight — built-in dedup, no manual singleflight | 5/5 | 0/5 | |
| 5.1 | Uses .WithLoaders() in the builder chain | ✓ | ✗ manual singleflight |
| 5.2 | Explains built-in singleflight deduplication | ✓ | ✗ suggests external package |
| 5.3 | Does NOT implement manual singleflight on top | ✓ | ✗ uses x/sync/singleflight |
| 5.4 | Loader signature matches func(keys []K) (map[K]V, error) | ✓ | ✗ manual fetch pattern |
| 5.5 | Recommends WithJitter() for additional thundering-herd mitigation | ✓ | ✗ manual rand jitter |
| 6. stale-while-revalidate — two-threshold revalidation pattern | 5/5 | 2/5 | |
| 6.1 | Uses .WithTTL(5 * time.Minute) for fresh duration | ✓ | ✓ |
| 6.2 | Uses .WithRevalidation(2 * time.Minute, loader) | ✓ | ✗ invented API |
| 6.3 | Explains fresh -> stale -> expired model | ✓ | ✓ |
| 6.4 | Uses WithRevalidationErrorPolicy() | ✓ | ✗ not mentioned |
| 6.5 | Does NOT use a single 7min TTL | ✓ | ✓ |
| 7. sharding-for-concurrency — WithSharding for lock contention | 4/4 | 2/4 | |
| 7.1 | Uses .WithSharding() in the builder | ✓ | ✗ custom sharded wrapper |
| 7.2 | Shard count is a power of 2 | ✓ | ✓ |
| 7.3 | Provides or discusses a hash function of type func(K) uint64 | ✓ | ✓ |
| 7.4 | Does NOT recommend WithoutLocking() | ✓ | ✓ |
| 8. copy-on-read-mutable-values — WithCopyOnRead for pointer safety | 4/4 | 2/4 | |
| 8.1 | Identifies callers mutating shared cached pointers | ✓ | ✓ |
| 8.2 | Uses .WithCopyOnRead(fn) | ✓ | ✗ manual copy wrapper |
| 8.3 | Copy function creates shallow or deep copy | ✓ | ✓ |
| 8.4 | Does NOT suggest only external mutexes as primary solution | ✓ | ✓ |
| 9. loader-chain-semantics — later overwrites earlier, error stops chain | 5/5 | 1/5 | |
| 9.1 | Later loaders overwrite earlier values (PostgreSQL wins) | ✓ | ✗ says Redis wins |
| 9.2 | Any loader error stops the entire chain | ✓ | ✗ unclear |
| 9.3 | Partial results from earlier loaders discarded on error | ✓ | ✗ not mentioned |
| 9.4 | Shows WithLoaders(redisLoader, dbLoader) | ✓ | ✗ invented API |
| 9.5 | Later loaders only receive keys NOT found by previous | ✓ | ✓ |
| 10. algorithm-scan-resistance — LRU scan pollution, switch to scan-resistant | 4/4 | 4/4 | |
| 10.1 | Identifies LRU scan pollution problem | ✓ | ✓ |
| 10.2 | Recommends scan-resistant algorithm | ✓ | ✓ |
| 10.3 | Explains why recommended algorithm resists scan pollution | ✓ | ✓ |
| 10.4 | Does NOT suggest only increasing cache size | ✓ | ✓ |
| 11. withoutlocking-janitor-conflict — mutually exclusive, panics | 4/4 | 3/4 | |
| 11.1 | Warns WithoutLocking() + WithJanitor() are mutually exclusive | ✓ | ✓ |
| 11.2 | Recommends removing one of the two | ✓ | ✓ |
| 11.3 | Suggests manual cleanup or dropping WithoutLocking() | ✓ | ✓ |
| 11.4 | Does NOT produce code chaining both | ✓ | ✗ initial code chains both |
| 12. warmup-before-traffic — WithWarmUp in builder | 4/4 | 1/4 | |
| 12.1 | Uses .WithWarmUp(fn) or .WithWarmUpWithTimeout(timeout, fn) | ✓ | ✗ wrong API guess |
| 12.2 | Warm-up function returns (map[K]V, []K, error) | ✓ | ✗ wrong signature |
| 12.3 | Warm-up happens before Build() returns | ✓ | ✓ |
| 12.4 | Does NOT manually loop calling Set() | ✓ | ✗ loops calling Get() |
| 13. prometheus-monitoring-setup — built-in WithPrometheusMetrics | 4/4 | 1/4 | |
| 13.1 | Uses .WithPrometheusMetrics(cacheName) | ✓ | ✗ custom counters |
| 13.2 | Registers with prometheus.MustRegister(cache) | ✓ | ✗ registers custom metrics |
| 13.3 | Mentions hit rate as key metric (target >80%) | ✓ | ✓ |
| 13.4 | Does NOT create custom Prometheus counters | ✓ | ✗ full custom instrumentation |
| 14. get-error-handling — three return values (V, bool, error) | 4/4 | 2/4 | |
| 14.1 | Get() returns (V, bool, error) | ✓ | ✗ wrong semantics |
| 14.2 | Checks error first before bool | ✓ | ✓ |
| 14.3 | Handles all three cases: err, !found, found | ✓ | ✗ no not-found handling |
| 14.4 | Does NOT ignore error return value | ✓ | ✓ |
| 15. peek-vs-get-distinction — Peek() for side-effect-free inspection | 4/4 | 3/4 | |
| 15.1 | Recommends Peek() or PeekMany() | ✓ | ✓ |
| 15.2 | Explains Peek() does not trigger loaders | ✓ | ✓ |
| 15.3 | Explains Peek() ignores expiration (returns expired entries) | ✓ | ✗ no expired-entry detail |
| 15.4 | Does NOT recommend Get() for inspection | ✓ | ✓ |
golang-samber-lo — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 83/86 (97%) | 49/86 (57%) | +40pp |
Full breakdown (86 assertions)
Model: Claude Opus 4.6 | Runs: 17 evals × 2 configs = 34 subagents | Grading: LLM-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. lop-for-io-trap — Should recommend errgroup over lop for HTTP fan-out | 5/5 | 1/5 | |
| 1.1 | Recommends errgroup or similar I/O concurrency pattern | ✓ | ✗ used lop.Map as primary |
| 1.2 | Explains lop is for CPU-bound parallelism, not I/O | ✓ | ✗ |
| 1.3 | Mentions lack of context cancellation as lop limitation | ✓ | ✗ |
| 1.4 | Does NOT use lop.Map as primary HTTP fan-out solution | ✓ | ✗ |
| 1.5 | Shows working Go code for concurrent fetch | ✓ | ✓ |
| 2. premature-lom-optimization — Should advise profiling before switching to lom | 6/6 | 3/6 | |
| 2.1 | Recommends profiling (pprof, alloc_objects) before switching | ✓ | ✓ |
| 2.2 | Warns lom mutates input slice (breaks immutability) | ✓ | ✗ claims lom doesn't exist |
| 2.3 | Does NOT blindly refactor all lo calls to lom | ✓ | ✓ |
| 2.4 | Explains performance issue may not be lo allocations | ✓ | ✓ |
| 2.5 | Mentions lom only for hot paths confirmed by profiling | ✓ | ✗ claims lom doesn't exist |
| 2.6 | Warns about concurrency safety of mutable ops | ✓ | ✗ |
| 3. stdlib-vs-lo-preference — Should prefer stdlib for Contains/Sort/Keys | 6/6 | 5/6 | |
| 3.1 | Recommends slices.Contains from stdlib | ✓ | ✓ |
| 3.2 | Recommends slices.Sort from stdlib | ✓ | ✓ |
| 3.3 | Recommends maps.Keys from stdlib | ✓ | ✓ |
| 3.4 | Mentions Go 1.21+ availability | ✓ | ✓ |
| 3.5 | Explains rationale: prefer stdlib when available | ✓ | ✗ says "either works fine" |
| 3.6 | Acknowledges lo useful for ops stdlib lacks | ✓ | ✓ |
| 4. loi-go-version-constraint — Should warn lo/it requires Go 1.23+ | 5/5 | 2/5 | |
| 4.1 | Warns loi requires Go 1.23+, not available in Go 1.20 | ✓ | ✗ claims lo/it doesn't exist |
| 4.2 | Suggests upgrading Go version to 1.23+ | ✓ | ✗ |
| 4.3 | Provides alternative approaches for Go 1.20 | ✓ | ✓ |
| 4.4 | Does NOT provide loi code that won't compile on Go 1.20 | ✓ | ✓ |
| 4.5 | Mentions range-over-func as Go 1.23 feature | ✓ | ✗ |
| 5. must-in-production-handler — Should warn against Must in HTTP handlers | 5/5 | 5/5 | |
| 5.1 | Warns Must panics on error, dangerous in handlers | ✓ | ✓ |
| 5.2 | Explains Must for tests/init only, not production | ✓ | ✓ |
| 5.3 | Provides refactored version with proper error handling | ✓ | ✓ |
| 5.4 | Shows returning appropriate HTTP error status codes | ✓ | ✓ |
| 5.5 | Does NOT approve Must-based approach for production | ✓ | ✓ |
| 6. import-aliases-knowledge — Should know lop/lom/loi aliases | 5/5 | 2/5 | |
| 6.1 | Lists lo/parallel with alias lop | ✓ | ✓ |
| 6.2 | Lists lo/mutable with alias lom | ✓ | ✗ unaware of mutable pkg |
| 6.3 | Lists lo/it with alias loi | ✓ | ✗ unaware of iterator pkg |
| 6.4 | Mentions lo/it requires Go 1.23+ | ✓ | ✗ |
| 6.5 | Lists core package with alias lo | ✓ | ✓ |
| 7. immutability-trap — lo.Filter does NOT modify input | 5/5 | 3/5 | |
| 7.1 | States lo.Filter returns new slice, doesn't modify input | ✓ | ✓ |
| 7.2 | Correctly says len(users) prints 2 | ✓ | ✓ |
| 7.3 | Explains lo is immutable by default as design principle | ✓ | ✗ no design principle framing |
| 7.4 | Points out return value of lo.Filter is discarded | ✓ | ✓ |
| 7.5 | Mentions lom.Filter as alternative for in-place mutation | ✓ | ✗ no mention of lom |
| 8. error-variant-awareness — Should recommend lo.MapErr | 4/5 | 1/5 | |
| 8.1 | Recommends lo.MapErr as error-aware variant | ✓ | ✗ |
| 8.2 | Explains MapErr stops on first error, returns (result, error) | ✓ | ✗ |
| 8.3 | Shows MapErr usage example | ✓ | ✗ |
| 8.4 | Mentions most lo functions have Err suffixes as pattern | ✗ | ✗ |
| 8.5 | Suggests lo.FilterMap for skipping errors | ✓ | ✓ |
| 9. simd-production-stability — Should warn simd is experimental | 4/5 | 2/5 | |
| 9.1 | Warns lo/exp/simd is experimental, API may break | ✓ | ✗ claims simd doesn't exist |
| 9.2 | States not covered by semver stability guarantees | ✓ | ✗ |
| 9.3 | Recommends benchmarking first | ✓ | ✓ |
| 9.4 | Suggests version pinning if used | ✗ | ✗ |
| 9.5 | Does NOT unconditionally recommend simd for production | ✓ | ✓ |
| 10. streaming-redirect-to-ro — Should redirect to samber/ro | 5/5 | 3/5 | |
| 10.1 | Recommends samber/ro for reactive/streaming pipelines | ✓ | ✗ suggests samber/hot instead |
| 10.2 | Explains lo is for finite/batch, not infinite streams | ✓ | ✓ |
| 10.3 | Mentions golang-samber-ro skill or samber/ro by name | ✓ | ✗ mentions samber/hot not ro |
| 10.4 | Does NOT attempt to use lo for infinite streaming | ✓ | ✓ |
| 10.5 | Explains conceptual difference: lo = batch, ro = reactive | ✓ | ✓ |
| 11. lop-small-dataset-trap — lop.Map wasteful for 10 items | 5/5 | 5/5 | |
| 11.1 | Recommends lo.Map instead of lop.Map for 10 items | ✓ | ✓ |
| 11.2 | Explains goroutine overhead exceeds benefit | ✓ | ✓ |
| 11.3 | Mentions threshold (~1000+ items) | ✓ | ✓ |
| 11.4 | Notes field access is trivial, no parallelism benefit | ✓ | ✓ |
| 11.5 | Does NOT recommend lop.Map for this case | ✓ | ✓ |
| 12. filtermap-vs-filter-then-map — Should use lo.FilterMap | 5/5 | 5/5 | |
| 12.1 | Recommends lo.FilterMap as single-pass alternative | ✓ | ✓ |
| 12.2 | Shows FilterMap with (R, bool) return signature | ✓ | ✓ |
| 12.3 | Explains avoids intermediate filtered slice | ✓ | ✓ |
| 12.4 | Provides working code parsing strings to ints | ✓ | ✓ |
| 12.5 | Does NOT chain Filter+Map as primary solution | ✓ | ✓ |
| 13. channel-dispatcher-strategies — Should list ChannelDispatcher strategies | 6/6 | 6/6 | |
| 13.1 | Mentions lo.ChannelDispatcher | ✓ | ✓ |
| 13.2 | Lists RoundRobin strategy | ✓ | ✓ |
| 13.3 | Lists Random strategy | ✓ | ✓ |
| 13.4 | Lists WeightedRandom/First/Least/Most | ✓ | ✓ |
| 13.5 | Explains when to choose different strategies | ✓ | ✓ |
| 13.6 | Shows code example using ChannelDispatcher | ✓ | ✓ |
| 14. v2-version-trap — Should state v2 does not exist | 4/4 | 0/4 | |
| 14.1 | States samber/lo v2 does not exist | ✓ | ✗ fabricated v2 migration guide |
| 14.2 | Mentions v1 follows semver, no breaking changes before v2 | ✓ | ✗ fabricated breaking changes |
| 14.3 | Provides correct install: go get github.com/samber/lo@v1 | ✓ | ✗ shows go get lo/v2 |
| 14.4 | Does NOT fabricate v2 migration steps | ✓ | ✗ full fabricated migration |
| 15. lazy-chain-intermediate-allocations — Should recommend loi | 5/5 | 1/5 | |
| 15.1 | Recommends lo/it (loi) for lazy evaluation | ✓ | ✗ claims loi doesn't exist |
| 15.2 | Explains loi processes on-demand without buffering | ✓ | ✗ |
| 15.3 | Mentions Go 1.23+ requirement | ✓ | ✗ |
| 15.4 | Shows lazy pipeline using loi functions | ✓ | ✗ |
| 15.5 | Contrasts eager vs lazy memory allocation | ✓ | ✓ |
| 16. lo-zero-dependencies — Should state zero external deps | 4/4 | 1/4 | |
| 16.1 | States zero external/runtime dependencies | ✓ | ✗ claims golang.org/x/exp dep |
| 16.2 | Mentions relies only on Go stdlib | ✓ | ✗ mentions x/exp |
| 16.3 | Addresses supply chain concern | ✓ | ✓ |
| 16.4 | Does NOT claim lo has external dependencies | ✓ | ✗ claims x/exp dependency |
| 17. ternary-evaluation-trap — lo.Ternary evaluates both branches | 4/4 | 4/4 | |
| 17.1 | Warns lo.Ternary evaluates both arguments | ✓ | ✓ |
| 17.2 | Recommends lo.TernaryF for lazy evaluation | ✓ | ✓ |
| 17.3 | Shows TernaryF usage with closures | ✓ | ✓ |
| 17.4 | Explains expensiveCompute runs even when condition is false | ✓ | ✓ |
golang-samber-slog — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 57/62 (92%) | 45/62 (73%) | +19pp |
Full breakdown (62 assertions)
Model: Claude Opus 4.6 | Runs: 12 evals × 2 configs = 24 subagents | Grading: LLM-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. pipeline-ordering — sampling + PII + Sentry composition | 5/5 | 3/5 | |
| 1.1 | Sampling is outermost/first handler in pipeline | ✓ | ✗ PII first, sampling second |
| 1.2 | Uses slog-sampling library for sampling | ✓ | ✓ |
| 1.3 | Uses slog-formatter for PII scrubbing | ✓ | ✓ |
| 1.4 | Uses slogmulti.Router or Fanout for routing | ✓ | ✓ |
| 1.5 | Explains why sampling should be first | ✓ | ✗ no ordering rationale |
| 2. fanout-vs-router — level-based routing to Sentry/Slack/stdout | 5/5 | 5/5 | |
| 2.1 | Uses slogmulti.Router() for level-based routing | ✓ | ✓ |
| 2.2 | stdout receives all log levels | ✓ | ✓ |
| 2.3 | Sentry receives only errors | ✓ | ✓ |
| 2.4 | Slack receives only warnings | ✓ | ✓ |
| 2.5 | Explains Router vs Fanout distinction | ✓ | ✓ |
| 3. missing-close-batch — Datadog handler graceful shutdown | 5/5 | 4/5 | |
| 3.1 | Calls handler.Close() or defer Close() | ✓ | ✓ |
| 3.2 | Uses Option{}.NewDatadogHandler() pattern | ✓ | ✓ |
| 3.3 | Close happens during graceful shutdown | ✓ | ✓ |
| 3.4 | Mentions batching / data loss risk | ✓ | ✓ |
| 3.5 | Import path uses versioned module (v2+) | ✓ | ✗ unversioned import |
| 4. sampling-strategy-selection — reduce 90% volume, keep errors | 5/5 | 5/5 | |
| 4.1 | Uses Threshold or Absolute (not Uniform) | ✓ | ✓ |
| 4.2 | Sampling only affects Debug/Info levels | ✓ | ✓ |
| 4.3 | Warn/Error pass through unsampled | ✓ | ✓ |
| 4.4 | Uses slog-sampling library | ✓ | ✓ |
| 4.5 | Explains why Uniform is wrong | ✓ | ✓ |
| 5. router-missing-default — DEBUG logs disappearing | 5/5 | 5/5 | |
| 5.1 | Identifies missing catch-all handler | ✓ | ✓ |
| 5.2 | Explains unmatched records silently dropped | ✓ | ✓ |
| 5.3 | Suggests adding default handler or LevelIs(Debug) | ✓ | ✓ |
| 5.4 | Uses slogmulti.Router() API correctly | ✓ | ✓ |
| 5.5 | Does NOT suggest changing log level as fix | ✓ | ✓ |
| 6. http-middleware-filters — slog-gin with path filters | 5/5 | 5/5 | |
| 6.1 | Uses sloggin.NewWithConfig() | ✓ | ✓ |
| 6.2 | Configures IgnorePath for /health and /metrics | ✓ | ✓ |
| 6.3 | Sets WithRequestBody: true | ✓ | ✓ |
| 6.4 | Config includes level fields | ✓ | ✓ |
| 6.5 | Uses IgnorePath (not custom filter) | ✓ | ✓ |
| 7. pipe-middleware-chain — trace injection + email scrubbing | 5/5 | 4/5 | |
| 7.1 | Uses slogmulti.Pipe() for chaining | ✓ | ✓ |
| 7.2 | Creates middleware for trace_id injection | ✓ | ✓ |
| 7.3 | Creates middleware for email scrubbing | ✓ | ✓ |
| 7.4 | Pipe wraps the final handler | ✓ | ✓ |
| 7.5 | Does NOT implement custom slog.Handler struct | ✓ | ✗ full Handler struct |
| 8. failover-handler — Loki fallback to file | 5/5 | 4/5 | |
| 8.1 | Uses slogmulti.Failover() | ✓ | ✓ |
| 8.2 | Loki is primary handler | ✓ | ✓ |
| 8.3 | File handler is fallback | ✓ | ✓ |
| 8.4 | Does NOT implement custom retry logic | ✓ | ✓ |
| 8.5 | Uses slog-loki library | ✓ | ✗ placeholder function |
| 9. pool-vs-fanout-latency — fix sequential Fanout latency | 2/5 | 2/5 | |
| 9.1 | Identifies sequential Fanout as root cause | ✓ | ✓ |
| 9.2 | Recommends slogmulti.Pool() | ✗ custom goroutines | ✗ misuses Pool |
| 9.3 | Explains latency reduction (sum to max) | ✓ | ✗ claims 0ms latency |
| 9.4 | Does NOT suggest raw goroutines as primary | ✗ | ✓ |
| 9.5 | Shows code change from Fanout to Pool | ✗ | ✗ |
| 10. formatter-pii-scrubbing — cross-cutting PII protection | 5/5 | 3/5 | |
| 10.1 | Uses slog-formatter library | ✓ | ✗ custom middleware |
| 10.2 | Uses PIIFormatter and/or IPAddressFormatter | ✓ | ✗ custom regex |
| 10.3 | Applied as Pipe middleware wrapping all handlers | ✓ | ✓ |
| 10.4 | Applied once before Router/Fanout | ✓ | ✓ |
| 10.5 | Does NOT implement custom per-handler PII logic | ✓ | ✓ |
| 11. backend-option-pattern — Sentry + Loki setup | 5/5 | 4/5 | |
| 11.1 | Uses slogsentry.Option{}.NewSentryHandler() | ✓ | ✓ |
| 11.2 | Uses slogloki.Option{}.NewLokiHandler() | ✓ | ✓ |
| 11.3 | Uses Router for level-based routing | ✓ | ✗ uses Fanout |
| 11.4 | Versioned import paths | ✓ | ✓ |
| 11.5 | Calls Close() on Loki handler | ✓ | ✓ |
| 12. attrfromcontext-without-middleware — empty request_id diagnosis | 7/7 | 6/7 | |
| 12.1 | Identifies no middleware populating context | ✓ | ✓ |
| 12.2 | Recommends adding HTTP middleware | ✓ | ✓ |
| 12.3 | Explains middleware injects attributes into context | ✓ | ✓ |
| 12.4 | Does NOT suggest changing context key as primary fix | ✓ | ✓ |
| 12.5 | Shows connection between middleware and AttrFromContext | ✓ | ✓ |
| 12.6 | Mentions WithRequestID config option | ✓ | ✗ |
| 12.7 | Does NOT blame slog-multi/sentry config | ✓ | ✓ |
golang-samber-ro — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 113/113 (100%) | 57/113 (50%) | +50pp |
Full breakdown (113 assertions)
Model: Claude Opus 4.6 | Runs: 25 evals × 2 configs = 50 subagents | Grading: LLM-as-judge (self-grading subagents)
| # | Assertion | With | Without |
|---|---|---|---|
| 1. typed-pipe-vs-untyped — chain 3 operators with type safety | 4/4 | 3/4 | |
| 1.1 | Uses ro.Pipe3 instead of untyped ro.Pipe for compile-time type safety | ✓ | ✗ uses untyped Pipe |
| 1.2 | Uses ro.Filter with func(int) bool predicate | ✓ | ✓ |
| 1.3 | Uses ro.Map with func(int) string transform | ✓ | ✓ |
| 1.4 | Uses ro.Takestring with correct generic type | ✓ | ✗ omits generic param |
| 2. lo-vs-ro-boundary — finite slice should use lo, not ro | 4/4 | 0/4 | |
| 2.1 | Recommends samber/lo instead of samber/ro | ✓ | ✗ uses ro Observable |
| 2.2 | Explains WHY lo is better: synchronous, no stream overhead | ✓ | ✗ |
| 2.3 | Does NOT create an Observable pipeline for a slice transform | ✓ | ✗ creates Observable |
| 2.4 | Uses lo.Filter and lo.Map | ✓ | ✗ uses ro operators |
| 3. observer-error-handling — full observer, not just OnNext | 4/4 | 1/4 | |
| 3.1 | Uses ro.NewObserver with all 3 callbacks | ✓ | ✗ uses ro.OnNext |
| 3.2 | Includes an error handler | ✓ | ✓ |
| 3.3 | Includes a completion handler | ✓ | ✗ infinite stream, skipped |
| 3.4 | Mentions risk of OnNext alone (silent error dropping) | ✓ | ✗ |
| 4. infinite-stream-shutdown — graceful SIGTERM shutdown | 5/5 | 0/5 | |
| 4.1 | Uses TakeUntil or context+ThrowOnContextCancel | ✓ | ✗ generic context approach |
| 4.2 | Mentions signal plugin (plugins/signal) | ✓ | ✗ uses manual os/signal |
| 4.3 | Calls .Wait() to block until shutdown | ✓ | ✗ doesn't know .Wait() |
| 4.4 | Does NOT suggest manual channel/goroutine killing | ✓ | ✗ uses channel-based approach |
| 4.5 | Mentions Unsubscribe() as alternative | ✓ | ✗ not in ro API context |
| 5. subject-config-store — BehaviorSubject for config | 5/5 | 3/5 | |
| 5.1 | Recommends BehaviorSubject | ✓ | ✓ |
| 5.2 | Explains replay of last value to new subscribers | ✓ | ✓ |
| 5.3 | Shows NewBehaviorSubjectConfig constructor | ✓ | ✗ wrong constructor syntax |
| 5.4 | Shows .Send() and .Subscribe() | ✓ | ✗ uses .Next() from RxJS |
| 5.5 | Does NOT recommend PublishSubject | ✓ | ✓ |
| 6. subject-chat-room — ReplaySubject for chat history | 4/4 | 3/4 | |
| 6.1 | Recommends ReplaySubject with buffer 50 | ✓ | ✓ |
| 6.2 | Shows NewReplaySubjectMessage constructor | ✓ | ✗ wrong constructor syntax |
| 6.3 | Explains N-value replay for late subscribers | ✓ | ✓ |
| 6.4 | Does NOT recommend BehaviorSubject | ✓ | ✓ |
| 7. share-websocket — hot observable for shared WebSocket | 5/5 | 5/5 | |
| 7.1 | Uses Share() or ShareReplay() | ✓ | ✓ |
| 7.2 | Creates only ONE WebSocket connection | ✓ | ✓ |
| 7.3 | Does NOT create 3 separate connections | ✓ | ✓ |
| 7.4 | Shows 3 separate .Subscribe() calls | ✓ | ✓ |
| 7.5 | Explains cold vs hot distinction | ✓ | ✓ |
| 8. combinelatest-vs-zip — latest values from different-rate sources | 6/6 | 5/6 | |
| 8.1 | Recommends CombineLatest2 | ✓ | ✓ |
| 8.2 | Explains re-emit on either source update | ✓ | ✓ |
| 8.3 | Explains why NOT Zip | ✓ | ✓ |
| 8.4 | Explains why NOT Merge | ✓ | ✓ |
| 8.5 | Shows lo.Tuple2 return type | ✓ | ✗ generic tuple |
| 8.6 | Uses Map to compute product | ✓ | ✗ generic combiner function |
| 9. retry-with-backoff — RetryWithConfig + cached fallback | 6/6 | 1/6 | |
| 9.1 | Uses RetryWithConfig (not infinite Retry) | ✓ | ✗ uses Retry with count param |
| 9.2 | Sets Max: 3 | ✓ | ✗ wrong API |
| 9.3 | Sets Delay: 500ms and BackoffMultiplier: 2.0 | ✓ | ✗ wrong field names |
| 9.4 | Sets MaxDelay: 10s | ✓ | ✗ field not included |
| 9.5 | Chains fallback after RetryWithConfig | ✓ | ✗ uses RxJS-style naming |
| 9.6 | Correct operator order: retry before fallback | ✓ | ✓ |
| 10. scan-vs-reduce — running average for dashboard | 5/5 | 5/5 | |
| 10.1 | Recommends Scan | ✓ | ✓ |
| 10.2 | Explains intermediate emission | ✓ | ✓ |
| 10.3 | Explains Reduce only emits final | ✓ | ✓ |
| 10.4 | Shows accumulator with count/sum | ✓ | ✓ |
| 10.5 | Notes infinite stream + Reduce = never emits | ✓ | ✓ |
| 11. collect-synchronous — Collect for finite observable to slice | 4/4 | 0/4 | |
| 11.1 | Uses ro.Collect returning ([]int, error) | ✓ | ✗ manual subscribe+accumulate |
| 11.2 | Checks error return value | ✓ | ✗ |
| 11.3 | Does NOT manually subscribe and accumulate | ✓ | ✗ |
| 11.4 | Mentions Collect blocks until complete | ✓ | ✗ |
| 12. context-propagation — wire timeout into pipeline | 4/4 | 1/4 | |
| 12.1 | Uses ContextWithTimeout or ContextReset | ✓ | ✗ passes ctx to Subscribe |
| 12.2 | Uses ThrowOnContextCancel | ✓ | ✗ |
| 12.3 | Chains context operators in pipeline | ✓ | ✗ |
| 12.4 | Handles cancellation error in onError | ✓ | ✓ |
| 13. plugin-fsnotify — file watcher with debounce | 5/5 | 3/5 | |
| 13.1 | Knows fsnotify plugin (plugins/fsnotify) | ✓ | ✗ uses manual fsnotify |
| 13.2 | Uses plugin to create observable | ✓ | ✗ wraps channel manually |
| 13.3 | Uses ThrottleTime for debounce | ✓ | ✓ |
| 13.4 | Filters for Write events | ✓ | ✓ |
| 13.5 | Shows Map to reload config | ✓ | ✓ |
| 14. plugin-cron — daily midnight schedule | 4/4 | 0/4 | |
| 14.1 | Knows cron plugin (plugins/cron) | ✓ | ✗ uses Interval |
| 14.2 | Uses cron expression 0 0 * * * |
✓ | ✗ |
| 14.3 | Shows correct import path | ✓ | ✗ |
| 14.4 | Chains with Map/FlatMap for report generation | ✓ | ✗ |
| 15. maperr-fallible-transform — MapErr for JSON parsing | 5/5 | 1/5 | |
| 15.1 | Uses MapErr (not Map) | ✓ | ✗ uses Map |
| 15.2 | Shows func(string) (MyStruct, error) signature | ✓ | ✗ |
| 15.3 | Explains error propagation through pipeline | ✓ | ✗ |
| 15.4 | Does NOT suggest panic/recover pattern | ✓ | ✓ |
| 15.5 | Mentions JSON encoding plugin as alternative | ✓ | ✗ |
| 16. buffer-batching — BufferWithTimeOrCount for DB writes | 4/4 | 3/4 | |
| 16.1 | Uses BufferWithTimeOrCount | ✓ | ✗ generic buffer approach |
| 16.2 | Sets count=100, duration=5s | ✓ | ✓ |
| 16.3 | Chains with Map/MapErr for batch processing | ✓ | ✓ |
| 16.4 | Explains why both conditions matter | ✓ | ✓ |
| 17. unicast-subject-queue — single-consumer job queue | 5/5 | 0/5 | |
| 17.1 | Recommends UnicastSubject | ✓ | ✗ uses channel or PublishSubject |
| 17.2 | Shows NewUnicastSubjectTask | ✓ | ✗ |
| 17.3 | Explains exactly one subscriber | ✓ | ✗ |
| 17.4 | Explains pre-subscribe buffering | ✓ | ✗ |
| 17.5 | Does NOT recommend PublishSubject | ✓ | ✗ |
| 18. share-vs-sharereplay — late subscribers missing data | 4/4 | 3/4 | |
| 18.1 | Recommends ShareReplay(1) | ✓ | ✓ |
| 18.2 | Explains Share doesn't buffer | ✓ | ✓ |
| 18.3 | Explains ShareReplay buffers last N | ✓ | ✓ |
| 18.4 | Shows correct syntax ro.ShareReplayT | ✓ | ✗ wrong generic syntax |
| 19. error-recovery-cascade — retry then fallback then default | 5/5 | 1/5 | |
| 19.1 | Correct order: RetryWithConfig, Catch, OnErrorReturn | ✓ | ✗ wrong operator names |
| 19.2 | Uses RetryWithConfig with Max: 2 | ✓ | ✗ generic retry approach |
| 19.3 | Uses Catch for secondary source | ✓ | ✗ uses RxJS-style naming |
| 19.4 | Uses OnErrorReturn for default | ✓ | ✗ wrong operator name |
| 19.5 | Correct Pipe order | ✓ | ✓ |
| 20. channel-bridge — FromChannel for legacy chan | 4/4 | 4/4 | |
| 20.1 | Uses ro.FromChannelEvent | ✓ | ✓ |
| 20.2 | Observable completes when channel closes | ✓ | ✓ |
| 20.3 | Chains standard operators | ✓ | ✓ |
| 20.4 | Does NOT create custom NewObservable | ✓ | ✓ |
| 21. tap-for-observability — logging without modifying stream | 5/5 | 3/5 | |
| 21.1 | Uses Tap/TapOnNext/TapOnError/TapOnComplete | ✓ | ✓ |
| 21.2 | Does NOT use Map with side effect | ✓ | ✓ |
| 21.3 | Shows TapOnError for error monitoring | ✓ | ✗ generic Tap only |
| 21.4 | Explains Tap/Do observe without altering | ✓ | ✓ |
| 21.5 | Mentions logging plugins (slog, zap, etc.) | ✓ | ✗ |
| 22. connectable-precise-control — wait for all subscribers | 5/5 | 4/5 | |
| 22.1 | Recommends Connectable | ✓ | ✓ |
| 22.2 | Shows ro.ConnectableT | ✓ | ✗ wrong constructor |
| 22.3 | Sets up subscribers before Connect | ✓ | ✓ |
| 22.4 | Explains Connect() starts execution | ✓ | ✓ |
| 22.5 | Explains why Share() is wrong | ✓ | ✓ |
| 23. flatmap-vs-map — flatten Observable[Observable[T]] | 4/4 | 3/4 | |
| 23.1 | Recommends FlatMap or MergeMap | ✓ | ✓ |
| 23.2 | Shows FlatMap code | ✓ | ✗ wrong generic syntax |
| 23.3 | Explains map+flatten behavior | ✓ | ✓ |
| 23.4 | Does NOT suggest Map+MergeAll as primary | ✓ | ✓ |
| 24. async-subject-final — last value on completion | 5/5 | 4/5 | |
| 24.1 | Recommends AsyncSubject | ✓ | ✓ |
| 24.2 | Shows NewAsyncSubject[T]() constructor |
✓ | ✗ wrong constructor |
| 24.3 | Explains last value only on completion | ✓ | ✓ |
| 24.4 | Does NOT recommend BehaviorSubject | ✓ | ✓ |
| 24.5 | Does NOT recommend ReplaySubject(1) | ✓ | ✓ |
| 25. version-stability — v0 pre-release warnings | 4/4 | 3/4 | |
| 25.1 | Mentions v0.x (pre-v1.0.0) | ✓ | ✓ |
| 25.2 | Warns about breaking changes | ✓ | ✓ |
| 25.3 | Mentions SemVer | ✓ | ✗ not confident about SemVer |
| 25.4 | Does NOT present as fully stable | ✓ | ✓ |
golang-uber-dig — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 20/20 (100%) | 18/20 (90%) | +10pp |
Full breakdown (20 assertions)
Model: Claude Opus 4.7 | Runs: 4 evals × 2 configs = 8 subagents | Grading: human (assertion-by-assertion)
Note: Preliminary subset of the 11-eval suite in
skills/golang-uber-dig/evals/evals.json(53 assertions total). The remaining 7 evals cover named values, dig.As, scopes for request-locals, DryRun graph validation, RecoverFromPanics, group flatten, and the fx-vs-dig recommendation. Re-run the full suite via/skill-creatorfor a complete report.
| # | Assertion | With | Without |
|---|---|---|---|
| 1. param-objects-many-deps — dig.In for 4+ deps | 5/5 | 5/5 | |
| 1.1 | Embeds dig.In in the parameter struct | ✓ | ✓ |
| 1.2 | Constructor takes the params struct as a single argument | ✓ | ✓ |
| 1.3 | Does NOT keep the long parameter list | ✓ | ✓ |
| 1.4 | Does NOT use a plain struct without dig.In | ✓ | ✓ |
| 1.5 | Mentions readability/maintainability benefit | ✓ | ✓ |
| 2. value-groups-for-handlers — group:"routes" | 5/5 | 5/5 | |
| 2.1 | Each handler returns dig.Out tagged group:"routes" | ✓ | ✓ |
| 2.2 | NewRouter consumes dig.In with []slice tagged group:"routes" | ✓ | ✓ |
| 2.3 | Handlers added with c.Provide; no manual slice in main() | ✓ | ✓ |
| 2.4 | Does NOT manually assemble the handler slice | ✓ | ✓ |
| 2.5 | Mentions group order is not guaranteed | ✓ | ✓ |
| 6. container-not-passed-around — composition-root only | 5/5 | 5/5 | |
| 6.1 | Advises against passing *dig.Container into business code | ✓ | ✓ |
| 6.2 | Container only at composition root (main / startup) | ✓ | ✓ |
| 6.3 | UserHandler takes typed dependencies as constructor parameters | ✓ | ✓ |
| 6.4 | Mentions service locator anti-pattern OR explains downside | ✓ | ✓ |
| 6.5 | Does NOT show example with container injected into handler | ✓ | ✓ |
| 8. decorate-not-rewrite — Decorate vs scope-shadow Provide | 5/5 | 3/5 | |
| 8.1 | Uses Decorate (c.Decorate / scope.Decorate) on *zap.Logger | ✓ | ✗ uses scope.Provide override |
| 8.2 | Decorator returns log.Named("worker") | ✓ | ✓ |
| 8.3 | Decorate at the worker scope/module, not globally | ✓ | ✓ |
| 8.4 | Does NOT modify the original NewLogger constructor | ✓ | ✓ |
| 8.5 | Mentions decorator scope semantics (applies to scope and descendants) | ✓ | ✗ describes scope-override semantics instead |
Analyst pass: evals 1, 2, and 6 score 5/5 in both configurations — they test knowledge the base model already has, hitting the "common knowledge" anti-pattern flagged in CLAUDE.md. Only eval 8 differentiates the skill: the without-skill agent reaches for scope.Provide shadowing (which would create a circular dependency at resolution time), while the skill steers to the correct Decorate API. Future iterations should redesign evals 1, 2, 6 to target subtler guidance and add cases the model gets wrong without the skill.
golang-uber-fx — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 21/21 (100%) | 20/21 (95%) | +5pp |
Full breakdown (21 assertions)
Model: Claude Opus 4.7 | Runs: 4 evals × 2 configs = 8 subagents | Grading: human (assertion-by-assertion)
Note: Preliminary subset of the 11-eval suite in
skills/golang-uber-fx/evals/evals.json(56 assertions total). The remaining 7 evals cover fx.Annotate vs fx.Out, fx.Module organization, fx.Supply, value groups, fxevent.ZapLogger, manual lifecycle for CLI embedding, and the fx-vs-dig recommendation. Re-run the full suite via/skill-creatorfor a complete report.
| # | Assertion | With | Without |
|---|---|---|---|
| 1. lifecycle-not-init — OnStart with goroutine | 6/6 | 6/6 | |
| 1.1 | Injects fx.Lifecycle into NewHTTPServer | ✓ | ✓ |
| 1.2 | lc.Append with fx.Hook (OnStart starts server, OnStop calls Shutdown) | ✓ | ✓ |
| 1.3 | OnStart launches srv.Serve in a goroutine | ✓ | ✓ |
| 1.4 | Does NOT call srv.Serve directly inside the constructor | ✓ | ✓ |
| 1.5 | Does NOT use init() to start the server | ✓ | ✓ |
| 1.6 | OnStop calls srv.Shutdown(ctx) for graceful shutdown | ✓ | ✓ |
| 5. fxtest-with-populate — fxtest.New + fx.Populate | 5/5 | 5/5 | |
| 5.1 | Uses fxtest.New(t, ...) instead of fx.New | ✓ | ✓ |
| 5.2 | Uses fx.Populate(&svc) to extract *UserService | ✓ | ✓ |
| 5.3 | Calls app.RequireStart() and app.RequireStop() | ✓ | ✓ |
| 5.4 | Provides a fake Database (interface, not real DB) | ✓ | ✓ |
| 5.5 | Does NOT use fx.Invoke as the primary extraction mechanism | ✓ | ✓ |
| 6. replace-for-fakes — fx.Replace with fx.Annotate(fx.As) | 5/5 | 4/5 | |
| 6.1 | Uses fx.Replace (or fx.Decorate) inside fxtest.New | ✓ | ✓ uses fx.Decorate |
| 6.2 | Composes ProductionModule alongside the override; module unchanged | ✓ | ✓ |
| 6.3 | Uses fx.Annotate with fx.As(new(Database)) for interface binding | ✓ | ✗ mentions but does not use |
| 6.4 | Does NOT modify or duplicate the production module | ✓ | ✓ |
| 6.5 | Mentions Replace/Decorate is appropriate for tests | ✓ | ✓ |
| 10. onstart-non-blocking — long-running work in goroutine | 5/5 | 5/5 | |
| 10.1 | OnStart launches the long-running method in a goroutine | ✓ | ✓ |
| 10.2 | OnStart returns nil quickly without waiting | ✓ | ✓ |
| 10.3 | OnStop signals stop and waits for drain (with timeout) | ✓ | ✓ |
| 10.4 | Does NOT call the long-running method synchronously in OnStart | ✓ | ✓ |
| 10.5 | Mentions a blocking OnStart hangs the boot | ✓ | ✓ |
Analyst pass: 3 of 4 evals score equally with and without the skill — the model's baseline knowledge of fx is very strong (lifecycle hooks, fxtest.New, OnStart/goroutine pattern). Only eval 6 differentiates: without the skill the agent picks fx.Decorate and skips the fx.As interface binding that the production graph requires. This is consistent with a well-known framework where the skill mainly adds value on subtle API choices. Future iterations should target less common patterns: fx.Annotate vs fx.Out trade-offs, fx.Module decorator scoping, fxevent customization, manual lifecycle for CLI embedding.
golang-google-wire — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 49/50 (98%) | 41/50 (82%) | +16pp |
Full breakdown (50 assertions)
Model: Claude Sonnet 4.6 | Runs: 10 evals × 2 configs = 20 subagents | Grading: LLM-as-Judge
| # | Assertion | With | Without |
|---|---|---|---|
1. build-constraint-on-injector — missing //go:build wireinject causes duplicate symbols |
5/5 | 3/5 | |
| 1.1 | Identifies the missing //go:build wireinject build tag as the root cause |
✓ | ✗ suggested //go:build ignore instead |
| 1.2 | Shows //go:build wireinject as the first line of the injector file |
✓ | ✗ showed //go:build ignore |
| 1.3 | Explains that the tag prevents the stub from being compiled into the binary | ✓ | ✓ |
| 1.4 | Does NOT suggest renaming the function or reorganizing packages as the fix | ✓ | ✓ |
| 1.5 | Does NOT suggest deleting wire_gen.go as the fix | ✓ | ✓ |
2. interface-binding-required — wire.Bind must be explicit; wire never auto-resolves |
5/5 | 4/5 | |
| 2.1 | Explains that wire never auto-resolves interface satisfaction — bindings must be explicit | ✓ | ✓ |
| 2.2 | Shows wire.Bind(new(UserStore), new(*PostgresUserRepo)) added to the provider set |
✓ | ✓ |
| 2.3 | Places wire.Bind inside the same wire.NewSet (or adds it to a set in wire.Build) |
✓ | ✓ |
| 2.4 | Explains WHY wire requires explicit bindings (predictability — avoids surprise rebinding) | ✓ | ✗ cited ambiguity concern, not predictability rationale |
| 2.5 | Does NOT suggest changing NewUserRepo to return UserStore directly as the primary fix |
✓ | ✓ |
| 3. duplicate-type-named-wrapper — named types disambiguate same underlying type | 5/5 | 4/5 | |
| 3.1 | Introduces distinct named types (e.g., type PrimaryDSN string and type ReplicaDSN string) |
✓ | ✓ |
| 3.2 | Updates NewPrimaryDSN to return PrimaryDSN and NewReplicaDSN to return ReplicaDSN |
✓ | ✓ |
| 3.3 | Updates NewPrimaryDB and NewReplicaDB signatures to accept the named types |
✓ | ✓ |
| 3.4 | Explains that wire enforces one provider per type, so distinct named types are the correct fix | ✓ | ✓ |
| 3.5 | Does NOT suggest using a single Config struct with both DSNs as the primary fix | ✓ | ✗ presented Config struct as co-primary recommendation |
4. cleanup-signature — (T, func(), error) cleanup propagated through injector |
5/5 | 4/5 | |
| 4.1 | Changes NewDB to return (*sql.DB, func(), error) where cleanup calls db.Close() |
✓ | ✓ |
| 4.2 | Changes the injector function to return (*App, func(), error) to propagate the cleanup chain |
✓ | ✓ |
| 4.3 | Shows main calling defer cleanup() after the nil-check |
✓ | ✗ called defer cleanup() directly without nil guard |
| 4.4 | Explains that wire chains cleanup functions in reverse construction order | ✓ | ✓ |
| 4.5 | Does NOT suggest passing db as an extra return value from InitApp alongside *App | ✓ | ✓ |
5. no-edit-wire-gen — never edit wire_gen.go; re-run wire ./... |
5/5 | 5/5 | |
| 5.1 | Explicitly says NOT to edit wire_gen.go (it is always overwritten) | ✓ | ✓ |
| 5.2 | Instructs running wire ./... to regenerate wire_gen.go |
✓ | ✓ |
| 5.3 | Explains that *zap.Logger must be provided in the graph (via a provider or wire.Value) |
✓ | ✓ |
| 5.4 | Shows how to add NewLogger (or wire.Value) to the appropriate wire.NewSet |
✓ | ✓ |
| 5.5 | Does NOT present editing wire_gen.go as an option | ✓ | ✓ |
6. provider-set-organization — per-package wire.NewSet, not one giant set in main |
5/5 | 5/5 | |
| 6.1 | Introduces per-package wire.NewSet variables (e.g., InfraSet, RepoSet, ServiceSet) |
✓ | ✓ |
| 6.2 | Each set lives in its own package's wire.go file (not all in main) | ✓ | ✓ |
| 6.3 | The injector wire.Build references the set variables rather than individual providers |
✓ | ✓ |
| 6.4 | wire.Bind declarations move into the relevant package's set (not into wire.Build directly) |
✓ | ✓ |
| 6.5 | Explains the benefit: per-package sets are independently composable | ✓ | ✓ |
7. injector-parameter-vs-value-provider — injector parameters vs wire.Value for pre-built values |
4/5 | 5/5 | |
| 7.1 | Shows the injector-parameter approach: func InitApp(cfg *Config) (*App, func(), error) |
✓ | ✓ |
| 7.2 | OR shows wire.Value(cfg) inside wire.Build — both are valid answers |
✗ | ✓ |
| 7.3 | Explains that injector parameters are treated as pre-built providers by wire | ✓ | ✓ |
| 7.4 | Does NOT use a global variable as the recommended solution | ✓ | ✓ |
| 7.5 | Does NOT suggest using init() to set the value |
✓ | ✓ |
8. fields-of-struct — wire.FieldsOf promotes struct fields without manual extractors |
5/5 | 1/5 | |
| 8.1 | Uses wire.FieldsOf(new(Config), "DatabaseDSN", "CacheAddress", "APIKey") or a subset |
✓ | ✗ actively discouraged wire.FieldsOf |
| 8.2 | Places wire.FieldsOf inside the provider set or wire.Build |
✓ | ✗ wire.FieldsOf not used |
| 8.3 | Updates NewDB, NewCache, NewExternalClient to accept the string fields as parameters |
✓ | ✓ |
| 8.4 | Explains that wire.FieldsOf promotes struct fields as individual graph nodes |
✓ | ✗ dismissed rather than explained |
| 8.5 | Does NOT suggest three separate extractor functions as the primary recommendation | ✓ | ✗ ProvideDatabaseDSN etc. were the only recommendation |
9. test-injector-pattern — test-only wire.NewSet with fake bindings |
5/5 | 5/5 | |
| 9.1 | Creates a test-only provider set with NewFakeMailer and wire.Bind(new(Mailer), new(*FakeMailer)) |
✓ | ✓ |
| 9.2 | Creates a test injector function in a _test.go file with //go:build wireinject |
✓ | ✓ |
| 9.3 | The test injector's wire.Build composes production sets with the test-only set |
✓ | ✓ |
| 9.4 | Does NOT suggest global variables, monkey-patching, or a runtime DI container for tests | ✓ | ✓ |
| 9.5 | Mentions that wire ./... (or go generate) must be run to produce the test-injector code |
✓ | ✓ |
| 10. wire-vs-fx-for-daemon — wire lacks lifecycle; recommend fx for long-running services | 5/5 | 5/5 | |
| 10.1 | Identifies that wire has no built-in lifecycle management (no OnStart/OnStop hooks) | ✓ | ✓ |
| 10.2 | Identifies that wire has no built-in signal handling (SIGINT/SIGTERM) | ✓ | ✓ |
| 10.3 | Recommends uber-go/fx (or flags it as the better fit) for a long-running HTTP daemon | ✓ | ✓ |
| 10.4 | Does NOT recommend wire as sufficient for a service requiring graceful shutdown | ✓ | ✓ |
| 10.5 | Mentions that with wire the developer must implement shutdown and signal handling manually | ✓ | ✓ |
Analyst pass: 5 of 10 evals score equally with and without the skill — evals 5, 6, 9, 10 reflect strong model priors for well-known wire patterns (never-edit wire_gen.go, per-package sets, test injectors, fx-vs-wire comparison). The clearest skill uplift is eval 8 (wire.FieldsOf): without the skill, the model actively discourages wire.FieldsOf and recommends manual extractor functions instead. Eval 1 also differentiates: without the skill, the model suggests //go:build ignore rather than the correct //go:build wireinject tag. Eval 7 produced a false negative for the with-skill run — the skill correctly notes wire.Value only accepts constant expressions and rejected it for a runtime-parsed config; this is technically correct behavior but the assertion accepts both forms as valid. Future iterations should redesign assertion 7.2 to credit the correct injector-parameter approach when wire.Value is rightly excluded.
golang-spf13-cobra — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 50/50 (100%) | 49/50 (98%) | +2pp |
Full breakdown (50 assertions)
Model: Claude Sonnet 4.6 | Runs: 12 evals × 2 configs = 24 self-graded evaluations | Grading: self-graded (⚠️ biased — model sees assertions in prompt; without-skill baseline inflated; human-as-judge evaluation recommended for accurate uplift)
| # | Assertion | With | Without |
|---|---|---|---|
| 1. rune-vs-run — RunE vs Run, defer bypass | 5/5 | 5/5 | |
| 1.1 | Identifies that Run cannot return an error | ✓ | ✓ |
| 1.2 | Explains only escape from Run is os.Exit or panic, bypassing defers | ✓ | ✓ |
| 1.3 | Recommends switching to RunE | ✓ | ✓ |
| 1.4 | Shows RunE signature: func(cmd *cobra.Command, args []string) error | ✓ | ✓ |
| 1.5 | Does NOT suggest goroutine or recover() as the fix | ✓ | ✓ |
| 2. args-validator — len(args) in RunE vs Args field | 4/4 | 3/4 | |
| 2.1 | Recommends cobra.ExactArgs(1) in the Args field | ✓ | ✓ |
| 2.2 | Shows the Args field on cobra.Command, not inside RunE | ✓ | ✓ |
| 2.3 | Explains cobra generates standard error messages automatically | ✓ | ✓ |
| 2.4 | Does NOT suggest keeping the len(args) check inside RunE | ✓ | ✗ |
| 3. outstdout — cmd.OutOrStdout vs os.Stdout | 4/4 | 4/4 | |
| 3.1 | Identifies os.Stdout (via fmt.Println) as root cause — can't redirect | ✓ | ✓ |
| 3.2 | Recommends cmd.OutOrStdout() as the io.Writer for output | ✓ | ✓ |
| 3.3 | Shows cmd.SetOut(buf) in test setup to redirect output | ✓ | ✓ |
| 3.4 | Does NOT suggest using os.Pipe() as the fix | ✓ | ✓ |
| 4. persistentprerun-chain — PersistentPreRunE replacement not chaining | 4/4 | 4/4 | |
| 4.1 | Explains cobra does NOT chain PersistentPreRunE — child replaces parent | ✓ | ✓ |
| 4.2 | Recommends explicitly calling parent's PersistentPreRunE from child | ✓ | ✓ |
| 4.3 | Shows calling parent hook: rootCmd.PersistentPreRunE(cmd, args) | ✓ | ✓ |
| 4.4 | Does NOT suggest using PreRunE instead as the fix | ✓ | ✓ |
| 5. silenceusage — SilenceUsage and SilenceErrors configuration | 4/4 | 4/4 | |
| 5.1 | Recommends SilenceUsage: true on the root command | ✓ | ✓ |
| 5.2 | Recommends SilenceErrors: true if app prints errors itself | ✓ | ✓ |
| 5.3 | Explains SilenceUsage prevents full usage wall on every error | ✓ | ✓ |
| 5.4 | Shows setting fields on root cobra.Command struct | ✓ | ✓ |
| 6. addgroup-ordering — AddGroup before AddCommand | 4/4 | 4/4 | |
| 6.1 | Identifies root cause: AddGroup must be called BEFORE AddCommand | ✓ | ✓ |
| 6.2 | Explains cobra does not retroactively assign groups | ✓ | ✓ |
| 6.3 | Shows correct order: AddGroup, GroupID, AddCommand | ✓ | ✓ |
| 6.4 | Does NOT suggest AddCommand before AddGroup as workaround | ✓ | ✓ |
| 7. validargsfunction — dynamic positional arg completion | 4/4 | 4/4 | |
| 7.1 | Uses ValidArgsFunction (not just ValidArgs []string) | ✓ | ✓ |
| 7.2 | Shows correct signature with args and toComplete parameters | ✓ | ✓ |
| 7.3 | Returns ShellCompDirectiveNoFileComp to suppress file fallback | ✓ | ✓ |
| 7.4 | Does NOT suggest parsing shell completion scripts manually | ✓ | ✓ |
| 8. registerflagcompletionfunc — flag value completion | 4/4 | 4/4 | |
| 8.1 | Uses RegisterFlagCompletionFunc (not ValidArgs on the command) | ✓ | ✓ |
| 8.2 | Shows function registered for the specific flag name "format" | ✓ | ✓ |
| 8.3 | Returns the static list ["json", "yaml", "table"] | ✓ | ✓ |
| 8.4 | Does NOT suggest modifying shell rc files or completion scripts | ✓ | ✓ |
| 9. test-isolation — fresh command tree per test | 4/4 | 4/4 | |
| 9.1 | Identifies cobra accumulates flag state across Execute() calls | ✓ | ✓ |
| 9.2 | Recommends building a fresh command tree per test | ✓ | ✓ |
| 9.3 | Suggests factory function (e.g. newRootCmd()) for fresh trees | ✓ | ✓ |
| 9.4 | Does NOT recommend -count=1 as the primary fix | ✓ | ✓ |
| 10. matchall — composing cobra validators | 4/4 | 4/4 | |
| 10.1 | Uses cobra.MatchAll() to compose validators | ✓ | ✓ |
| 10.2 | Shows cobra.MatchAll(cobra.ExactArgs(2), cobra.OnlyValidArgs) | ✓ | ✓ |
| 10.3 | Sets ValidArgs on the command for OnlyValidArgs to work | ✓ | ✓ |
| 10.4 | Does NOT suggest custom validator duplicating ExactArgs logic | ✓ | ✓ |
| 11. cobra-vs-viper — cobra does not handle config files | 5/5 | 5/5 | |
| 11.1 | Clearly states cobra does NOT handle config file parsing | ✓ | ✓ |
| 11.2 | Recommends viper (spf13/viper) for config file support | ✓ | ✓ |
| 11.3 | Explains cobra handles command tree/flags, viper handles config resolution | ✓ | ✓ |
| 11.4 | Mentions BindPFlag as the integration point between cobra and viper | ✓ | ✓ |
| 11.5 | Does NOT suggest cobra alone is sufficient for YAML/JSON config files | ✓ | ✓ |
| 12. cobra-cli-scaffolder — cobra-cli scaffolding tool | 4/4 | 4/4 | |
| 12.1 | Mentions cobra-cli as the scaffolding tool | ✓ | ✓ |
| 12.2 | Shows cobra-cli init or cobra-cli add commands | ✓ | ✓ |
| 12.3 | Explains cobra-cli generates command boilerplate | ✓ | ✓ |
| 12.4 | Does NOT suggest writing boilerplate from scratch as the only option | ✓ | ✓ |
Analyst pass: Almost all 12 evals score equally with and without the skill because the model's baseline cobra knowledge is strong — RunE, Args validators, OutOrStdout, SilenceUsage, and completions are well-documented. The only differentiating eval is 2.4 (explicitly discouraging len(args) checks in RunE as a backup). The self-grading methodology inflates the without-skill baseline — a human-as-judge rerun is recommended to measure true uplift, particularly on the subtler evals (4: PersistentPreRunE non-chaining, 6: AddGroup ordering, 9: test isolation with factory function).
golang-spf13-viper — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 52/53 (98%) | 53/53 (100%) | -2pp |
Full breakdown (53 assertions)
Model: Claude Sonnet 4.6 | Runs: 12 evals × 2 configs = 24 self-graded evaluations | Grading: self-graded (⚠️ biased — model sees assertions in prompt; without-skill baseline inflated; human-as-judge evaluation recommended for accurate uplift. Negative delta on eval 8.5: with-skill model mentions GetDuration as alternative, which fails the "Does NOT suggest" assertion; without-skill model omits it, which passes.)
| # | Assertion | With | Without |
|---|---|---|---|
| 1. env-key-replacer-nested-keys — SetEnvKeyReplacer for nested keys | 5/5 | 5/5 | |
| 1.1 | Identifies the root cause as missing SetEnvKeyReplacer | ✓ | ✓ |
| 1.2 | Explains viper preserves the dot in 'database.host' when looking up env var | ✓ | ✓ |
| 1.3 | Provides the fix: viper.SetEnvKeyReplacer(strings.NewReplacer(".", "_")) | ✓ | ✓ |
| 1.4 | Shows full setup requires prefix + replacer + AutomaticEnv together | ✓ | ✓ |
| 1.5 | Does NOT suggest renaming the config key to avoid dots | ✓ | ✓ |
| 2. sub-returns-nil — viper.Sub() returns nil for missing keys | 5/5 | 5/5 | |
| 2.1 | Identifies that viper.Sub() returns nil when the key doesn't exist | ✓ | ✓ |
| 2.2 | Shows nil check: if sub := viper.Sub("database"); sub != nil { ... } | ✓ | ✓ |
| 2.3 | Suggests returning error or using defaults when sub is nil | ✓ | ✓ |
| 2.4 | Does NOT suggest checking err from Sub() (returns no error, only nil) | ✓ | ✓ |
| 2.5 | Optionally suggests UnmarshalKey("database", &dbCfg) as an alternative | ✓ | ✓ |
| 3. config-file-not-found-graceful — ConfigFileNotFoundError handling | 5/5 | 5/5 | |
| 3.1 | Uses errors.As(err, ¬Found) with *viper.ConfigFileNotFoundError | ✓ | ✓ |
| 3.2 | Only propagates errors that are NOT ConfigFileNotFoundError | ✓ | ✓ |
| 3.3 | Continues execution normally when config file is not found | ✓ | ✓ |
| 3.4 | Does NOT use os.Stat or file existence check as the solution | ✓ | ✓ |
| 3.5 | Does NOT ignore all errors from ReadInConfig (real errors still propagate) | ✓ | ✓ |
| 4. global-viper-test-pollution — viper.New() for test isolation | 5/5 | 5/5 | |
| 4.1 | Identifies root cause as shared global viper state across tests | ✓ | ✓ |
| 4.2 | Recommends viper.New() per test for isolated instance | ✓ | ✓ |
| 4.3 | Shows v := viper.New() with v.SetConfigFile, v.ReadInConfig, etc. | ✓ | ✓ |
| 4.4 | Does NOT recommend viper.Reset() as the primary solution | ✓ | ✓ |
| 4.5 | Explains why tests are order-dependent (state persists to next test) | ✓ | ✓ |
| 5. unmarshal-mapstructure-tags — mapstructure struct tags required | 4/4 | 4/4 | |
| 5.1 | Identifies missing mapstructure struct tag as the root cause | ✓ | ✓ |
| 5.2 | Shows adding mapstructure:"max_conn" to the MaxConn field |
✓ | ✓ |
| 5.3 | Explains mapstructure does case-insensitive matching but not underscore-camelcase | ✓ | ✓ |
| 5.4 | Does NOT suggest using viper.GetInt as the fix | ✓ | ✓ |
| 6. bind-pflag-timing — BindPFlag must be before Execute() | 4/4 | 4/4 | |
| 6.1 | Identifies BindPFlag must be called before Execute() / before RunE runs | ✓ | ✓ |
| 6.2 | Recommends moving BindPFlag to init() or PersistentPreRunE | ✓ | ✓ |
| 6.3 | Explains cobra parses flags before RunE — binding after misses Changed state | ✓ | ✓ |
| 6.4 | Shows correct pattern: define flag + BindPFlag in init() | ✓ | ✓ |
| 7. viper-key-case-insensitivity — viper keys always lowercase | 4/4 | 4/4 | |
| 7.1 | Explains viper normalizes all keys to lowercase internally | ✓ | ✓ |
| 7.2 | Recommends using lowercase keys consistently in viper.Get* calls | ✓ | ✓ |
| 7.3 | Identifies viper.GetString("database_host") as the correct form | ✓ | ✓ |
| 7.4 | Notes DATABASE_HOST works via lowercasing but is a source of confusion | ✓ | ✓ |
| 8. duration-decode-hook — StringToTimeDurationHookFunc for time.Duration | 4/5 | 5/5 | |
| 8.1 | Identifies mapstructure cannot decode duration string without a hook | ✓ | ✓ |
| 8.2 | Shows viper.Unmarshal with DecodeHook using StringToTimeDurationHookFunc | ✓ | ✓ |
| 8.3 | Uses mapstructure.StringToTimeDurationHookFunc() as part of the hook | ✓ | ✓ |
| 8.4 | Does NOT suggest changing config value to nanoseconds | ✓ | ✓ |
| 8.5 | Does NOT suggest using viper.GetDuration as a workaround | ✗ | ✓ |
| 9. watch-config-atomic-rename-trap — fsnotify inode replacement by editors | 4/4 | 4/4 | |
| 9.1 | Explains vim writes atomically via rename (write-to-temp, then rename) | ✓ | ✓ |
| 9.2 | Explains fsnotify watches inode, replaced by rename-based writes | ✓ | ✓ |
| 9.3 | Recommends testing with direct writes (os.WriteFile or echo >>) not editor saves | ✓ | ✓ |
| 9.4 | Does NOT suggest downgrading vim or switching editors as the fix | ✓ | ✓ |
| 10. viper-alone-no-cobra — viper usable without cobra | 4/4 | 4/4 | |
| 10.1 | Clearly states viper can be used without cobra | ✓ | ✓ |
| 10.2 | Explains cobra is for command trees/flags, not needed for HTTP service | ✓ | ✓ |
| 10.3 | Shows viper setup without any cobra imports | ✓ | ✓ |
| 10.4 | Does NOT recommend adding cobra just for configuration purposes | ✓ | ✓ |
| 11. unmarshal-key-vs-sub — UnmarshalKey vs Sub+Unmarshal | 4/4 | 4/4 | |
| 11.1 | Recommends viper.UnmarshalKey("database", &dbCfg) as the simpler alternative | ✓ | ✓ |
| 11.2 | Explains it avoids the nil check required with Sub() | ✓ | ✓ |
| 11.3 | Shows correct usage: viper.UnmarshalKey("database", &dbCfg) | ✓ | ✓ |
| 11.4 | If mentioning Sub+Unmarshal, notes the nil risk | ✓ | ✓ |
| 12. allow-empty-env — AllowEmptyEnv for empty string env vars | 4/4 | 4/4 | |
| 12.1 | Explains viper treats empty string env vars as 'not set' by default | ✓ | ✓ |
| 12.2 | Introduces viper.AllowEmptyEnv(true) as the fix | ✓ | ✓ |
| 12.3 | Explains with AllowEmptyEnv(true), empty env var overrides config file value | ✓ | ✓ |
| 12.4 | Does NOT suggest using viper.Set() as a workaround | ✓ | ✓ |
Analyst pass: All 12 evals score equally or better without the skill because the model's baseline viper knowledge is strong — ConfigFileNotFoundError, viper.New() for tests, mapstructure tags, and UnmarshalKey are all well-documented. The single differentiating point is eval 8.5 (negative): the skill makes the model mention viper.GetDuration as an alternative, which fails the "Does NOT suggest" assertion. The skill's reference to the unmarshal section may be over-helpful on this eval. The self-grading methodology inflates the without-skill baseline throughout. A human-as-judge rerun is recommended targeting evals 1 (env key replacer — the subtlest viper gotcha), 8 (decode hook), and 12 (AllowEmptyEnv — lesser-known API).
golang-swagger — v1.0.0
| With Skill | Without Skill | Delta | |
|---|---|---|---|
| Overall | 58/60 (97%) | 43/60 (72%) | +25pp |
Full breakdown (60 assertions)
Model: Claude Sonnet 4.6 | Runs: 12 evals × 2 configs = 24 subagents | Grading: Human-as-judge
| # | Assertion | With | Without |
|---|---|---|---|
| 1. gin-blank-import — blank import + ginSwagger wire-up | 5/5 | 4/5 | |
| 1.1 | Adds a blank import of the docs package (e.g., _ "<module>/docs") |
✓ | ✗ |
| 1.2 | Imports github.com/swaggo/gin-swagger | ✓ | ✓ |
| 1.3 | Imports github.com/swaggo/files | ✓ | ✓ |
| 1.4 | Registers a GET route matching /swagger/*any using ginSwagger.WrapHandler | ✓ | ✓ |
| 1.5 | Does not suggest running swag init again (it was already done) | ✓ | ✓ |
| 2. map-response-type — map[string]bool cannot be used directly in @Success | 5/5 | 3/5 | |
| 2.1 | Does NOT use {object} map[string]bool directly in @Success |
✓ | ✗ |
| 2.2 | Defines a named struct for the response OR acknowledges a swaggertype workaround is needed | ✓ | ✗ |
| 2.3 | @Success annotation uses a named type (not a raw map literal) | ✓ | ✗ |
| 2.4 | @Router annotation is present with [get] method | ✓ | ✓ |
| 2.5 | @Produce annotation specifies json | ✓ | ✓ |
| 3. swaggertype-overrides — time.Time and []byte need swaggertype tags | 5/5 | 1/5 | |
| 3.1 | Adds swaggertype tag to CreatedAt field | ✓ | ✗ |
| 3.2 | Adds swaggertype tag to UpdatedAt field with the same treatment | ✓ | ✗ |
| 3.3 | Adds swaggertype:"string" and format:"base64" to the Payload []byte field | ✓ | ✗ |
| 3.4 | Preserves the json tags (does not remove them) | ✓ | ✓ |
| 3.5 | Does not leave time.Time fields without any swaggertype override | ✓ | ✗ |
| 4. chi-dynamic-basepath — Chi router + BasePath from env var | 5/5 | 5/5 | |
| 4.1 | Imports github.com/swaggo/http-swagger | ✓ | ✓ |
| 4.2 | Uses r.Get (chi method) to register the swagger route with a wildcard pattern | ✓ | ✓ |
| 4.3 | Sets docs.SwaggerInfo.BasePath using os.Getenv("API_BASE_PATH") or equivalent | ✓ | ✓ |
| 4.4 | Includes the blank docs import | ✓ | ✓ |
| 4.5 | Does not suggest rebuilding or running swag init per environment | ✓ | ✓ |
| 5. and-security-condition — && vs two separate @Security lines | 5/5 | 4/5 | |
| 5.1 | Uses && between security schemes on a single @Security annotation line | ✓ | ✗ |
| 5.2 | Does NOT write two separate @Security lines for AND semantics | ✓ | ✓ |
| 5.3 | References valid security definition names (ApiKeyAuth, BasicAuth, or similar) | ✓ | ✓ |
| 5.4 | Explains or implies that two separate @Security lines would mean OR, not AND | ✓ | ✓ |
| 5.5 | @Security line appears inside the handler doc comment block | ✓ | ✓ |
| 6. tag-exclusion — swag init --tags with ! prefix to exclude tags | 5/5 | 4/5 | |
| 6.1 | Uses the --tags flag (or -t) with swag init | ✓ | ✓ |
| 6.2 | Uses ! prefix to exclude tags (e.g., --tags '!Internal,!Admin' or similar) | ✓ | ✗ |
| 6.3 | Shows a complete swag init command | ✓ | ✓ |
| 6.4 | Does not suggest manually editing the generated swagger.json | ✓ | ✓ |
| 6.5 | Does not require writing custom Go code to filter endpoints | ✓ | ✓ |
| 7. godoc-comment-swag-fmt — godoc comment line required for swag fmt | 5/5 | 4/5 | |
| 7.1 | Adds // CreateOrder godoc as the first line of the comment block |
✓ | ✗ |
| 7.2 | godoc comment appears before any @ annotation | ✓ | ✓ |
| 7.3 | At minimum includes @Summary, @Router annotations | ✓ | ✓ |
| 7.4 | @Router specifies both path and HTTP method | ✓ | ✓ |
| 7.5 | Annotation block is placed directly above the function signature | ✓ | ✓ |
| 8. csv-vs-multi — collectionFormat(multi) for ids, collectionFormat(csv) for fields | 4/5 | 2/5 | |
| 8.1 | @Param for ids uses collectionFormat(multi) | ✓ | ✗ |
| 8.2 | @Param for fields uses collectionFormat(csv) or omits it (csv is the default) | ✗ | ✗ |
| 8.3 | Both params use []string or []int as data type | ✓ | ✓ |
| 8.4 | Both params are marked as not required (false) | ✓ | ✓ |
| 8.5 | @Router annotation is present with [get] method | ✓ | ✗ |
| 9. conditional-swagger — env-based swagger toggle without build tags | 5/5 | 5/5 | |
| 9.1 | Uses os.Getenv (or equivalent) to read APP_ENV at runtime | ✓ | ✓ |
| 9.2 | Conditionally registers the swagger route only when not in production | ✓ | ✓ |
| 9.3 | Does not require separate builds or build tags | ✓ | ✓ |
| 9.4 | The blank docs import is still present (or acknowledged as needed) | ✓ | ✓ |
| 9.5 | Solution works without recompiling between environments | ✓ | ✓ |
| 10. nested-composition — Envelope{data=model.User} nested composition syntax | 5/5 | 3/5 | |
| 10.1 | @Success annotation uses nested composition syntax with curly braces (e.g., Envelope{data=model.User}) | ✓ | ✗ |
| 10.2 | The inner type is the User struct (or equivalent named type) | ✓ | ✗ |
| 10.3 | Does not create a new wrapper struct just for documentation purposes | ✓ | ✓ |
| 10.4 | @Param for the id path parameter is present with path location | ✓ | ✓ |
| 10.5 | @Router specifies the correct path and [get] method | ✓ | ✗ |
| 11. enums-min-max-tags — enums/minimum/maximum struct tags | 4/5 | 4/5 | |
| 11.1 | Adds enums:"admin,editor,viewer" struct tag to Role field |
✓ | ✓ |
| 11.2 | Adds minimum:"0" and maximum:"100" struct tags to Score field |
✓ | ✓ |
| 11.3 | Adds json tags to all fields | ✓ | ✓ |
| 11.4 | Adds example tags to at least one field | ✗ | ✗ |
| 11.5 | Does not only describe constraints in a comment — they must be machine-readable struct tags | ✓ | ✓ |
| 12. swaggerignore-serialization — swaggerignore:"true" vs json:"-" | 5/5 | 4/5 | |
| 12.1 | Uses swaggerignore:"true" on the LastModified field | ✓ | ✗ |
| 12.2 | Keeps a valid json tag on LastModified (NOT json:"-") | ✓ | ✓ |
| 12.3 | Does NOT suggest removing the field from the struct | ✓ | ✓ |
| 12.4 | Other struct fields retain their json and swagger documentation | ✓ | ✓ |
| 12.5 | Explains or implies why json:"-" would be wrong here (breaks serialization) | ✓ | ✓ |
Analyst pass: The skill shows clear uplift on annotation mechanics that require knowing non-obvious swag behaviors: blank import requirement (eval 1), map type limitations (eval 2), swaggertype overrides for time.Time and []byte (eval 3), nested composition syntax (eval 10), and swaggerignore vs json:"-" (eval 12). Evals 4 (Chi+env) and 9 (conditional toggle) scored 5/5 in both conditions — the model already knows how to read env vars and register Chi routes, so these evals test common knowledge rather than skill uplift. Eval 11 shows a consistent miss on the example: struct tag in both conditions, indicating a coverage gap rather than discrimination; the skill should add an explicit example on example: struct tag usage. Future iterations should also target less-known swag behaviors: @extensions, @x- custom properties, multi-file init patterns (--dir flag), and OpenAPI 3.0 output.