mirror of
https://github.com/pbakaus/impeccable.git
synced 2026-09-19 09:36:59 +03:00
Fix critique's close on the right mechanism
The earlier fix in this branch was built on a wrong diagnosis. It assumed a
structured question hides any prose sharing its message, so it split report and
question across two turns. A controlled check showed prose before a question
renders fine; what hides a report is emitting it AFTER the question. The split
therefore fixed nothing and introduced a worse failure: a turn that ends on the
report is a turn that ends, and the questions never arrived at all.
Persistence returns to main's ordering, byte for byte, and the boundary prose is
gone. What replaces it is a position rule: the question is the last thing in the
response.
The trace test added here found two failures beyond the reported one. Critique
can fail to land in three ways, and they are now all asserted:
1. Question emitted before the report, hiding it behind the picker.
2. No close at all: no questions and no skip line, so polish inherits nothing.
3. Report authored into the persistence heredoc and never written to chat,
leaving a perfect snapshot and a user who sees nothing.
Mode 3 predates this branch entirely. Persistence step 1 now says the temp file
is an archive copy, not delivery.
The Codex final-question gate is promoted out of its <codex> fence, where it was
stripped for three of four providers, and the skip branch is now a countable
threshold (fewer than 3 Priority Issues) rather than a judgment call.
Known floor, recorded in the suite README: gpt-5.6-luna passes 1 run in 6 and
deepseek-v4-flash is flaky. claude-sonnet-5 and gemini-3.5-flash are consistent.
Prepared with AI assistance (Claude Code).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
d4e1b0902f
commit
ebc63f071a
@@ -59,9 +59,18 @@ The trace is the source of truth, not the model's free-form reply.
|
||||
| 15 | same iOS fixture; prompt is `/impeccable audit` | agent loads `reference/audit.native.md` (the Commands-table native variant, routed instead of `audit.md`) |
|
||||
|
||||
The workflow-contract file adds end-to-end assertions for attended fresh init,
|
||||
an initialized natural build request, replacement-world redesign, and scope-preserving bolder
|
||||
refinement. It checks question order and context/artifact writes rather than
|
||||
only reference-file loading.
|
||||
an initialized natural build request, replacement-world redesign, scope-preserving bolder
|
||||
refinement, and critique's closing question. It checks question order and
|
||||
context/artifact writes rather than only reference-file loading.
|
||||
|
||||
`critique closes with the question or an explicit skip line` is a regression
|
||||
guard, not a routing check. A critique that prints its report and then stops,
|
||||
asking nothing and printing no `Questions skipped: <reason>` line, is an
|
||||
incomplete run: the close is half the deliverable, and `polish` downstream has
|
||||
no priorities to inherit without it. The fixture page is deliberately broken
|
||||
enough to put the report past the three-Priority-Issue threshold, so the run
|
||||
cannot reach the skip branch on merit. The assertion is deliberately loose about
|
||||
*how* the run closes, because either close is valid; what it forbids is neither.
|
||||
|
||||
## Workflow-contract baseline (2026-08-13, current lineup)
|
||||
|
||||
@@ -77,6 +86,7 @@ failure beyond these.
|
||||
| initialized natural build | not measured | not measured | not measured | not measured |
|
||||
| redesign replaces DESIGN | flaky | not measured | not measured | not measured |
|
||||
| bolder refinement | not measured | pass | pass | **fail** |
|
||||
| critique closes | pass (2 of 2) | **fail (1 of 6)** | pass (2 of 2) | flaky (1 of 2) |
|
||||
|
||||
`not measured` means exactly that: the cell was never run in isolation on this
|
||||
lineup. Only the two failing scenarios were scoped per model, because the
|
||||
@@ -91,6 +101,32 @@ the 16-step cap. Confirmed identical on HEAD with `bolder.md` reverted, so it is
|
||||
not a skill-text problem. Same shape as the gpt-5.4-mini scenario 6/7 failures
|
||||
below: the model consumes the references and then declines to act.
|
||||
|
||||
**`critique closes`, gpt-5.6-luna.** Five runs while tuning the instruction text
|
||||
produced one pass. It fails in three distinct ways, which is why the scenario
|
||||
asserts on emission order rather than only on the presence of a question:
|
||||
|
||||
1. *No close.* Report lands, no question, no skip line. The failure this
|
||||
scenario was written for.
|
||||
2. *Question before report.* The question is emitted first and the report after
|
||||
it, so the report is withheld until the user answers. Observed directly, and
|
||||
the reason the invariant is stated as a position rule ("the question is the
|
||||
LAST thing in the response") rather than as prose order.
|
||||
3. *Report never spoken.* The full report is authored into the persistence
|
||||
heredoc, archived, and never written to chat. The snapshot is perfect and the
|
||||
user sees nothing. This is why persistence step 1 says the temp file is an
|
||||
archive copy, not delivery.
|
||||
|
||||
claude-sonnet-5 and gemini-3.5-flash pass consistently. deepseek-v4-flash has
|
||||
hit mode 3 once in two runs, so it is flaky here, not clean. Strengthening the
|
||||
instruction text moved luna from consistently failing to occasionally passing,
|
||||
and further prose tuning stopped paying.
|
||||
|
||||
Read the counts in the table as what they are: small samples on a nondeterministic
|
||||
system, gathered while the instruction text was being changed between runs. They
|
||||
say the close is reliable on the two strongest models and unreliable on the two
|
||||
cheapest ones. They do not support a finer claim than that. Re-measure rather
|
||||
than assuming when the lineup changes.
|
||||
|
||||
**`redesign replaces DESIGN`, flaky.** It has failed on two different assertions
|
||||
across runs (`designWrite > question` and `implementation > designWrite`), and on
|
||||
one run claude-sonnet-5 exhausted the 300s per-test timeout instead of asserting.
|
||||
|
||||
Reference in New Issue
Block a user