Codex re-review P2s on the fix wave, both verified:
- A finding AUQ rendering within TAIL_SCAN_BYTES of the gate (model waiting,
no further output) was vetoed by the blanket tail exclusion until timeout.
The veto is now ACTIVE-RENDER-aware: parseNumberedOptions anchors the last
cursor menu, so only a pending GATE menu vetoes; the judge fallback shares
the same check. Residual (documented): prose gate + prose finding inside
one tail — floors run the native-menu path in practice.
- The four demoted periodic tests are not in evals-periodic.yml's explicit
matrix (a named instance of the pre-existing periodic-orphans TODO), so
they run locally/manually until the PTY-capable periodic job lands.
CHANGELOG claim softened accordingly; TODO filed with the wiring recipe.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adversarial-review wording fixes (Claude adversarial F1-F8 + Codex
cross-confirmation), applied to both gate templates + regen:
- Host-anchored mode signal: only the host's own system messages (plan-mode
reminder or active plan file path) arm the auto-select; plan-shaped text
inside pasted documents, tool results, or fetched pages does NOT count —
injected content can't disarm the consent gate or nominate the target.
- Multiple plan candidates: the host-referenced plan file wins; still
ambiguous means ask.
- The DIFFERENT-target override carries the passing-mention guard.
- Plan mode + explicitly named target + no drafted plan resolves to the
named target instead of a contradictory re-ask.
- The numbered ask-path rules are qualified ('When no exception above
applied:') so they no longer restate an unconditional MUST-ask that
contradicts the exceptions.
- 'Whenever this gate does ask — in any mode — it is a hard STOP.'
- Shared preamble: 'any AskUserQuestion the skill fires is the workflow
operating within plan mode' (was 'the first AskUserQuestion is the
workflow entering plan mode', which framed the opposite of the bypass);
regenerates every skill.
- Ceilings ratcheted with attribution: plan-eng union ratio 1.10,
investigate 1.10 (the ~250B shared-preamble reword lands the
closest-to-ceiling skill at 1.092).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The extended plan-mode-no-op smoke invokes /plan-eng-review and
/plan-design-review, but the fresh CI containers registered only
office-hours and plan-ceo-review — both new runs would return
'Unknown command' and fail every PR's gate job (Codex structured
review P1, verified against evals.yml). Registration loops, the
dangling-target fail-fast list, and the frontmatter checks (now a
loop over the same skill list, so the lists can't drift) all cover
the two skills.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- no-op regression: gate-must-ask is now UNCONDITIONAL for eng/design (the
outcome==='asked' conditional let a silent-bypass plan_ready run sail
through); eng/design cases force --disallowedTools so the pinned prose
shape is contractual rather than hoping native AUQ renders match; the
named-target case uses trackTokens for consumption and lists
wrote_findings_before_asking in its diagnostic throw branch.
- tier-alignment invariant: both quote styles matched; zero-self-gate,
mixed-tier, and owning-keys-without-E2E_TIERS-entries are all REPORTED
instead of silently skipped (the fail-open holes three reviewers found).
- drift-guard: the generated gate menus must carry the exact question/option
strings the PTY question detector anchors on — free CI fails before the
paid smokes can go vacuous on a menu reword.
- touchfiles: corrected the no-op cost note for CI concurrency + retry
semantics.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review-army + adversarial findings on the scope-gate observability work,
all verified before fixing:
- Floor check: acceptance scanned the CUMULATIVE buffer while the scope-gate
exclusion scanned only the 1500-byte tail, so an early gate render satisfied
the floor vacuously once ~1.5KB of output accumulated (found independently
by 4 review passes; predicate reproduced). Acceptance now scans only content
APPENDED after the first gate render (positional anchor), and the LLM-judge
'waiting' shortcut no longer fires while the gate menu is the pending render.
- High-water flags are built once and spread at every return path — the
hand-spread pattern had already drifted (judge-waiting return omitted two
flags), which made must-stay-false asserts vacuous on those paths.
- isScopeGateAutoSelectVisible: tense-tolerant selected/selecting/selects
token (must-be-TRUE asserts shouldn't fail semantically-perfect paraphrases)
and quoted-occurrence rejection (a model verbatim-quoting the announcement
while declining must not trip must-stay-FALSE asserts). Fixtures added for
both directions.
- PlanSkillObservation outcome union gains 'wrote_findings_before_asking'
(returned at runtime via classifyVisible but missing from the type).
- trackTokens/tokensObserved: cumulative-buffer token high-water for
consumption asserts (the 2KB evidence tail is lossy and the plan-file
fallback is unreachable outside plan mode).
- New scope-gate-floor unit pins (from the ship coverage audit): both gate
render forms trip acceptance and exclusion; a genuine finding AUQ is not
excluded; tail-scoping semantics pinned.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
autoplan Step 3 reads plan-eng-review / plan-design-review SKILL.md verbatim,
and its section skip list omitted the scope gate — so autoplan ingested a
hard-STOP AskUserQuestion that contradicts its every-question-auto-decides
contract. One skip-list line fixes it; a static toContain pin in
skill-validation keeps the entry load-bearing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
In plan mode the scope gate's "What should I review? A/B/C" question is pure
friction: there is no branch diff and the target is the plan being drafted.
Both gates gain an ordered exceptions block, checked BEFORE asking:
1. Plan mode → auto-select B: review the active plan (in context or pasted),
announce it in one line ("Scope gate: plan mode — auto-selected B
(reviewing <target>)") so the user can interrupt; an explicitly different
user-named target still wins; no plan drafted yet → ask as normal.
2. User-named target (outside plan mode): explicit-only — a path, a pasted
doc, or the literal words "branch diff". A passing mention is not naming;
when in doubt, ask.
Outside plan mode with no explicitly-named target, nothing changes. Plan-mode
is checked FIRST because the PTY harness seeds drafts as pasted user messages
(claude-pty-runner.ts:1600) — ordering makes the seeded smokes deterministic.
Pinning: seeded plan-mode smokes assert no gate render + announcement rendered
(eng test 2; new design seeded test); plan-mode-no-op extends to eng/design
(bypass must not misfire outside plan mode; first question must be the gate)
plus a named-target case proving the pasted target is consumed; a drift-guard
asserts the two hand-duplicated exceptions blocks stay identical modulo the
two variant slots and carry the announcement string the detectors pin.
Skeleton ceilings ratcheted with comments (eng 68k, design 89k; eng union
ratio 1.08→1.09) — measured 67,006 B / 88,226 B after regen.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two render-shape-anchored detectors (whitespace-squished, like the Pattern-4/5
collapsed-form handling): isScopeGateQuestionVisible requires the question text
PLUS option-body text (native AskUserQuestion renders numbered options, prose
fallback renders lettered — the option body appears in both; narration doesn't),
and isScopeGateAutoSelectVisible requires the announcement prefix PLUS the
selected-B token.
runPlanSkillObservation gains scopeGateQuestionObserved /
scopeGateAutoSelectObserved high-water flags (attached at every return path) so
paid smokes can assert gate behavior across the whole run instead of the lossy
2KB evidence tail. runPlanSkillFloorCheck no longer counts a scope-gate render
toward auqObserved (tail-scoped exclusion) — the floor measures FINDING-driven
questions, and the gate could fire inside the 3s pre-target window.
Unit fixtures pin clean/native/collapsed positives, narration negatives, and
the verbatim template announcement string (template rewording fails here first,
before the paid smokes degrade to vacuous asserts).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The #2077 demotion of these four stochastic tests to 'periodic' was inert:
E2E_TIERS declared periodic but the files self-gated on EVALS_TIER === 'gate',
so they kept running in the blocking gate lane and never in the weekly lane.
Flip the four self-gates to 'periodic' (headers/describe labels updated), add
a free static tier-alignment invariant test (dep-list filename mapping;
unmapped self-gated files are reported, never silently skipped), and name the
two plan-mode test files in their own touchfiles dep lists so the invariant
binds for them.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(careful): warn on chained rm even when the last target is safe
The safe-exception block whitelisted rm -rf of build artifacts by
extracting targets with a single greedy match (.*rm ...), which only ever
inspects the LAST rm in the command. A chain like 'rm -rf /; rm -rf
node_modules' was therefore judged solely by its trailing safe target and
allowed without warning, waving through the destructive 'rm -rf /'.
Gate the shortcut to single rm invocations: when any shell separator
(; | & newline, incl. JSON-escaped \n/\r from the grep extraction path)
is present, fall through to the destructive-pattern check, which warns on
any recursive rm. Single-command artifact cleanups still allow.
Adds 3 regression tests covering semicolon and && chains in both orders.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* harden(careful): substitution separators + capital -R recursive flag (#2039)
Two residual fail-opens in the same guard PR #2040 hardened, both verified
by executing the script pre-fix:
- rm -rf $(./wipe-all)/node_modules silently allowed: the substitution token
ends in a whitelisted suffix and the safe-exception early exit skipped ALL
downstream checks. $( and backtick now count as chain separators; plain
$VAR expansion stays allowed.
- rm -R / silently allowed: both greps required a lowercase r in the flag
cluster; capital -R is the documented BSD/macOS recursive flag. Both greps
now match -[a-zA-Z]*[rR].
Six new tests: substitution x2 -> ask, capital-R x2 -> ask, rm -Rf
node_modules single-command -> still allowed, escaped-newline branch
(existing code, previously untested), and a pinned deliberate FP
(cd app && rm -rf node_modules -> ask) documenting the fail-closed
direction on chains.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(context-restore): prefer the current branch's own checkpoint (#2052)
All worktrees of a repo share one origin-derived slug, so they share one
`~/.gstack/projects/<slug>/checkpoints/` dir. `/context-restore` loaded the
newest checkpoint across the whole dir, so in one worktree it could silently
restore a *sibling worktree's* newer checkpoint.
Step 1 now orders candidates current-branch-first (read from each file's
`branch:` frontmatter), keeping other branches as a fallback. A branch is
checked out in at most one worktree, so this stops cross-worktree contamination
while preserving Conductor cross-branch handoff: when the current branch has no
checkpoint of its own, the full newest-first set is still used.
- scan the 200 newest before partitioning so a current-branch checkpoint sitting
below a burst of sibling saves is still found; output still capped at 20
- non-git / detached HEAD / branchless legacy saves fall back to the old
newest-first behavior (back-compat)
- +5 regression tests in context-save-hardening.test.ts (the #2052 bug case
fails on the old pipeline); regenerated SKILL.md + proactive-suggestions.json
Fixes#2052
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(gbrain): pass --confirm-destructive on drift re-register (#1985)
ensureSourceRegistered() handles match-but-different-path by removing the
old source then re-adding it at the new path. The remove was issued as
`gbrain sources remove <id> --yes`, but gbrain >= 0.42 gates `sources
remove` behind `--confirm-destructive` (`--yes` alone no longer suppresses
the data-loss prompt). The remove therefore fails with "To proceed, pass
--confirm-destructive", which ensureSourceRegistered surfaces as "source
registration failed" — aborting the entire /sync-gbrain code stage for any
already-registered source whose path has drifted. The memory and brain-sync
stages still pass, so the code index silently stops refreshing.
The orchestrator's own safeSourcesRemove() already passes
--confirm-destructive; this brings the lib helper in line with that
convention. Keeps --yes for older gbrain.
Tests: extend the fake gbrain shim in gbrain-sources.test.ts to simulate
the gbrain >= 0.42 guard (remove without --confirm-destructive exits 1),
update the drift re-register assertion, and add a regression test that
proves the drift path no longer throws. Both fail on main with the exact
"To proceed, pass --confirm-destructive" error and pass with the fix.
Fixes#1985
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* harden(gbrain-sources): route drift remove through #1734 guards + realpath drift check
Absorbing #2031 un-blocked a destructive remove that bypassed the #1734
data-loss guards: ensureSourceRegistered's drift path issued
`gbrain sources remove` directly, without the detectAutopilot +
decideSourceRemove checks every other remove routes through via
safeSourcesRemove. gbrain >= 0.42's own prompt was accidentally blocking
that path; with --confirm-destructive passed it is live again.
- Drift remove now refuses LOUDLY (throws, actionable message) while an
autopilot is active or when decideSourceRemove disallows; a silent
changed=false would hide the drifted registration.
- decideSourceRemove's extraArgs (--keep-storage when supported) propagate
to the remove call, matching safeSourcesRemove.
- Drift is realpath-normalized before being declared: a symlink alias of the
same directory (macOS /tmp -> /private/tmp) is a match, not drift — the
probable cause of #1985's reporter hitting the remove on an unmoved repo.
- Drift fires a loud stderr line (old -> new path); perpetual drift in logs
is the trigger for promoting #1985's reindex-in-place design.
Tests: autopilot-active refusal (no remove in call log), fail-closed refusal
on unreadable sources list, --keep-storage propagation, symlink-alias
no-drift; existing drift tests pin the guard probes so a live autopilot on
the dev machine can't flip them.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(developer-profile): exclude mode:resources rows from SESSION_COUNT, TIER, NUDGE_ELIGIBLE (#2067)
Every /office-hours run appends a mode:"resources" bookkeeping row alongside
the real session row, so --read double-counted sessions (~2x): tiers promoted
early and the builder-to-founder nudge armed prematurely. The file already
filtered resources rows for LAST_*/CROSS_PROJECT; the same realSessions
filter now feeds SESSION_COUNT/TIER, and the nudge predicate is the faithful
allowlist (mode === 'builder') so a future mode #4 fails closed instead of
re-opening this bug.
8 regression tests: count vs resources noise, tier boundaries both sides,
nudge false-with-noise / true-at-3-builders, cross-project trailing row.
Absorbed from PR #1991 by @mvann (fix + tests commits; the PR's version-bump
commit is superseded by this wave's consolidated release commit).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(hooks): passThrough() two-branch contract — never emit permissionDecision:'defer' (#2035, #2006)
Every AskUserQuestion died with "Tool result missing due to internal error"
on current Claude Code builds (Desktop 1.14271.0, CC 2.1.177). Root cause:
the question-preference-hook emitted permissionDecision:'defer' on every
pass-through path. 'defer' is a real PreToolUse value, but since CC v2.1.89
its semantics are "pause this tool call for external resumption" (headless
resume) — never "abstain". Interactive sessions have nothing to resume the
paused call, so the tool orphaned. Pre-2.1.89 builds ignored the unknown
value, which is why the hook worked when it shipped and broke later.
The fix is the two-branch pass-through contract:
- no context -> exit 0 with EXACTLY empty stdout
- memory nuggets present -> hookSpecificOutput with hookEventName +
additionalContext ONLY (the documented shape; plan-tune Layer 8 memory
injection ships through this branch and keeps working)
defer() is renamed passThrough() so the function says what it does, and
docs/spikes/claude-code-hook-mutation.md's protocol contract (cited by the
hook header) is corrected in the same commit — it taught '"defer" — let
permission flow continue' and was the reintroduction vector.
Test contract rewritten in the same commit (13 assertions across 3 files,
verified fail-first against the unfixed hook): pass-through paths assert
exact-empty stdout (a garbage/partial write cannot slip past an
optional-chained parse), the nugget path asserts permissionDecision is
ABSENT while additionalContext survives, and a new tripwire asserts no
non-deny path ever puts the string "permissionDecision" on stdout. The
deny (auto-decide) and Conductor prose-redirect paths are unchanged.
Deployment: no migration needed — settings.json points at the absolute
bash shim which execs the .ts live; /gstack-upgrade delivers the fix.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(one-way-doors): unify credential noun net + wire it into the runtime (#2024)
Library fix: revoke/reset/rotate now share ONE noun alternation (api key,
token, secret, credential, access key, password) with optional plural s?.
Pre-fix leaks: "reset my secret", "reset my access key", "revoke my secret"
(mismatched per-verb lists) and every plural form ("rotate the credentials",
"revoke all tokens" — \b(...)\b cannot match a trailing s).
Runtime wiring — the regexes could never fire in production before:
- gstack-question-preference --check gains --summary-stdin: the question
text pipes via stdin (never argv — summaries carry quotes/newlines/shell
metacharacters) and feeds isOneWayDoor alongside the id, so an ad-hoc
destructive question with a stored never-ask preference now forces
ASK_NORMALLY. Empty/absent stdin keeps exact id-only semantics.
- question-preference-hook falls back to classifyQuestion(question text)
when the registry lookup misses, so unregistered destructive questions
pass through to a human instead of auto-deciding.
- question-tuning resolver prose shows the piped form (SKILL.md regen lands
in the wave's release commit).
Tripwires (verified fail-first): full verbs x nouns x singular/plural matrix
with the #2024 repro rows, benign-summary no-over-match rows, stdin
transport survival (quotes/newlines), empty-stdin fail-safe, and hook
fallback both directions (destructive -> pass-through, benign -> deny).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(design): loud integer-flag contract for --count/--retry/--timeout (#2032)
design variants --count abc silently generated ZERO variants and exited 0:
parseInt(NaN) flowed through Math.min into the generation loop bound. The
same NaN class was live on the two sibling flags in the same file:
--retry abc made generate() a silent no-op (attempt <= NaN never true, null
output, exit 0) and --timeout abc killed the serve board ~immediately
(setTimeout(NaN)).
New design/src/flag-utils.ts: parseIntFlag (pure, unit-testable) +
normalizeIntFlag (CLI wrapper). Contract matches the --viewports precedent
(error loudly on nonsense — these commands spend real image-API money, a
silent fixup hides typos from calling agents): undefined -> default; bare
flag/empty/non-integer ("3.7" rejected, not truncated)/below-min -> exit 1
with usage hint; above-max -> clamp with stderr warning. --count normalizes
at the variants() consumption site so programmatic callers are covered, with
the ceiling derived from STYLE_VARIATIONS.length instead of a magic 7; the
CLI passes the raw flag through (a pre-parseInt would truncate "3.7").
Tripwires live in test/design-flag-utils.test.ts — deliberately under test/,
not design/test/, which is invisible to the bun test glob, TEST_ROOTS, and
every workflow (wiring design/test/ into CI is a captured TODO).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(gbrain): thin-client state — remote-MCP brains no longer classify as broken-config (#2051)
A thin client (remote-HTTP MCP brain, no local engine by design) probed
`gbrain sources list`, which gbrain's dispatch guard REFUSES on thin clients
(exit 1, no recognized error string), so the classifier fell to its
defensive broken-config default and every suppression gate silently hid
brain-aware blocks from exactly the users on a shared team brain.
New 'thin-client' state, detected PRE-probe from gbrain's own remote_mcp
config marker via the existing gbrainConfigPath() helper (mirrors gbrain's
isThinClient(); honors GBRAIN_HOME; zero network, immune to error-string
drift), with a /thin[- ]client/ stderr backstop in the probe catch. Remote
reachability is deliberately NOT probed by the classifier — that is the
#1964 pathology; gbrain calls degrade gracefully at use time, and the detect
JSON says so honestly (gbrain_thin_client: {probed: false}).
The state is admitted at every suppression gate — gstack-gbrain-detect
--is-ok (drives setup + gbrain-refresh), gen-skill-docs' detection override,
gstack-config gbrain-refresh — while the sync stages (code/memory/dream)
SKIP with an accurate reason: code indexing runs on the brain server, memory
syncs via the remote brain's artifacts pull. The two consumer classes need
opposite answers, which is why this is a distinct state and not a
skip-the-probe special case. sync-gbrain Step 1.5 and setup-gbrain prose
route thin-client to proceed, never into broken-config remediation.
detectMcpMode secondary generalization: url-match against the config's
remote_mcp.mcp_url (deterministic — gbrain mounts at the generic /mcp path)
-> name pattern gbrain[-_]* -> stdio command token; gbrain_mcp_mode stays a
3-value enum.
Tripwires: end-to-end --is-ok exits 0 on a thin-client fixture AND still
exits 1 on broken-config (the gate didn't widen); pre-probe + stderr-fallback
classifier paths; 4 detectMcpMode identification cases incl. a non-matching
url that must NOT false-positive.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* release: v1.60.0.0 — regen SKILL.md, VERSION, CHANGELOG, TODOS follow-ups
- Regenerate all SKILL.md from templates (question-tuning --summary-stdin
prose from #2024, context-restore branch preference from PR #2054,
sync-gbrain/setup-gbrain thin-client prose from #2051) + llms.txt.
- VERSION + package.json -> 1.60.0.0 (bin/gstack-next-version, queue-aware:
#1815 claims 1.59.0.0, #2213 claims 1.59.1.0).
- CHANGELOG release summary + itemized entry crediting @jbetala7 (x3) and
@mvann.
- TODOS.md: three eng-review follow-ups (design/test CI wiring + documented
pre-existing retry-after flake, /context-save worktree identity, gbrain
reindex-in-place conditional on the new drift log).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(resolvers): compress --summary-stdin preamble prose to fit parity budget; re-bless ship goldens
The v1.57.7.0 parity suite caps investigate's generated size at 1.09x
baseline; the #2024 question-tuning prose (duplicated into every tier->=2
skill) tipped it to 1.092. Compressed to a single inline command + short
pointer (the full rationale lives in bin/gstack-question-preference's
header and the one-way-doors module docs). Ship goldens re-blessed against
the final resolver text (conscious template-change acknowledgment, per the
golden-file regression contract).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(e2e): office-hours-spec-review turn budget fits the carved skill layout (#2473)
The test failed deterministically with error_max_turns at 9 turns on main
and this branch alike (CI attempt logs + local main repro). Root cause from
the failing transcript: the Spec Review Loop content is carved out of
office-hours/SKILL.md into office-hours/sections/, so the agent needs
discovery hops (grep SKILL.md -> ls sections/ -> read the section) before it
can write — 8 tool turns + the closing text turn = 9 > the 8-turn budget,
which predates the carve. Observed failures wrote a CORRECT summary on tool
turn 8 and died on the closing turn.
maxTurns 8 -> 12. Verified: PASS locally post-fix (7 turns this run — the
extra headroom absorbs discovery-path nondeterminism).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(e2e): review-dashboard-via session budget survives runner contention (#2473)
The test failed on CI (and its baseline run) with the timeout signature:
0 turns, $0.00, exactly 183s, 3/3 attempts — the spawned claude -p session
never emitted a single stream event before the 180s inner timeout. The
file's tests run concurrently on one runner; session startup queues behind
sibling sessions, and this test had the tightest budget in the file (the
240s-budget tests in the same job passed). A clean local run takes 270s
wall for 4 turns, confirming 180s was too tight even without contention.
Inner timeout 180s -> 300s; outer bun timeout 240s -> 360s to keep headroom
over the inner budget. Verified: PASS locally post-fix (4 turns, 270s).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(e2e): retro-base-branch session budget survives runner contention (#2473)
Same class as review-dashboard-via, one test over in the same file: /retro
is a long multi-step flow whose clean pass measures 225-239s — a coin flip
against the 240s inner budget. First CI run passed at 225s; the rerun timed
out at the 240s line on all 3 attempts (exitReason "timeout"); the local
verification run passed at 239s, ONE second under the old cap.
Inner timeout 240s -> 360s; outer bun timeout 300s -> 480s for headroom.
Verified: PASS locally post-fix (17 turns, 239s).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Jayesh Betala <jayesh.betala7@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Michael Vann <9221873+mvann@users.noreply.github.com>
* fix(test): eval-list CLI spawns from neutral cwd so slug detection can't dodge the fixture store
getProjectEvalDir() probes cwd-relative .claude/skills/gstack/bin/gstack-slug;
with cwd=ROOT on a dev machine the self-symlink makes it succeed, routing reads
to an empty project-scoped dir instead of the seeded legacy ~/.gstack-dev/evals.
Neutral cwd + absolute script path fails both probes deterministically — same
behavior as CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): case-insensitive remediation-hint match for reworded Gemini NOT-READY message
The message now leads with 'Export GEMINI_API_KEY...' (free-tier OAuth
deprecation); the old pattern only knew lowercase 'export'.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(test): observability check 11 floor 6 -> 5 after shell-free spawn removed promptFile unlink
aa3bd6f0 deleted the prompt temp file (and its /* non-fatal */ marker) when it
dropped shell interpolation. The invariant — every runner I/O path wrapped
non-fatally — still holds at the 5 remaining sites, now named in the comment.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(todos): P1 — free-suite exit code masked by in-process force-exits
Five browse test files setTimeout(() => process.exit(0), 500) inside the shared
bun process; the suite can exit 0 before the summary with real failures masked.
Receipts + fix path filed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(release): v1.60.2.0
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
urlBlocklistFilter compared URLs against the exfiltration blocklist with
case-sensitive substring checks, so an uppercased sink (https://WEBHOOK.SITE/x)
bypassed the guard and reached the scoped browser agent. Normalize the page URL
and extracted content URLs to lowercase before comparing, and make URL
extraction scheme-insensitive so HTTPS:// links are still caught.
Fixes#2190
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Parses package.json overrides + every resolved basic-ftp specifier in bun.lock
(including nested paths like get-uri/basic-ftp) and fails if any is below 5.3.1.
Deterministic and offline. Fires on the pre-fix tree (basic-ftp@5.2.0) and would
also catch a direct-dependency-only bump that leaves a nested vulnerable copy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
basic-ftp reaches the tree only transitively:
puppeteer-core > @puppeteer/browsers > proxy-agent > pac-proxy-agent > get-uri > basic-ftp
Versions <= 5.3.0 carry four HIGH advisories, all fixed in 5.3.1:
- GHSA-chqc-8p9q-pq6q (CVE-2026-39983) FTP command injection via CRLF
- GHSA-6v7q-wjvx-w8wg incomplete CRLF protection (USER/PASS + MKD bypass)
- GHSA-rpmf-866q-6p89 DoS via unbounded multiline control-response buffering
- GHSA-rp42-5vxx-qpwr DoS via unbounded memory in Client.list()
Pin via a bun `overrides` entry rather than a phantom direct dependency.
Overriding forces every basic-ftp in the tree to 5.3.1, including get-uri's
nested copy; `bun audit` then reports zero basic-ftp advisories (total 37 -> 33,
HIGH 13 -> 9). A direct-dependency bump leaves get-uri/basic-ftp at the
vulnerable version and audit still flags all four.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
endpoint-hash returns "local" for stdio/PGLite, but validators only
allowed hex suffixes, so setup-gbrain could not persist trust policy.
Co-authored-by: Cursor <cursoragent@cursor.com>
Exercise resultFromGeminiStream against current stream-json fixtures
(including empty-success hardening). Recognize GEMINI_API_KEY and map
IneligibleTierError to auth — personal OAuth free-tier is no longer
supported by gemini CLI.
Co-authored-by: Cursor <cursoragent@cursor.com>
GeminiAdapter was reading message.text and result.usage, so current CLI
content/stats events produced empty $0 success rows. Accept content with
an assistant role guard, stats token fallbacks, init model, and treat
empty exit-0 output as an error (#2159).
Co-authored-by: Cursor <cursoragent@cursor.com>
Root cause of months of silent local failure: the sandbox copied skill dirs to
the repo root, but claude >= 2.x resolves slash commands strictly from
registered skills, so /autoplan short-circuited with 'Unknown command' (0
turns, ~1s) on every attempt. Install /autoplan + review skills at
project-level .claude/skills/ (same pattern as skill-routing-e2e).
Also: the transcript filter matched entry.type === 'tool_use', a shape that
never appears at the top level of raw stream-json, so assertions only ever saw
the final result text; filter on assistant/user events instead. Hang
protection accepts the Phase 1 review dispatch (Agent/Task tool call carrying
review instructions) as progress evidence, since full Phase 1 completion is
15+ min of subagent work. Budget raised to 10 min / 40 turns.
Invisible in CI: the file is in neither evals.yml nor evals-periodic.yml
matrices (coverage decision filed in TODOS.md).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
proc.kill() only signals the sh -c wrapper; the claude child survives as an
orphan that inherited our stdout/stderr pipes, blocking the stream drain until
it exits (observed: a 600s spawn timeout stretching to 1431s and tripping bun's
per-test timeout with no result). On timeout, cancel the stdout reader; race
the stderr drain against child exit + 5s grace. Streamed transcript lines
survive the cancel, so callers still get their evidence.
Regression test: test/session-runner-timeout.test.ts (fake claude spawns a
pipe-holding orphan; fails in 30s without the fix, passes in 8s with it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: first-run activation — project-aware scaffold, router front door, onboarding nudges
Adds the activation system that drives a new install toward a concrete first move:
- bin/gstack-first-task-detect: local-git+filesystem repo classifier emitting one
validated enum bucket (greenfield/code_<lang>/branch_ahead/dirty_default/clean_default),
portable timeouts, fail-safe empty output.
- generate-first-run-guidance.ts: unified preamble section — first-run project-aware
scaffold + returning-session plan->review->ship tip, gated on a persistent .activated
marker and never run in headless. Detection wired lazily in generate-preamble-bash.ts.
- SKILL.md.tmpl: top-level gstack skill is now a pure router (browse body removed; it
lives in /browse), routing any request and sending browser/QA work to /browse.
- setup: first-move nudge on first install. office-hours: closing handoff that launches
the next review via the Skill tool.
- telemetry-ingest: accept onboarding/first_task_scaffold_shown/handoff/route event types.
* test: cover first-run detection + repoint browse-content assertions to /browse
- New unit tests for every detection bucket, the eval-safe enum contract, and the
first-run gating (test/preamble-first-task-scaffold.test.ts); periodic E2E that runs
the detector through the real harness (test/skill-e2e-first-task-scaffold.test.ts).
- Repoint browse-content assertions (gen-skill-docs, audit-compliance, skill-validation,
LLM-judge eval) from the root skill to browse/SKILL.md following the router split;
add a regression pinning that the router carries no browse body.
- Register first-task-scaffold touchfiles + periodic tier; bump parity/carve size caps
~1-2KB per skill for the shared first-run-guidance preamble section.
- Refresh ship golden fixtures for the preamble addition.
* chore: regenerate SKILL.md + llms.txt for first-run activation
* chore: bump version and changelog (v1.58.5.0)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(test): repoint bws skillmd-* setup-block assertions to browse/SKILL.md
The skillmd-setup-discovery / -no-local-binary / -outside-git E2E tests extracted
the `## SETUP`→`## IMPORTANT` browse binary-discovery block from the root SKILL.md.
P2 moved that block to browse/SKILL.md (end anchor is now `## Core QA Patterns`),
so the slice came back empty and the `browse/dist/browse` guard failed. Repoint to
browse/SKILL.md. Verified: 7/7 e2e-browse pass locally.
* fix(test): tolerate skill-discovery race in PTY plan-mode smoke
The e2e-pty-plan-smoke suite (office-hours / plan-mode-no-op) failed in CI with
`Unknown command: /office-hours` (claude exited ~10s) while passing locally. Root
cause: a cold CI container's overlay-FS scan of the symlinked ~/.claude/skills
registry finishes AFTER the runner's 8s boot grace, so the first `/skill` send
reaches claude before the skill is indexed and is rejected as unknown. The runner
gave up on the first "Unknown command:" line.
runPlanSkillObservation now re-sends the skill command up to 3x (6s apart),
re-marking the buffer each time so stale scrollback can't re-trip the check,
before concluding the skill is genuinely unregistered. A real dangling-symlink /
missing-skill still surfaces as 'exited' (after retries), preserving the original
diagnostic. Pure-helper contract unchanged: 95/95 unit tests pass.
This is a pre-existing harness bug (fails identically on #2077's own branch, which
introduced the suite) surfaced while shipping the activation feature.
* debug(ci): temporarily instrument pty-smoke skill discovery
Capture claude version, env, registry tree, and a claude -p discovery probe to
pin why /office-hours isn't discovered in CI (retries proved it's not a race).
Temporary — revert once the registry fix is identified.
* chore: revert pty-smoke harness experiments (race-retry + CI debug step)
Diagnosis is conclusive and the experiments aren't the fix, so restore the
harness to its original state (net-zero diff vs main for both files).
What the CI debug step proved: `claude -p` returns READY — claude v2.1.187 fully
DISCOVERS /office-hours from the symlinked registry. Only the interactive PTY TUI
rejects it as "Unknown command" (and it received the full command text). So the
e2e-pty-plan-smoke failure is a claude 2.1.187 interactive-TUI regression (skills
discovered by `claude -p` aren't exposed as TUI slash commands), pre-existing in
the #2077 harness and failing identically on its own origin branch — unrelated to
this activation PR. The race-retry can't help (the TUI genuinely lacks the
command); the debug step also tripped actionlint (shellcheck SC2012). Both reverted.
* fix(ci): copy SKILL.md as real files in pty-smoke registry (cross-mount symlink)
The e2e-pty-plan-smoke suite failed with "Unknown command: /office-hours" in CI
while passing locally. Root cause (proven, not guessed): claude 2.1.187's
interactive-TUI skill scanner does not follow the /github/home -> /__w cross-mount
symlink the registry used for per-skill SKILL.md. Evidence: a CI debug step showed
`claude -p` discovered the skill (printed READY), and a local macOS repro with the
identical symlinked registry recognized /office-hours — isolating the failure to
the container's cross-mount symlink, not registration content, claude version,
duplicate names, or a race.
Fix: register the per-skill SKILL.md + sections as REAL copies (same mount as
$HOME) so the TUI reads them directly. The gstack root stays a symlink — the
preamble's runtime bash resolves bin/* and sections/* through it and bash follows
cross-mount symlinks fine.
* fix(ci): guard rm expansion in pty-smoke registry (shellcheck SC2115)
* fix(ci): also register pty-smoke skills project-scoped (cwd/.claude/skills)
The real-file user-dir registration still left the TUI rejecting /office-hours in
the container. claude's interactive TUI surfaces /slash commands from the PROJECT
dir (<cwd>/.claude/skills); the smokes run with cwd=$REPO whose .claude/skills is
gitignored (absent on a fresh CI checkout), so the user-dir registry feeds
`claude -p` (READY) but not the TUI. Populate $REPO/.claude/skills with real
SKILL.md + sections copies (no gstack symlink there — it would point at its own
parent; runtime paths use the user-dir gstack symlink).
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(gbrain): stop forcing GBRAIN_PREPARE on transaction-mode poolers (#1965)
buildGbrainEnv auto-set GBRAIN_PREPARE=true whenever DATABASE_URL targeted
port 6543, and the /sync-gbrain capability check exported it for the rest
of the skill run. Both had the semantics inverted: gbrain auto-disables
prepared statements on transaction-mode poolers because they break every
write there ("prepared statement does not exist"); GBRAIN_PREPARE=true is
gbrain's documented override for SESSION-mode poolers on 6543, not a
requirement for transaction mode. The #1435 search symptom the auto-set
worked around was fixed gbrain-side.
Remove both force-sets. A caller-set GBRAIN_PREPARE (either value) still
passes through untouched, preserving the session-mode-on-6543 escape hatch.
isTransactionModePooler stays exported.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(gbrain): classify probe timeout as its own status; sync proceeds instead of skipping (#1964)
The 5s engine probe misclassified healthy-but-slow engines (cold Supabase
pooler connections measured at 6.9-10.7s) as broken-config, so /sync-gbrain
silently skipped code+memory and told the user their config was malformed.
- New "timeout" status: probe killed at the deadline with no recognized
stderr pattern. Default deadline is now 15s, overridable via
GSTACK_GBRAIN_PROBE_TIMEOUT_MS (tests set 300ms against a fake that
sleeps 2s).
- Sync stages PROCEED on timeout with a stderr warning naming the env knob;
a genuinely-dead engine surfaces its real error at the first operation
instead of a false config diagnosis.
- Consistency everywhere "ok" gated behavior: gstack-gbrain-detect --is-ok
exits 0 on timeout, and gen-skill-docs' detection gate accepts it, so a
slow engine no longer silently suppresses brain-aware features.
- Status cache: key now includes the effective probe timeout (raising it
invalidates a cached timeout) and GBRAIN_HOME; config detection honors
GBRAIN_HOME so relocated-home users stop being misclassified as
missing-config.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(bins): cygpath-normalize SCRIPT_DIR for bun imports; surface learnings-log errors (#1950)
Under Windows git-bash, pwd yields a POSIX path (/c/Users/...) that Bun on
Windows cannot resolve as an ES module specifier. gstack-learnings-log
interpolates SCRIPT_DIR into a bun -e import, so every invocation died with
"Cannot find module" — and 2>/dev/null swallowed the error, silently
dropping every AI-logged learning for Windows users.
- 3-line cygpath -m guard in gstack-learnings-log and gstack-question-log
(which gains the same import shape in the next commit). Matches the
duplicated IS_WINDOWS convention in setup; no shared shell lib exists.
- learnings-log adopts question-log's set +e / TMPERR capture pattern
wholesale: validation errors now print to stderr. The old
`if [ $? -ne 0 ]` check was dead code under set -euo pipefail — the
script exited at the failing assignment before reaching it.
- New test/bin-windows-bun-import-paths.test.ts: static invariant (any
bash bin interpolating $SCRIPT_DIR into a bun -e import must carry the
guard) + behavioral end-to-end run invoked via `bash <bin>` — added to
the windows-free-tests workflow list so the conversion is proven on the
only platform where the bug exists.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(question-log): dedupe INJECTION_PATTERNS via lib/jsonl-store (#1934)
bin/gstack-question-log carried a local copy of the injection-pattern list,
so pattern fixes to lib/jsonl-store.ts never propagated — including the
/override[:\s]/i false-positive fix arriving via community PR #1940.
Import the shared hasInjection instead (enabled by the previous commit's
cygpath guard). question-log also gets the lib's stricter superset
(human:, disregard, from-now-on, approve-all patterns).
Tests pin the contract in a #1940-order-independent way: an "Override:
ignore all previous instructions" header is rejected, "prose overrides the
deterministic table" is accepted, and a static invariant keeps local
INJECTION_PATTERNS duplicates out of the bin.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): community-pulse + both dashboards never report fake zeros (#1947)
The security-signaling surface failed open at three layers — every failure
mode read as a reassuring "0 attacks" / "0 installs":
- community-pulse edge function: supabase-js returns {data,error} without
throwing, and all five queries discarded `error` — a DB outage produced
real-looking zeros via the SUCCESS path, and the catch (also returning
zeros with HTTP 200) was unreachable for query failures. Every query now
destructures and throws; the catch serves the stale cache (marked
"stale": true) when one exists, else 503 {"error":"pulse_unavailable"}.
Success responses carry "status":"ok" so clients can distinguish
authoritative data from legacy backends. NOTE: the edge function deploys
out-of-band (supabase functions deploy community-pulse).
- gstack-security-dashboard: captures the HTTP status; non-200 / network
failure / error body / missing section → "unknown — backend error";
jq missing → "unknown — install jq" (the lossy grep fallback broke on
nested arrays and under-reported attacks as zero — removed); a 200
without the new marker shows figures with an "unverified (legacy
backend)" note. Also fixes a latent display bug: the TOTAL grep matched
the digit 7 inside "attacks_last_7_days" and misreported every count.
- gstack-community-dashboard: same class — curl || echo "{}" plus
grep || echo "0" printed "Weekly active installs: 0" on any failure.
Now "unknown — backend error (HTTP N)".
test/security-dashboard-fallback.test.ts pins the matrix (200+marker,
200-legacy, 503, network failure) x (jq present, jq absent) for both bins:
"unknown" states never render as 0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(telemetry): redact error_message spans before they leave the machine (#1947)
error_message was uploaded with only quote/newline escaping — stack traces
and failed-API errors can embed credentials, private paths, and hostnames,
and the sync path strips only _repo_slug/_branch.
New lib/redact-engine.ts export redactFindingSpans(): replaces EVERY
finding's span with <REDACTED-{id}> regardless of tier (applyRedactions is
the interactive PII-only path and exits nonzero on credential findings, so
it can't serve machine egress). Returns null when a span can't be located —
callers drop the whole payload rather than risk a leak.
gstack-telemetry-log pipes error_message through it at LOG time, so the
local JSONL at rest is clean too; surrounding text survives for crash
triage. FAIL CLOSED: bun missing, engine error, or non-JSON-string output
all null the field. Tests pin: embedded ghp_ token → <REDACTED-github.pat>
with context intact; redactor unavailable → null; raw bytes on disk never
contain the token.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(redact): prepush guard fails closed on git failure; /ship owns hook install (#1946)
Two gaps closed:
1. Fail closed. The git() helper returned "" on ANY non-zero exit or
maxBuffer overflow (status null), addedLinesFor produced an empty
string, and the push sailed through unscanned — fail-open on exactly
the oversized-diff case where a large secret-bearing blob is most
likely. The diff call now uses a strict variant that throws; main
blocks with a clear message naming the GSTACK_REDACT_PREPUSH=skip
escape valve. Probe calls (symbolic-ref, rev-parse, merge-base) keep
the permissive helper — their failures are normal control flow.
2. Install path. The hook was installed by nothing ("opt-in, installed by
nothing" was the issue's words). ./setup runs in the gstack checkout —
the wrong repo for a per-project hook — so it gets a one-line hint
only. /ship owns per-repo install: config redact_prepush_hook=true +
hook missing → silent install (consent already given); config unset +
no ~/.gstack/.redact-prepush-prompted marker → one-time machine-wide
AskUserQuestion offer, answer persisted. ship/SKILL.md regenerated in
this same commit (check-freshness bisect discipline).
Tests: unscannable diff (bogus SHAs) → exit 1 + valve named; empty-but-
successful diff → exit 0; static asserts pin setup as hint-only and the
ship template as the installer surface.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(redact): six new credential patterns — GitLab, HuggingFace, npm, DigitalOcean, Bearer, GCP SA (#1946)
Coverage gaps from the #1946 security review, including token types for
tooling gstack itself drives (glab):
HIGH (block): gitlab.token (glpat-/glptt-/gldt-), huggingface.token (hf_),
npm.token (npm_), digitalocean.token (dop_v1_), gcp.service_account (the
JSON-escaped "private_key" form that dodges pem.private_key's literal-block
match when minified, confirmed by "private_key_id" proximity).
MEDIUM (warn): auth.bearer — the most FP-prone shape in the set (docs are
full of "Authorization: Bearer <token>"), so it requires header-context
proximity and the same entropy>=3.0 + placeholder validator recipe as
env.kv. "Bearer YOUR_TOKEN_HERE" never fires; calibration over coverage,
per the cries-wolf principle.
All shapes are linear-time; test/redact-pattern-lint.test.ts covers them
automatically. Engine tests add positive + placeholder-negative cases per
pattern.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: coverage-audit additions for the fix wave
Ship Step 7 gap-fill (all passing, 248 tests across the touched suites):
memory + dream stage probe-timeout proceeds, gbrain-detect override paths,
stale-flag passthrough, 200-body-missing-.security fail-closed case,
telemetry redaction edges, and credential-pattern edge cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: pre-landing review fixes
Review army findings (1 critical, auto-fixed with regression tests):
- CRITICAL (security specialist, verified live): redactFindingSpans spliced
only the regex capture span, and pem.private_key / gcp.service_account
capture just the BEGIN-header — the key body survived "redaction" and
shipped via telemetry. Marker-only patterns now drop the whole payload
(null, fail closed). Overlapping spans (Bearer+JWT on the same bytes) are
coalesced before splicing so stale offsets can't leave partial secret
bytes behind.
- gitStrict: drop the dead `|| r.status === null` disjunct (null !== 0
already covers it); add the signal-kill/null-status regression test the
docstring promised.
- security-dashboard human mode flags stale snapshots ("figures may be out
of date") instead of presenting frozen counts as current.
- community-dashboard marker check uses jq when available — the grep-only
variant misclassified whitespaced/reserialized bodies as legacy.
- telemetry fail-closed test now shadows bun with a failing stub
(deterministic on any host layout); stale "five status cases" describe
title renamed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: adversarial review fixes (Claude + Codex cross-model passes)
Both adversarial passes ran against the wave; every FIXABLE finding landed
with a regression test:
- probeTimeoutMs clamps to >=1ms: a fractional override floored to 0, and
execFileSync treats timeout:0 as NO timeout — the probe that exists to
bound hangs could hang forever (found by both models independently).
- /ship silent hook install now requires the hooks dir to live inside
.git: with core.hooksPath (husky's COMMITTED .husky/), the chaining
installer would have renamed the team's committed pre-push and written a
machine-local wrapper into the working tree (found by both models).
- gstack-config gbrain-refresh accepts the "timeout" status — the last
consumer still gating on literal "ok" (Codex); gstack-gbrain-detect's
config-derived fields honor GBRAIN_HOME so the detection JSON can't
report status ok alongside config_exists false (Codex).
- prepush: a remote sha absent locally (shallow clone / stale fetch) falls
back to the merge-base/empty-tree range — scans MORE, never blocks a
legitimate push into training users toward --no-verify.
- dashboards: curl's own 000 no longer doubles to "HTTP 000000"; the
community dashboard flags stale snapshots like the security one; array
sections parse via jq (the sed/grep loops truncated at the first ']');
the no-jq marker grep tolerates whitespace.
- telemetry: multi-line redactor output nulls the field instead of
corrupting the JSONL record; setup's hint fires only when the config key
is genuinely unset (an explicit false is a recorded decline); the /ship
prompt marker honors GSTACK_HOME.
Kept as designed (cross-model tension noted): Bearer stays MEDIUM in the
prepush gate — a HIGH Bearer would block every docs example; the entropy
validator can't eliminate that FP class, and MEDIUM warns visibly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: bump version and changelog (v1.57.11.0)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: P1 TODO — eval harness live progress + incremental persistence
Root-caused during this ship: a killed eval run was indistinguishable from a
healthy one for hours (per-file output buffering across mega test files, no
incremental eval-store writes, no honest liveness signal). Full context and
starting points in the entry.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: fix operational-learning E2E fixture — copy lib/jsonl-store.ts
Pre-existing breakage, proven on main: gstack-learnings-log has imported
lib/jsonl-store.ts (shared injection patterns) since v1.57.5.0 / #1910, but
the fixture copies only the bin scripts — the bin exits 1 before writing
anything, on main silently (stderr swallowed) and on this branch loudly
(the #1950 error-surfacing made the four-day-old failure visible). A real
install always ships bin/ and lib/ together; the fixture now does too.
Verified: the fixture-shaped invocation writes the learning (exit 0) with
lib present, exits 1 on both main and this branch without it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ios-qa): isolate E2E tests under --concurrent (3 real races)
The ios-qa E2E file failed intermittently under `bun test --concurrent`
(the eval harness default). Three distinct shared-state races, all fixed:
1. Shared pidfile: a module-level `workDir` reassigned in beforeEach was
clobbered by parallel tests, so concurrent daemons collided on the same
pidfile and the loser returned `already_running`. Each test now gets its
own dir via makeWorkDir().
2. process.env path globals: tests set GSTACK_IOS_AUDIT_PATH /
_ATTEMPTS_PATH / _ALLOWLIST_PATH on the shared process env; concurrent
tests stomped each other's audit/attempts destinations. Threaded
auditPath/attemptsPath/allowlistPath through DaemonOptions (and
mintForCaller) as explicit args — env is no longer load-bearing.
3. afterEach cleanup race: the per-test cleanup drained a shared dir array,
so the first test to finish deleted still-running tests' workDirs
mid-assertion. Moved to afterAll (cleans once, after all settle).
Verified: 5/5 clean full-suite runs at --max-concurrency 15 (was
intermittent); daemon unit suite 91/91; daemon source compiles. The paths
default to the env-derived locations when options are omitted, so the
production CLI path is unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(pty): pin spawned claude to EVALS model chain (default claude-sonnet-4-6)
launchClaudePty spawned the interactive `claude` TUI with no --model flag, so
the child inherited the operator's ~/.claude/settings.json model. On a
slow-thinking model that meant 5+ min of extended thinking on empty plan-mode
context, timing out the plan-mode smoke tests regardless of contention. Pin the
model via opts.model ?? EVALS_MODEL ?? 'claude-sonnet-4-6' — byte-identical to
session-runner.ts:144, so PTY and `claude -p` evals always agree.
Pushed before extraArgs (last flag wins, so a per-test --model still overrides).
Placement leaves the spawn region byte-stable for a clean merge with the
in-flight hermetic-env branch. Plumbed model through the three plan-skill
wrappers. Static-grep tripwires guard the pin, its fallback chain, the
before-extraArgs ordering, and all three wrapper forwards.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(pty): detect markdown bold-bullet prose AUQs (fixes office-hours smoke)
office-hours auto-mode renders its mode question as `- **Building a startup**`
markdown bullets (office-hours/SKILL.md.tmpl:102) with no letter/number marker.
isProseAUQVisible only matched `A)`-style lettered or `1.`-style numbered
options, so the question went undetected: the model surfaced it at ~2m19s
(well under the 300s budget) but the harness kept scoring the run "working"
off the spinner glyphs and timed out — a false timeout on a question that was
already on screen.
Add Pattern 3: when an interrogative line ('?') is present AND 3+ bold-bullet
markers (`- **`) appear in the 4KB tail, classify as a prose AUQ. Bold is the
discriminator vs incidental prose bullets; the line anchor is dropped (stripAnsi
can collapse option lines) and the existing `❯ 1.` cursor gate still defers to a
live native list. Wires through the existing classifyVisible 'asked' path and the
timeout high-water-mark, so office-hours now classifies 'asked' instead of
'timeout'. Five unit cases: the office-hours render passes; no-'?', <3-bullet,
plain-bullet, and native-cursor cases stay false.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(pty): detect stripAnsi-collapsed prose AUQs + judge spinner-precedence
The plan-eng/plan-design plan-mode + finding-floor smokes timed out even when
the skill HAD rendered a complete prose AskUserQuestion and was waiting: the PTY
strips cursor-positioning escapes, collapsing the option newlines/spaces so
"A) ..." arrives as "A(recommended)" / "-B:" and "Reply with A, B, or C" as
"ReplywithA,B,orC". Every line-anchored detector (Patterns 1-3) returns false on
those bytes, so proseAUQEverObserved never latched and the run timed out on a
question that was already on screen.
Add Pattern 4/5: a two-signal collapsed-form detector — a reply/recommendation
marker (space-insensitive "reply with [A-D]", "Recommendation:", or
"(recommended)") AND 2+ distinct A-D letters each punctuated by ) : or (. The
conjunction is what separates a real AUQ from incidental report prose; verified
true on the verbatim failing-run buffers where Patterns 1-3 return false.
Also fix the Haiku judge spinner bias: of 614 verdicts, 569 were 'working' and
95 of those noted a question was visible — Claude Code keeps the spinner
animating at an idle prose decision, so the judge coin-flipped. Add a precedence
override: when an option list AND a Recommendation/Reply instruction are both
visible, classify WAITING even with spinner glyphs. Kept the strict dual-signal
gate (never option-list-alone) so auto-decide-preserved doesn't flip.
5 unit tests pin the two-signal contract (2 true on real collapsed bytes, 3
false guards). 90 -> 95 pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(plan-review): ask-first scope gate for plan-eng + plan-design review
On an empty/cold invocation, plan-eng-review and plan-design-review would dive
straight into repo exploration (plan-eng) or a 7-pass mockup+audit (plan-design)
and only ask the user much later, if at all. plan-ceo-review already asks first
via an unconditional Step-0 gate and behaves well; these two did not.
Add a hard-STOP scope gate as the FIRST operational instruction in each skill
(above the design-doc check / pre-review audit / mockup defaults it explicitly
overrides): the first tool call must be AskUserQuestion confirming the review
target, before any git/Read/Grep/Glob/Bash or mockup generation. Under
--disallowedTools the options render as plain column-0 lettered prose with a
Recommendation + "Reply with A, B, or C" line so the answer is detectable.
This is correct cold-start UX (confirm what to review before grinding a full
review on nothing) and it is the product half of the plan-mode smoke fix; the
harness collapsed-form detector is the deterministic half that catches the ask
however it renders. Templates + regenerated SKILL.md (default variant).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(tiers): reclassify stochastic plan-eng/plan-design ask-first smokes as periodic
plan-eng-review and plan-design-review run a long explore/audit before their
first AskUserQuestion, so whether the plan-mode + finding-floor smokes reach a
terminal outcome within the 300s/600s budget depends on stochastic ask-first
compliance (measured ~50-67%/run even with the hardened gate). Per the
"non-deterministic -> periodic" tiering rule, move the four affected smokes
(plan-eng/plan-design review-plan-mode + finding-floor) to periodic.
The deterministic harness fix (collapsed-form detector + judge precedence) and
the ask-first gate lift these from always-failing to mostly-passing and are the
real product+harness improvements; periodic monitoring tracks the rate weekly
without blocking PRs on an LLM coin-flip. plan-ceo/plan-devex ask-first reliably
and stay gate-tier.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(evals): gate the deterministic PTY plan-mode smokes in CI
The real-PTY plan-mode smokes never ran in CI — the gate was local-only. Add an
e2e-pty-plan-smoke matrix suite running the two deterministically-reliable ones
(office-hours-auto-mode, plan-mode-no-op) so a regression there blocks PRs. The
stochastic plan-eng/plan-design ask-first smokes stay periodic (touchfiles
E2E_TIERS) and are not CI-gated.
A fresh CI container has no ~/.claude.json, so the spawned interactive `claude`
would wedge on the onboarding + API-key-approval dialog. Add a scoped seed step
(hasCompletedOnboarding + key approval, its own ANTHROPIC_API_KEY env) before the
run — mirrors what the hermetic E2E child env seeds. Per-suite timeout override
(35 min) via matrix.suite.timeout so the PTY suite has headroom for --retry 2
without bumping the other 12 suites. Report runner count 12 -> 13.
Validate via workflow_dispatch before relying on the gate (PTY-in-CI is new).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(evals): install gstack skill registry for the PTY smoke suite
The first dry-run of e2e-pty-plan-smoke failed: the spawned interactive `claude`
printed "Unknown command: /plan-ceo-review". .claude/skills is gitignored, so a
fresh CI checkout has no gstack skill registry and the TUI can't resolve
/office-hours or /plan-ceo-review.
Add a Register step (scoped to the suite, after Seed, before Run) that mirrors
setup's --no-prefix user-scoped registry minimally: $HOME/.claude/skills/gstack
-> repo (resolves the preambles' absolute ~/.claude/skills/gstack/bin/* and
<skill>/sections/* paths) + per-skill SKILL.md/sections symlinks for the two
skills these tests invoke. HOME is /github/home in this container and the runner
adds no HOME/CLAUDE_CONFIG_DIR override (no hermetic mode), so $HOME is the right
anchor — the Seed step already proved claude reads it. No ./setup (binary build
+ Chromium + fonts + /dev/tty prompt); SKILL.md + bin/ + sections/ are committed.
Self-validating: fails the step loudly on a dangling symlink or missing
`name:` frontmatter, so a moved target surfaces here instead of as a silent
35-min "Unknown command" timeout.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: bump version and changelog (v1.58.4.0)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>