gstack

History

Garry Tan d646bc12d8 test(opus-4.7): rewrite scratch-root helper + add afterAll cleanup First run of the Opus 4.7 eval exposed two test-setup gaps that made results misleading: - Only the root gstack SKILL.md was installed. Claude Code does auto-discovery per-directory under .claude/skills/{name}/SKILL.md, so without individual skill dirs the Skill tool had nothing to route to. Positive routing cases all failed. - `claude -p` does not load SKILL.md content as system context the way the Claude Code harness does. The overlay nudges in SKILL.md were invisible to the model, so the fanout A/B could not actually differ. New `mkEvalRoot(suffix, includeOverlay)` helper, modelled on the pattern in skill-routing-e2e.test.ts: - Installs per-skill SKILL.md under .claude/skills/ for ~14 key skills so the Skill tool has discoverable targets. - Writes an explicit routing block into project CLAUDE.md. - When includeOverlay is true, inlines the content of model-overlays/opus-4-7.md into CLAUDE.md too. This is what makes the fanout A/B observable in `claude -p`: arm ON gets the overlay in context, arm OFF does not. Plus an afterAll that re-runs gen-skill-docs at the default model so the working tree is not left with opus-4-7-generated SKILL.md files after the eval finishes (would break golden-file tests in the next `bun test` run otherwise). With this setup in place: routing went from 3/3 FAIL to 3/3 PASS (correct skill or clarification in every positive case, zero false positives on negatives). Fanout A/B is now a fair comparison; still shows 0 parallel in both arms under `claude -p` (tracked as a P0 TODO for re-measurement inside Claude Code's harness, where fanout may land differently). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>		2026-04-22 00:27:53 -07:00
..
fixtures	chore(opus-4.7): regenerate SKILL.md files + update golden fixtures	2026-04-21 23:40:07 -07:00
helpers	test(opus-4.7): E2E eval for fanout rate + routing precision	2026-04-22 00:11:38 -07:00
analytics.test.ts	feat: safety hook skills + skill usage telemetry (v0.7.1) (#189 )	2026-03-18 23:57:59 -05:00
audit-compliance.test.ts	feat(v1.3.0.0): open agents learnings + cross-model benchmark skill (#1040 )	2026-04-19 17:50:31 +08:00
benchmark-cli.test.ts	feat(v1.3.0.0): open agents learnings + cross-model benchmark skill (#1040 )	2026-04-19 17:50:31 +08:00
benchmark-runner.test.ts	feat(v1.3.0.0): open agents learnings + cross-model benchmark skill (#1040 )	2026-04-19 17:50:31 +08:00
builder-profile.test.ts	feat: relationship closing — office-hours adapts to repeat users (v0.16.2.0) (#937 )	2026-04-08 22:21:28 -10:00
codex-e2e.test.ts	feat: worktree isolation for E2E tests + infrastructure elegance (v0.11.12.0) (#425 )	2026-03-23 23:05:22 -07:00
codex-hardening.test.ts	codex + Apple Silicon hardening wave (v0.18.4.0) (#1056 )	2026-04-18 12:30:54 +08:00
context-save-hardening.test.ts	fix(checkpoint): rename /checkpoint → /context-save + /context-restore (v1.0.1.0) (#1064 )	2026-04-19 08:38:19 +08:00
diff-scope.test.ts	feat: Review Army — parallel specialist reviewers for /review (v0.14.3.0) (#692 )	2026-03-30 22:07:50 -06:00
explain-level-config.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
gemini-e2e.test.ts	feat: Confusion Protocol, Hermes + GBrain hosts, brain-first resolver (v0.18.0.0) (#1005 )	2026-04-16 10:41:38 -07:00
gen-skill-docs.test.ts	test(routing): assert slash-prefixed skills + new policy + current names	2026-04-21 23:41:31 -07:00
global-discover.test.ts	fix: close redundant PRs + friendly error on all design commands (v0.15.8.1) (#817 )	2026-04-05 02:02:06 -07:00
gstack-developer-profile.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
gstack-question-log.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
gstack-question-preference.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
hook-scripts.test.ts	feat: safety hook skills + skill usage telemetry (v0.7.1) (#189 )	2026-03-18 23:57:59 -05:00
host-config.test.ts	community wave: 6 PRs + hardening (v0.18.1.0) (#1028 )	2026-04-17 00:45:13 -07:00
jargon-list.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
learnings-injection.test.ts	fix: community security wave — 8 PRs, 4 contributors (v0.15.13.0) (#847 )	2026-04-06 00:47:04 -07:00
learnings.test.ts	feat: GStack Learns — per-project self-learning infrastructure (v0.13.4.0) (#622 )	2026-03-29 17:02:01 -06:00
migration-checkpoint-ownership.test.ts	fix(checkpoint): rename /checkpoint → /context-save + /context-restore (v1.0.1.0) (#1064 )	2026-04-19 08:38:19 +08:00
openclaw-native-skills.test.ts	community wave: 6 PRs + hardening (v0.18.1.0) (#1028 )	2026-04-17 00:45:13 -07:00
plan-tune.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
readme-throughput.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
relink.test.ts	fix: headed browser auto-shutdown + disconnect cleanup (v0.18.1.0) (#1025 )	2026-04-16 15:39:44 -07:00
review-log.test.ts	fix: community PRs + security hardening + E2E stability (v0.12.7.0) (#552 )	2026-03-26 23:21:27 -06:00
setup-codesign.test.ts	codex + Apple Silicon hardening wave (v0.18.4.0) (#1056 )	2026-04-18 12:30:54 +08:00
ship-version-sync.test.ts	fix(ship): detect + repair VERSION/package.json drift in Step 12 (v1.1.1.0) (#1063 )	2026-04-18 23:58:59 +08:00
skill-collision-sentinel.test.ts	fix(checkpoint): rename /checkpoint → /context-save + /context-restore (v1.0.1.0) (#1064 )	2026-04-19 08:38:19 +08:00
skill-e2e-autoplan-dual-voice.test.ts	fix(checkpoint): rename /checkpoint → /context-save + /context-restore (v1.0.1.0) (#1064 )	2026-04-19 08:38:19 +08:00
skill-e2e-benchmark-providers.test.ts	feat(v1.3.0.0): open agents learnings + cross-model benchmark skill (#1040 )	2026-04-19 17:50:31 +08:00
skill-e2e-bws.test.ts	fix: cookie picker auth token leak (v0.15.17.0) (#904 )	2026-04-08 10:10:13 -07:00
skill-e2e-context-skills.test.ts	fix(checkpoint): rename /checkpoint → /context-save + /context-restore (v1.0.1.0) (#1064 )	2026-04-19 08:38:19 +08:00
skill-e2e-cso.test.ts	feat: /cso v2 — infrastructure-first security audit (v0.11.6.0) (#384 )	2026-03-23 06:57:22 -07:00
skill-e2e-deploy.test.ts	feat: /land-and-deploy first-run dry run + staging-first + trust ladder (v0.12.2.0) (#518 )	2026-03-26 11:08:31 -07:00
skill-e2e-design.test.ts	feat: CI evals on Ubicloud — 12 parallel runners + Docker image (v0.11.10.0) (#360 )	2026-03-23 10:17:33 -07:00
skill-e2e-learnings.test.ts	feat: recursive self-improvement — operational learning + full skill wiring (v0.13.8.0) (#647 )	2026-03-31 23:08:22 -06:00
skill-e2e-office-hours.test.ts	feat: mode-posture energy fix for /plan-ceo-review and /office-hours (v1.1.2.0) (#1065 )	2026-04-19 05:44:39 +08:00
skill-e2e-opus-47.test.ts	test(opus-4.7): rewrite scratch-root helper + add afterAll cleanup	2026-04-22 00:27:53 -07:00
skill-e2e-plan-tune.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
skill-e2e-plan.test.ts	feat: mode-posture energy fix for /plan-ceo-review and /office-hours (v1.1.2.0) (#1065 )	2026-04-19 05:44:39 +08:00
skill-e2e-qa-bugs.test.ts	feat: CI evals on Ubicloud — 12 parallel runners + Docker image (v0.11.10.0) (#360 )	2026-03-23 10:17:33 -07:00
skill-e2e-qa-workflow.test.ts	feat: CI evals on Ubicloud — 12 parallel runners + Docker image (v0.11.10.0) (#360 )	2026-03-23 10:17:33 -07:00
skill-e2e-review-army.test.ts	feat: Review Army — parallel specialist reviewers for /review (v0.14.3.0) (#692 )	2026-03-30 22:07:50 -06:00
skill-e2e-review.test.ts	feat: Confusion Protocol, Hermes + GBrain hosts, brain-first resolver (v0.18.0.0) (#1005 )	2026-04-16 10:41:38 -07:00
skill-e2e-session-intelligence.test.ts	fix(checkpoint): rename /checkpoint → /context-save + /context-restore (v1.0.1.0) (#1064 )	2026-04-19 08:38:19 +08:00
skill-e2e-sidebar.test.ts	feat: declarative multi-host platform + OpenCode, Slate, Cursor, OpenClaw (v0.15.5.0) (#793 )	2026-04-04 15:32:20 -07:00
skill-e2e-workflow.test.ts	refactor: extract TabSession for per-tab state isolation (v0.15.16.0) (#873 )	2026-04-07 00:23:36 -07:00
skill-e2e.test.ts	feat: recursive self-improvement — operational learning + full skill wiring (v0.13.8.0) (#647 )	2026-03-31 23:08:22 -06:00
skill-llm-eval.test.ts	feat: voice directive for all skills (v0.12.3.0) (#520 )	2026-03-26 17:31:53 -06:00
skill-parser.test.ts	feat: SKILL.md template system, 3-tier testing, DX tools (v0.3.3) (#41 )	2026-03-13 21:08:12 -07:00
skill-routing-e2e.test.ts	feat: Confusion Protocol, Hermes + GBrain hosts, brain-first resolver (v0.18.0.0) (#1005 )	2026-04-16 10:41:38 -07:00
skill-validation.test.ts	test(binary-guard): replace xargs-per-file loops with fs.statSync + mode filter	2026-04-21 23:43:43 -07:00
taste-engine.test.ts	feat(v1.3.0.0): open agents learnings + cross-model benchmark skill (#1040 )	2026-04-19 17:50:31 +08:00
team-mode.test.ts	test(team-mode): give setup -q / setup --local tests a 3-minute budget	2026-04-21 23:48:48 -07:00
telemetry.test.ts	feat: community wave — 7 fixes, relink, sidebar Write, discoverability (v0.13.5.0) (#641 )	2026-03-29 21:43:36 -06:00
timeline.test.ts	feat: Session Intelligence Layer — /checkpoint + /health + context recovery (v0.15.0.0) (#733 )	2026-04-01 00:50:42 -06:00
touchfiles.test.ts	feat: mode-posture energy fix for /plan-ceo-review and /office-hours (v1.1.2.0) (#1065 )	2026-04-19 05:44:39 +08:00
uninstall.test.ts	feat: community PRs — faster install, skill namespacing, uninstall, Codex fallback, Windows fix, Python patterns (v0.12.9.0) (#561 )	2026-03-27 00:44:37 -06:00
upgrade-migration-v1.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
v0-dormancy.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00
worktree.test.ts	feat: content security — 4-layer prompt injection defense for pair-agent (#815 )	2026-04-06 14:41:06 -07:00
writing-style-resolver.test.ts	feat: gstack v1 — simpler prompts + real LOC receipts (v1.0.0.0) (#1039 )	2026-04-18 15:05:42 +08:00