gstack

History

Garry Tan a164597847 build(skills): T7 — atomic regenerate + capture v1.45.0.0 baseline Final regen pass across all hosts after T1-T6 work landed. Captures the v1.45.0.0 parity baseline at test/fixtures/parity-baseline-v1.45.0.0.json for diffing against the v1.44.1 reference. Measured deltas (real numbers from test/helpers/capture-parity-baseline.ts): Total SKILL.md corpus 2,847 KB → 2,813 KB (-1.2%) Catalog tokens (always-loaded) ~9,319 → ~4,045 tokens (-56.6%) Top 10 heaviest skills 0.5-1.0% drop each The catalog token cut is the headline. It's the always-loaded surface, i.e. tokens charged on every session start. Per-skill SKILL.md sizes barely moved because T4 catalog trim MOVES routing prose from frontmatter to a body "## When to invoke" section rather than deleting it — the catalog wins without amputating discoverability. The bigger per-skill compression lands in v2.0.0.0 (Phase B sections/ pattern on the 5 heavyweights). v1.45 is the foundation: eval-first infrastructure + cheap wins. scripts/proactive-suggestions.json regenerated with the latest 52 skills listed (one-time write per gen-skill-docs run; aggregated catalog parts). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>		2026-05-25 20:38:52 -07:00
..
golden	v1.43.2.0 fix wave: post-Daegu paper-cut — 18 fixes, 28 bisect commits (#1642 )	2026-05-21 21:21:07 -07:00
ios-qa/FixtureApp	v1.43.0.0 feat: iOS device-farm (5 skills, Mac daemon, Tailscale) (#1574 )	2026-05-21 16:09:26 -07:00
mode-posture	feat: mode-posture energy fix for /plan-ceo-review and /office-hours (v1.1.2.0) (#1065 )	2026-04-19 05:44:39 +08:00
plans	v1.15.0.0 feat: slim preamble + real-PTY plan-mode E2E harness (#1215 )	2026-04-26 13:55:13 -07:00
coverage-audit-fixture.ts	feat: test coverage catalog — shared audit across plan/ship/review (v0.10.1.0) (#259 )	2026-03-22 11:28:16 -07:00
eval-baselines.json	fix: rewrite session-runner to claude -p subprocess, lower flaky baselines	2026-03-14 02:34:10 -05:00
forcing-finding-seeds.ts	v1.31.0.0 fix: delete AskUserQuestion fallback (root cause of forever war) + harness primitives (#1390 )	2026-05-09 17:01:13 -07:00
golden-ship-claude.md	fix: community security wave — 8 PRs, 4 contributors (v0.15.13.0) (#847 )	2026-04-06 00:47:04 -07:00
overlay-nudges.ts	feat(v1.10.1.0): overlay efficacy harness + Opus 4.7 fanout nudge removal (#1166 )	2026-04-23 18:42:58 -07:00
parity-baseline-v1.44.1.json	test(parity): T0a — capture v1.44.1 baseline + capture helper + diff utility	2026-05-25 20:29:47 -07:00
parity-baseline-v1.45.0.0.json	build(skills): T7 — atomic regenerate + capture v1.45.0.0 baseline	2026-05-25 20:38:52 -07:00
qa-eval-checkout-ground-truth.json	fix: 100% E2E pass — isolate test dirs, restart server, relax FP thresholds	2026-03-14 07:17:17 -05:00
qa-eval-ground-truth.json	fix: 100% E2E pass — isolate test dirs, restart server, relax FP thresholds	2026-03-14 07:17:17 -05:00
qa-eval-spa-ground-truth.json	fix: 100% E2E pass — isolate test dirs, restart server, relax FP thresholds	2026-03-14 07:17:17 -05:00
review-army-migration.sql	feat: Review Army — parallel specialist reviewers for /review (v0.14.3.0) (#692 )	2026-03-30 22:07:50 -06:00
review-army-n-plus-one.rb	feat: Review Army — parallel specialist reviewers for /review (v0.14.3.0) (#692 )	2026-03-30 22:07:50 -06:00
review-eval-design-slop.css	feat: design review lite in /review and /ship + gstack-diff-scope (v0.6.3) (#142 )	2026-03-17 20:12:55 -05:00
review-eval-design-slop.html	feat: design review lite in /review and /ship + gstack-diff-scope (v0.6.3) (#142 )	2026-03-17 20:12:55 -05:00
review-eval-enum-diff.rb	feat: contributor mode, session awareness, recommendation format (#90 )	2026-03-16 01:45:50 -05:00
review-eval-enum.rb	feat: contributor mode, session awareness, recommendation format (#90 )	2026-03-16 01:45:50 -05:00
review-eval-vuln.rb	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00