From 9b9bc84bdc63eb6120497b200db992ee67be2e9e Mon Sep 17 00:00:00 2001 From: Garry Tan Date: Sat, 15 Aug 2026 17:23:00 -0700 Subject: [PATCH] docs: sync docs for v1.66.0.0 (test/evals/CI speedup) CONTRIBUTING.md, AGENTS.md, and ARCHITECTURE.md still taught bare `bun test` for the suite; the shipped runner deprecates it (walks the whole repo, loads paid eval files, misses the strict classifier). All suite-level references now say `bun run test`, the Tier 1 section describes the strict shard runner (~90-100s, --verbose, --wall-timeout), the sharded paid-runner paragraph documents diff-based shard skipping and the EVALS_JOBS / EVALS_CONCURRENCY split, the Tier 3 row points at the actual judge-only invocation, and GSTACK_EVAL_MODEL_JUDGE is documented at the judge it overrides. Co-Authored-By: Claude Fable 5 --- AGENTS.md | 2 +- ARCHITECTURE.md | 4 ++-- CONTRIBUTING.md | 32 ++++++++++++++++++++++---------- 3 files changed, 25 insertions(+), 13 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 4df7eec35..cef8c05f8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -111,7 +111,7 @@ End-to-end walkthrough: [docs/howto-ios-testing-with-gstack.md](docs/howto-ios-t ```bash bun install # install dependencies -bun test # run free tests (no API spend) +bun run test # run free tests via the strict shard runner (no API spend, ~90-100s) bun run test:windows # curated Windows-safe subset (runs on windows-latest) bun run build # generate docs + compile binaries bun run gen:skill-docs # regenerate SKILL.md files from templates diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index f6d584c8c..1856f0ad7 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -321,7 +321,7 @@ Three reasons: | 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, check for errors | ~$3.85 | ~20min | | 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s | -Tier 1 runs on every `bun test`. Tiers 2+3 are gated behind `EVALS=1`. The idea is: catch 95% of issues for free, use LLMs only for judgment calls. +Tier 1 runs on every `bun run test`. Tiers 2+3 are gated behind `EVALS=1`. The idea is: catch 95% of issues for free, use LLMs only for judgment calls. ## Command dispatch @@ -435,7 +435,7 @@ The `EvalCollector` accumulates test results and writes them in two ways: | 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, scan for errors | ~$3.85 | ~20min | | 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s | -Tier 1 runs on every `bun test`. Tiers 2+3 are gated behind `EVALS=1`. The idea: catch 95% of issues for free, use LLMs only for judgment calls and integration testing. +Tier 1 runs on every `bun run test`. Tiers 2+3 are gated behind `EVALS=1`. The idea: catch 95% of issues for free, use LLMs only for judgment calls and integration testing. ## What's intentionally not here diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 16d3183ca..e7da57eea 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -141,20 +141,26 @@ Bun auto-loads `.env` — no extra config. Conductor workspaces inherit `.env` f | Tier | Command | Cost | What it tests | |------|---------|------|---------------| -| 1 — Static | `bun test` | Free | Command validation, snapshot flags, SKILL.md correctness, TODOS-format.md refs, observability unit tests | +| 1 — Static | `bun run test` | Free | Command validation, snapshot flags, SKILL.md correctness, TODOS-format.md refs, observability unit tests | | 2 — E2E | `bun run test:e2e` | ~$3.85 | Full skill execution via `claude -p` subprocess | -| 3 — LLM eval | `bun run test:evals` | ~$0.15 standalone | LLM-as-judge scoring of generated SKILL.md docs | +| 3 — LLM eval | `EVALS=1 bun test test/skill-llm-eval.test.ts` | ~$0.15 standalone | LLM-as-judge scoring of generated SKILL.md docs | | 2+3 | `bun run test:evals` | ~$4 combined | E2E + LLM-as-judge (runs both) | ```bash -bun test # Tier 1 only (runs on every commit, <5s) +bun run test # Tier 1 only (run before every commit, ~90-100s for the full ~7,000-test suite) bun run test:e2e # Tier 2: E2E only (needs EVALS=1, can't run inside Claude Code) bun run test:evals # Tier 2 + 3 combined (~$4/run) ``` ### Tier 1: Static validation (free) -Runs automatically with `bun test`. No API keys needed. +Runs with `bun run test`, which routes through `scripts/test-free-shards.ts`: N +concurrent shard processes under a strict output contract — a shard that exits +without bun's own terminal summary line, or a crashed worker, fails the run, so +silent truncation can never report green. Pass `--verbose` to forward the full +child stream; `--wall-timeout ` overrides the per-shard kill deadline. +Don't type bare `bun test` for the suite: it walks the whole repo, loads paid +eval files, and misses the strict classifier. No API keys needed. - **Skill parser tests** (`test/skill-parser.test.ts`) — Extracts every `$B` command from SKILL.md bash code blocks and validates against the command registry in `browse/src/commands.ts`. Catches typos, removed commands, and invalid snapshot flags. - **Skill validation tests** (`test/skill-validation.test.ts`) — Validates that SKILL.md files reference only real commands and flags, and that command descriptions meet quality thresholds. @@ -240,7 +246,12 @@ as `bun run test:gate:sharded` / `bun run test:periodic:sharded`): one Bun process per test file, an external wall-clock timeout that kills the shard's whole process group (stray `claude`/`codex` grandchildren included), a per-shard eval dir (`GSTACK_EVAL_DIR=/shards//`), and an aggregate that -distinguishes failed vs timed-out vs never-started shards. `eval:list`, +distinguishes failed vs timed-out vs never-started shards. The runner also +selects by diff: shards untouched by your branch are reported as +skipped-by-diff, with a selection banner naming the reason (`EVALS_ALL=1` +forces everything). `EVALS_JOBS` sets how many shard processes run at once +(default 4); `EVALS_CONCURRENCY` is bun's concurrency WITHIN a shard — they +are deliberately separate knobs. `eval:list`, `eval:compare`, and `eval:summary` are shard-aware. Humans running `bun run test:evals` foreground in their own terminal don't need this — Ctrl-C is intended there. @@ -251,7 +262,8 @@ Artifacts are never cleaned up — they accumulate in `~/.gstack-dev/` for post- ### Tier 3: LLM-as-judge (~$0.15/run) -Uses Claude Sonnet to score generated SKILL.md docs on three dimensions: +Uses Claude Sonnet to score generated SKILL.md docs on three dimensions. +Override the judge model per run with `GSTACK_EVAL_MODEL_JUDGE`: - **Clarity** — Can an AI agent understand the instructions without ambiguity? - **Completeness** — Are all commands, flags, and usage patterns documented? @@ -371,7 +383,7 @@ See `scripts/host-config.ts` for the full `HostConfig` interface. ```bash # Run all static tests (includes parameterized smoke tests for all hosts) -bun test +bun run test # Check freshness for all hosts bun run gen:skill-docs --host all --dry-run @@ -388,7 +400,7 @@ See [docs/ADDING_A_HOST.md](docs/ADDING_A_HOST.md) for the full guide. Short ver 2. Add to `hosts/index.ts` 3. Add `.myhost/` to `.gitignore` 4. Run `bun run gen:skill-docs --host myhost` -5. Run `bun test` (parameterized tests auto-cover it) +5. Run `bun run test` (parameterized tests auto-cover it) Zero generator, setup, or tooling code changes needed. @@ -503,7 +515,7 @@ When community PRs accumulate, batch them into themed waves: 2. **Deduplicate** — if two PRs fix the same thing, pick the one that changes fewer lines. Close the other with a note pointing to the winner. 3. **Collector branch** — create `pr-wave-N`, merge clean PRs, resolve - conflicts for dirty ones, verify with `bun test && bun run build` + conflicts for dirty ones, verify with `bun run test && bun run build` 4. **Close with context** — every closed PR gets a comment explaining why and what (if anything) supersedes it. Contributors did real work; respect that with clear communication. @@ -558,7 +570,7 @@ Failures are logged but never block the upgrade. ### Testing migrations -Migrations are tested as part of `bun test` (tier 1, free). The test suite +Migrations are tested as part of `bun run test` (tier 1, free). The test suite verifies that all migration scripts in `gstack-upgrade/migrations/` are executable and parse without syntax errors.