mirror of https://github.com/garrytan/gstack.git
docs: sync docs for v1.66.0.0 (test/evals/CI speedup)
CONTRIBUTING.md, AGENTS.md, and ARCHITECTURE.md still taught bare `bun test` for the suite; the shipped runner deprecates it (walks the whole repo, loads paid eval files, misses the strict classifier). All suite-level references now say `bun run test`, the Tier 1 section describes the strict shard runner (~90-100s, --verbose, --wall-timeout), the sharded paid-runner paragraph documents diff-based shard skipping and the EVALS_JOBS / EVALS_CONCURRENCY split, the Tier 3 row points at the actual judge-only invocation, and GSTACK_EVAL_MODEL_JUDGE is documented at the judge it overrides. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
f627c676fc
commit
9b9bc84bdc
|
|
@ -111,7 +111,7 @@ End-to-end walkthrough: [docs/howto-ios-testing-with-gstack.md](docs/howto-ios-t
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bun install # install dependencies
|
bun install # install dependencies
|
||||||
bun test # run free tests (no API spend)
|
bun run test # run free tests via the strict shard runner (no API spend, ~90-100s)
|
||||||
bun run test:windows # curated Windows-safe subset (runs on windows-latest)
|
bun run test:windows # curated Windows-safe subset (runs on windows-latest)
|
||||||
bun run build # generate docs + compile binaries
|
bun run build # generate docs + compile binaries
|
||||||
bun run gen:skill-docs # regenerate SKILL.md files from templates
|
bun run gen:skill-docs # regenerate SKILL.md files from templates
|
||||||
|
|
|
||||||
|
|
@ -321,7 +321,7 @@ Three reasons:
|
||||||
| 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, check for errors | ~$3.85 | ~20min |
|
| 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, check for errors | ~$3.85 | ~20min |
|
||||||
| 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s |
|
| 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s |
|
||||||
|
|
||||||
Tier 1 runs on every `bun test`. Tiers 2+3 are gated behind `EVALS=1`. The idea is: catch 95% of issues for free, use LLMs only for judgment calls.
|
Tier 1 runs on every `bun run test`. Tiers 2+3 are gated behind `EVALS=1`. The idea is: catch 95% of issues for free, use LLMs only for judgment calls.
|
||||||
|
|
||||||
## Command dispatch
|
## Command dispatch
|
||||||
|
|
||||||
|
|
@ -435,7 +435,7 @@ The `EvalCollector` accumulates test results and writes them in two ways:
|
||||||
| 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, scan for errors | ~$3.85 | ~20min |
|
| 2 — E2E via `claude -p` | Spawn real Claude session, run each skill, scan for errors | ~$3.85 | ~20min |
|
||||||
| 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s |
|
| 3 — LLM-as-judge | Sonnet scores docs on clarity/completeness/actionability | ~$0.15 | ~30s |
|
||||||
|
|
||||||
Tier 1 runs on every `bun test`. Tiers 2+3 are gated behind `EVALS=1`. The idea: catch 95% of issues for free, use LLMs only for judgment calls and integration testing.
|
Tier 1 runs on every `bun run test`. Tiers 2+3 are gated behind `EVALS=1`. The idea: catch 95% of issues for free, use LLMs only for judgment calls and integration testing.
|
||||||
|
|
||||||
## What's intentionally not here
|
## What's intentionally not here
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -141,20 +141,26 @@ Bun auto-loads `.env` — no extra config. Conductor workspaces inherit `.env` f
|
||||||
|
|
||||||
| Tier | Command | Cost | What it tests |
|
| Tier | Command | Cost | What it tests |
|
||||||
|------|---------|------|---------------|
|
|------|---------|------|---------------|
|
||||||
| 1 — Static | `bun test` | Free | Command validation, snapshot flags, SKILL.md correctness, TODOS-format.md refs, observability unit tests |
|
| 1 — Static | `bun run test` | Free | Command validation, snapshot flags, SKILL.md correctness, TODOS-format.md refs, observability unit tests |
|
||||||
| 2 — E2E | `bun run test:e2e` | ~$3.85 | Full skill execution via `claude -p` subprocess |
|
| 2 — E2E | `bun run test:e2e` | ~$3.85 | Full skill execution via `claude -p` subprocess |
|
||||||
| 3 — LLM eval | `bun run test:evals` | ~$0.15 standalone | LLM-as-judge scoring of generated SKILL.md docs |
|
| 3 — LLM eval | `EVALS=1 bun test test/skill-llm-eval.test.ts` | ~$0.15 standalone | LLM-as-judge scoring of generated SKILL.md docs |
|
||||||
| 2+3 | `bun run test:evals` | ~$4 combined | E2E + LLM-as-judge (runs both) |
|
| 2+3 | `bun run test:evals` | ~$4 combined | E2E + LLM-as-judge (runs both) |
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bun test # Tier 1 only (runs on every commit, <5s)
|
bun run test # Tier 1 only (run before every commit, ~90-100s for the full ~7,000-test suite)
|
||||||
bun run test:e2e # Tier 2: E2E only (needs EVALS=1, can't run inside Claude Code)
|
bun run test:e2e # Tier 2: E2E only (needs EVALS=1, can't run inside Claude Code)
|
||||||
bun run test:evals # Tier 2 + 3 combined (~$4/run)
|
bun run test:evals # Tier 2 + 3 combined (~$4/run)
|
||||||
```
|
```
|
||||||
|
|
||||||
### Tier 1: Static validation (free)
|
### Tier 1: Static validation (free)
|
||||||
|
|
||||||
Runs automatically with `bun test`. No API keys needed.
|
Runs with `bun run test`, which routes through `scripts/test-free-shards.ts`: N
|
||||||
|
concurrent shard processes under a strict output contract — a shard that exits
|
||||||
|
without bun's own terminal summary line, or a crashed worker, fails the run, so
|
||||||
|
silent truncation can never report green. Pass `--verbose` to forward the full
|
||||||
|
child stream; `--wall-timeout <secs>` overrides the per-shard kill deadline.
|
||||||
|
Don't type bare `bun test` for the suite: it walks the whole repo, loads paid
|
||||||
|
eval files, and misses the strict classifier. No API keys needed.
|
||||||
|
|
||||||
- **Skill parser tests** (`test/skill-parser.test.ts`) — Extracts every `$B` command from SKILL.md bash code blocks and validates against the command registry in `browse/src/commands.ts`. Catches typos, removed commands, and invalid snapshot flags.
|
- **Skill parser tests** (`test/skill-parser.test.ts`) — Extracts every `$B` command from SKILL.md bash code blocks and validates against the command registry in `browse/src/commands.ts`. Catches typos, removed commands, and invalid snapshot flags.
|
||||||
- **Skill validation tests** (`test/skill-validation.test.ts`) — Validates that SKILL.md files reference only real commands and flags, and that command descriptions meet quality thresholds.
|
- **Skill validation tests** (`test/skill-validation.test.ts`) — Validates that SKILL.md files reference only real commands and flags, and that command descriptions meet quality thresholds.
|
||||||
|
|
@ -240,7 +246,12 @@ as `bun run test:gate:sharded` / `bun run test:periodic:sharded`): one Bun
|
||||||
process per test file, an external wall-clock timeout that kills the shard's
|
process per test file, an external wall-clock timeout that kills the shard's
|
||||||
whole process group (stray `claude`/`codex` grandchildren included), a per-shard
|
whole process group (stray `claude`/`codex` grandchildren included), a per-shard
|
||||||
eval dir (`GSTACK_EVAL_DIR=<evalDir>/shards/<slug>/`), and an aggregate that
|
eval dir (`GSTACK_EVAL_DIR=<evalDir>/shards/<slug>/`), and an aggregate that
|
||||||
distinguishes failed vs timed-out vs never-started shards. `eval:list`,
|
distinguishes failed vs timed-out vs never-started shards. The runner also
|
||||||
|
selects by diff: shards untouched by your branch are reported as
|
||||||
|
skipped-by-diff, with a selection banner naming the reason (`EVALS_ALL=1`
|
||||||
|
forces everything). `EVALS_JOBS` sets how many shard processes run at once
|
||||||
|
(default 4); `EVALS_CONCURRENCY` is bun's concurrency WITHIN a shard — they
|
||||||
|
are deliberately separate knobs. `eval:list`,
|
||||||
`eval:compare`, and `eval:summary` are shard-aware. Humans running
|
`eval:compare`, and `eval:summary` are shard-aware. Humans running
|
||||||
`bun run test:evals` foreground in their own terminal don't need this — Ctrl-C
|
`bun run test:evals` foreground in their own terminal don't need this — Ctrl-C
|
||||||
is intended there.
|
is intended there.
|
||||||
|
|
@ -251,7 +262,8 @@ Artifacts are never cleaned up — they accumulate in `~/.gstack-dev/` for post-
|
||||||
|
|
||||||
### Tier 3: LLM-as-judge (~$0.15/run)
|
### Tier 3: LLM-as-judge (~$0.15/run)
|
||||||
|
|
||||||
Uses Claude Sonnet to score generated SKILL.md docs on three dimensions:
|
Uses Claude Sonnet to score generated SKILL.md docs on three dimensions.
|
||||||
|
Override the judge model per run with `GSTACK_EVAL_MODEL_JUDGE`:
|
||||||
|
|
||||||
- **Clarity** — Can an AI agent understand the instructions without ambiguity?
|
- **Clarity** — Can an AI agent understand the instructions without ambiguity?
|
||||||
- **Completeness** — Are all commands, flags, and usage patterns documented?
|
- **Completeness** — Are all commands, flags, and usage patterns documented?
|
||||||
|
|
@ -371,7 +383,7 @@ See `scripts/host-config.ts` for the full `HostConfig` interface.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Run all static tests (includes parameterized smoke tests for all hosts)
|
# Run all static tests (includes parameterized smoke tests for all hosts)
|
||||||
bun test
|
bun run test
|
||||||
|
|
||||||
# Check freshness for all hosts
|
# Check freshness for all hosts
|
||||||
bun run gen:skill-docs --host all --dry-run
|
bun run gen:skill-docs --host all --dry-run
|
||||||
|
|
@ -388,7 +400,7 @@ See [docs/ADDING_A_HOST.md](docs/ADDING_A_HOST.md) for the full guide. Short ver
|
||||||
2. Add to `hosts/index.ts`
|
2. Add to `hosts/index.ts`
|
||||||
3. Add `.myhost/` to `.gitignore`
|
3. Add `.myhost/` to `.gitignore`
|
||||||
4. Run `bun run gen:skill-docs --host myhost`
|
4. Run `bun run gen:skill-docs --host myhost`
|
||||||
5. Run `bun test` (parameterized tests auto-cover it)
|
5. Run `bun run test` (parameterized tests auto-cover it)
|
||||||
|
|
||||||
Zero generator, setup, or tooling code changes needed.
|
Zero generator, setup, or tooling code changes needed.
|
||||||
|
|
||||||
|
|
@ -503,7 +515,7 @@ When community PRs accumulate, batch them into themed waves:
|
||||||
2. **Deduplicate** — if two PRs fix the same thing, pick the one that
|
2. **Deduplicate** — if two PRs fix the same thing, pick the one that
|
||||||
changes fewer lines. Close the other with a note pointing to the winner.
|
changes fewer lines. Close the other with a note pointing to the winner.
|
||||||
3. **Collector branch** — create `pr-wave-N`, merge clean PRs, resolve
|
3. **Collector branch** — create `pr-wave-N`, merge clean PRs, resolve
|
||||||
conflicts for dirty ones, verify with `bun test && bun run build`
|
conflicts for dirty ones, verify with `bun run test && bun run build`
|
||||||
4. **Close with context** — every closed PR gets a comment explaining
|
4. **Close with context** — every closed PR gets a comment explaining
|
||||||
why and what (if anything) supersedes it. Contributors did real work;
|
why and what (if anything) supersedes it. Contributors did real work;
|
||||||
respect that with clear communication.
|
respect that with clear communication.
|
||||||
|
|
@ -558,7 +570,7 @@ Failures are logged but never block the upgrade.
|
||||||
|
|
||||||
### Testing migrations
|
### Testing migrations
|
||||||
|
|
||||||
Migrations are tested as part of `bun test` (tier 1, free). The test suite
|
Migrations are tested as part of `bun run test` (tier 1, free). The test suite
|
||||||
verifies that all migration scripts in `gstack-upgrade/migrations/` are
|
verifies that all migration scripts in `gstack-upgrade/migrations/` are
|
||||||
executable and parse without syntax errors.
|
executable and parse without syntax errors.
|
||||||
|
|
||||||
|
|
|
||||||
Loading…
Reference in New Issue