gstack

History

Garry Tan 76803d789a feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1) Adds comprehensive eval infrastructure: - Tier 1 (free): 13 new static tests — cross-skill path consistency, QA structure validation, greptile format, planted-bug fixture validation - Tier 2 (Agent SDK E2E): /qa quick, /review with pre-built git repo, 3 planted-bug outcome evals (static, SPA, checkout — each with 5 bugs) - Tier 3 (LLM judge): QA workflow quality, health rubric clarity, cross-skill consistency, baseline score pinning New fixtures: 3 HTML pages with 15 total planted bugs, ground truth JSON, review-eval-vuln.rb, eval-baselines.json. Shared llm-judge.ts helper (DRY). Unified EVALS=1 flag replaces SKILL_E2E + ANTHROPIC_API_KEY checks. `bun run test:evals` runs everything that costs money (~$4/run). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>		2026-03-14 01:17:36 -05:00
..
basic.html	Initial release — gstack v0.0.1	2026-03-12 01:32:16 -07:00
cursor-interactive.html	feat: Phase 3.5 — cookie import, QA testing, team retro (v0.3.1) (#29 )	2026-03-13 00:31:41 -07:00
dialog.html	feat: Phase 3.5 — cookie import, QA testing, team retro (v0.3.1) (#29 )	2026-03-13 00:31:41 -07:00
empty.html	feat: Phase 3.5 — cookie import, QA testing, team retro (v0.3.1) (#29 )	2026-03-13 00:31:41 -07:00
forms.html	Initial release — gstack v0.0.1	2026-03-12 01:32:16 -07:00
qa-eval-checkout.html	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
qa-eval-spa.html	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
qa-eval.html	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
responsive.html	Initial release — gstack v0.0.1	2026-03-12 01:32:16 -07:00
snapshot.html	Initial release — gstack v0.0.1	2026-03-12 01:32:16 -07:00
spa.html	Initial release — gstack v0.0.1	2026-03-12 01:32:16 -07:00
states.html	feat: Phase 3.5 — cookie import, QA testing, team retro (v0.3.1) (#29 )	2026-03-13 00:31:41 -07:00
upload.html	feat: Phase 3.5 — cookie import, QA testing, team retro (v0.3.1) (#29 )	2026-03-13 00:31:41 -07:00