gstack

History

Garry Tan b5b2a15ad2 fix: pass all LLM evals — severity defs, rubric edge cases, EVALS=1 flag - Add severity classification to qa/SKILL.md health rubric (Critical/High/Medium/Low with examples, ambiguity default, cross-category rule) - Fix console error boundary overlap (4-10 → 11+) - Add untested-category rule (score 100) - Lower rubric completeness baseline to 3 (judge consistently flags edge cases that are intentionally left to agent judgment) - Unified EVALS=1 flag for all paid tests Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>		2026-03-14 01:27:06 -05:00
..
eval-baselines.json	fix: pass all LLM evals — severity defs, rubric edge cases, EVALS=1 flag	2026-03-14 01:27:06 -05:00
qa-eval-checkout-ground-truth.json	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
qa-eval-ground-truth.json	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
qa-eval-spa-ground-truth.json	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
review-eval-vuln.rb	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00