gstack/test
Garry Tan f1581e6ff7
chore: upgrade eval judge to Sonnet 4.6, update changelog
Switch LLM-as-judge evals from Haiku to Sonnet 4.6 for more stable,
nuanced scoring. Add changelog entry for all eval improvements.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 00:12:48 -05:00
..
helpers feat: SKILL.md template system, 3-tier testing, DX tools (v0.3.3) (#41) 2026-03-13 21:08:12 -07:00
gen-skill-docs.test.ts test: add usage consistency and pipe guard tests 2026-03-14 00:12:42 -05:00
skill-e2e.test.ts feat: SKILL.md template system, 3-tier testing, DX tools (v0.3.3) (#41) 2026-03-13 21:08:12 -07:00
skill-llm-eval.test.ts chore: upgrade eval judge to Sonnet 4.6, update changelog 2026-03-14 00:12:48 -05:00
skill-parser.test.ts feat: SKILL.md template system, 3-tier testing, DX tools (v0.3.3) (#41) 2026-03-13 21:08:12 -07:00
skill-validation.test.ts test: add usage consistency and pipe guard tests 2026-03-14 00:12:42 -05:00