gstack

History

Garry Tan f9c834f1f9 feat: add E2E eval for session awareness ELI16 mode Stubs _SESSIONS=4, gives agent a decision point on feature/add-payments branch, verifies the output re-grounds the user with project, branch, context, and RECOMMENDATION — the ELI16 mode behavior for 3+ sessions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>		2026-03-16 01:19:35 -05:00
..
fixtures	feat: add evals for RECOMMENDATION format, session awareness, and enum completeness	2026-03-16 01:15:45 -05:00
helpers	feat: QA restructure, browser ref staleness, eval efficiency metrics (v0.4.0) (#83 )	2026-03-15 23:55:39 -05:00
gen-skill-docs.test.ts	feat: contributor mode, session awareness, universal RECOMMENDATION format	2026-03-16 00:32:27 -05:00
skill-e2e.test.ts	feat: add E2E eval for session awareness ELI16 mode	2026-03-16 01:19:35 -05:00
skill-llm-eval.test.ts	fix: lower planted-bug detection baselines and LLM judge thresholds for reliability	2026-03-14 05:16:17 -05:00
skill-parser.test.ts	feat: SKILL.md template system, 3-tier testing, DX tools (v0.3.3) (#41 )	2026-03-13 21:08:12 -07:00
skill-validation.test.ts	feat: add evals for RECOMMENDATION format, session awareness, and enum completeness	2026-03-16 01:15:45 -05:00