gstack

History

Garry Tan 95d9116003 fix: Codex E2E test now validates all skills load without warnings - Install ALL skills to temp HOME (not just one) to catch missing SKILL.md - Pre-flight asserts every .agents/ dir has both SKILL.md and openai.yaml - Assert no "invalid SKILL.md" or "Skipped loading" in stderr - Add existingHome option to runCodexSkill for pre-populated temp HOMEs - Increase discover test timeout to 120s (all-skills load takes longer)		2026-03-23 09:14:17 -07:00
..
codex-session-runner.ts	fix: Codex E2E test now validates all skills load without warnings	2026-03-23 09:14:17 -07:00
e2e-helpers.ts	feat: /land-and-deploy, /canary, /benchmark + perf review (v0.7.0) (#183 )	2026-03-21 14:31:36 -07:00
eval-store.test.ts	feat: QA restructure, browser ref staleness, eval efficiency metrics (v0.4.0) (#83 )	2026-03-15 23:55:39 -05:00
eval-store.ts	feat: /land-and-deploy, /canary, /benchmark + perf review (v0.7.0) (#183 )	2026-03-21 14:31:36 -07:00
gemini-session-runner.test.ts	feat: Gemini CLI E2E tests (v0.9.2.0) (#252 )	2026-03-20 08:30:09 -07:00
gemini-session-runner.ts	feat: Gemini CLI E2E tests (v0.9.2.0) (#252 )	2026-03-20 08:30:09 -07:00
llm-judge.ts	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
observability.test.ts	fix: never clean up observability artifacts — partial file persists after finalize	2026-03-14 12:37:38 -05:00
session-runner.test.ts	feat: stream-json NDJSON parser for real-time E2E progress	2026-03-14 03:49:36 -05:00
session-runner.ts	feat: /land-and-deploy, /canary, /benchmark + perf review (v0.7.0) (#183 )	2026-03-21 14:31:36 -07:00
skill-parser.ts	feat: 3-tier eval suite with planted-bug outcome testing (EVALS=1)	2026-03-14 01:17:36 -05:00
touchfiles.ts	feat: /cso v2 — infrastructure-first security audit (v0.11.6.0) (#384 )	2026-03-23 06:57:22 -07:00