Commit Graph

1 Commits

Author SHA1 Message Date
renancloudwalk 95b4acabf2
docs(acceptance): live GLM-5.2 verdict = NO SHIP (gate 13 impossible, gate 21 fails)
Live 5-variant parallel run (1,747s, ~24k real LLM calls) on the clean 5,143
universe. Honest gate-19 winner = NONE.

- Gate 13 (bias symmetry) FAILS for both B and C, confirmed 3 ways
  (original prompt 76% bearish; bias-fixed 73% on committed; anonymized 98%).
  Model-level bias, not prompt-fixable -> rubric stop & report, not rewritten.
- Gate 16 PASS: deep-dive debate produced 1,427 real flips (14.3%) vs A's 1
  (0.47%), z=5.68 significant -- the debate is not theater.
- Gate 18 PASS: B beats A on all three (reinf 0<1, gap 3.17<18.11, flip sig).
- Gate 21 FAIL: r_B=0.6404 < r_A=0.786 -- B does not sharpen calls on GLM-5.2.
- Gate 20 FAIL: calibration off across A/B/C.

Recommendation: swap to a less-pessimistic model and re-run; machinery is
proven end-to-end on real calls (64 tests pass).

Adds: exp_anon_bias.py (confound check), ACCEPTANCE_V15_LIVE.md (real captured
terminal output, gate-by-gate), CEO_BRIEF_V15_LIVE.md.
2026-07-16 08:14:04 -03:00