Live 5-variant parallel run (1,747s, ~24k real LLM calls) on the clean 5,143
universe. Honest gate-19 winner = NONE.
- Gate 13 (bias symmetry) FAILS for both B and C, confirmed 3 ways
(original prompt 76% bearish; bias-fixed 73% on committed; anonymized 98%).
Model-level bias, not prompt-fixable -> rubric stop & report, not rewritten.
- Gate 16 PASS: deep-dive debate produced 1,427 real flips (14.3%) vs A's 1
(0.47%), z=5.68 significant -- the debate is not theater.
- Gate 18 PASS: B beats A on all three (reinf 0<1, gap 3.17<18.11, flip sig).
- Gate 21 FAIL: r_B=0.6404 < r_A=0.786 -- B does not sharpen calls on GLM-5.2.
- Gate 20 FAIL: calibration off across A/B/C.
Recommendation: swap to a less-pessimistic model and re-run; machinery is
proven end-to-end on real calls (64 tests pass).
Adds: exp_anon_bias.py (confound check), ACCEPTANCE_V15_LIVE.md (real captured
terminal output, gate-by-gate), CEO_BRIEF_V15_LIVE.md.