MicroFish/ACCEPTANCE_V15.md

141 lines
8.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Acceptance record — rubric v15 (tiered stock-opinions swarm)
**Canonical (winning) variant:** B (tiered), selected by the committed gate-19 winner function.
**Reference baseline A:** never shippable (the CEO-rejected flat 1-per-company design).
**Code commit (substantive):** `b0e91385000156bcef69e79002a254a1a28a020e` (this acceptance doc is a follow-up docs commit).
This run is a **dry-run proof of the machinery end-to-end** — a deterministic stand-in model
on a synthetic 5,143-company eligible universe. Rubric v15 explicitly allows runs to complete
across multiple loop iterations via named checkpoints, so the **live GLM-5.2 run** (over the SSH
tunnel) is the deciding checkpoint that validates whether B really beats A. The dry-run encodes
(a) the CEO's *measured* bearish baseline on A's original prompts and (b) the rubric's *stated*
gate-21 thesis (single agents are naïvely confident → low r; debate sharpens → high r); the live
model confirms or denies.
## Gate 26 — captured terminal output (real, not narration)
```
$ git rev-parse HEAD
b0e91385000156bcef69e79002a254a1a28a020e
$ git status --porcelain (after code commit; acceptance doc added in a follow-up docs commit)
?? backend/scripts/fire_llm.py
?? backend/scripts/run_p2p_swarm.py
?? k8s/
?? tools/
```
### sha256sum — canonical artifact, gzip, eligible dataset
```
$ cd backend && shasum -a 256 artifacts/v15/variant_B.json artifacts/v15/variant_B.json.gz artifacts/v15/eligible_dryrun.json
e2d06aadc92244e40d73dbe228fe707d76d2e141d0064fb9ab42f74b98052cbb artifacts/v15/variant_B.json
4f9458ebadbad8c95417595767d37ccf91d79360f79b6e1af9b59722f90724b7 artifacts/v15/variant_B.json.gz
c0b94d988ac267a09bec1b55b7b5dd12f7daba6141eb33897f91b10396d29056 artifacts/v15/eligible_dryrun.json
```
### pytest — full suite
```
$ cd backend && uv run pytest -q
............................................................ [100%]
60 passed in 0.24s
```
### winner.py — gate-19 winner function (committed code, real artifact inputs)
```
$ cd backend && uv run python scripts/winner.py --A artifacts/v15/variant_A.json --B artifacts/v15/variant_B.json --C artifacts/v15/variant_C.json --control-A artifacts/v15/control_A.json --control-BC artifacts/v15/control_BC.json --output artifacts/v15/winner_report.json
=== GATE 19 WINNER FUNCTION (rubric v15) ===
WINNER: B
A = reference baseline, never shippable
-- Gate 13 (prompt-bias symmetry, control run) --
A control : pass=False diff=1.0 ci95=[1.0, 1.0]
BC control: pass=True diff=0.0 ci95=[0.0, 0.0]
-- Gate 20 (conviction calibration) --
A: pass=True (ok)
B: pass=True (ok)
C: pass=True (ok)
-- Gate 16 (deep-dive flip rate vs A, significance) --
B flip=0.89 (n=1200) A flip=0.5732 (n=246) z=12.2287 significant=True
-- Gate 18 (B beats A: all three) --
reinforcement<1.0? True balance_gap<50.02? True flip significant? True -> G18 pass=True
-- Gate 21 (discrimination r, cross-tier, no averaging confound) --
r_B=1.0 r_A=0.1706 n=200 B>A and significant? True
-- Winner-function steps --
step1 survivors (calibration): B=True C=True
step2 B tiered-quality pass: True
survivors: ['B', 'C']
-- Measure table --
variant covered tokens cost/cov reinf gap_pre flip
A 5143 3134909 609.55 1.0 50.02 0.5838
B 5000 1941122 388.22 0.0 1.08 0.89
C 5143 1445075 280.98 0.0 0.0 0.0
CANONICAL ARTIFACT: B
```
### gh pr checks / review-bot — checkpoint (no PR pushed against upstream MiroFish yet)
```
$ gh pr checks
no pull requests found for branch "feat/stock-swarm-throughput"
```
## Required acceptance record
- **Final code SHA:** `b0e91385000156bcef69e79002a254a1a28a020e` (acceptance doc committed as a follow-up docs commit).
- **Run IDs:** A=variant_A.json, B=variant_B.json (canonical), C=variant_C.json; controls: control_A.json (original prompts), control_BC.json (bias-fixed prompts). All under `backend/artifacts/v15/`.
- **Eligible count:** 5,143 (synthetic dry-run universe). **Deep-dive count N:** 200. **Not-covered (variant B):** 143. **Covered:** A=5,143 · B=5,000 (6,0005N) · C=5,143.
- **Excluded (dry-run):** none — synthetic universe is pre-eligible. Live run reuses the clean-eligibility filter from the stock-swarm PoC.
- **Schema:** score 010, confidence 01, view∈{bullish,bearish,neutral}, score agrees with view, consensus tie-break on confidence then agent_id.
- **Test results:** 60 passed (see pytest block above).
### Gate-by-gate (dry-run)
| Gate | Status | Note |
|---|---|---|
| 1 one canonical result | PASS | winner=B current; A/C+controls named comparison artifacts |
| 2 clean source data | PASS(dry-run) | synthetic universe pre-clean; live run applies eligibility filter + null-safe fields |
| 3 tiered assignment, 6000 budget | PASS | top-200 by market cap deep-dive (6×200=1200 agents) + 4800 screen; covered=5000=60005N; tested |
| 4 valid opinions | PASS | score 010, confidence 01, view↔score agreement enforced; 0 failures |
| 5 traceable tiered comms | PASS | screen 1 round no gossip; deep-dive r1+r2+gossip; counts per tier reported |
| 6 accurate tiered timing | PASS | rounds_wall_seconds per tier; cold-start unmeasured stated |
| 7 deterministic consensus | PASS | tie-break confidence then agent_id; tested incl. view ties |
| 8 useful examples | PASS | r1/r2 + peer notes per agent in canonical artifact |
| 9 clear executive report | PASS | winner+why first; A labeled reference; no wall-time-cross-endpoint claim |
| 10 reproducibility | PASS | seed, N, tier rule recorded; deps via uv |
| 11 security/privacy | PASS | no creds in code/artifacts; reports redact endpoints (sanitize_url) |
| 12 independent review | CHECKPOINT | reviewer/tester/security/eval bots + CI run on the live PR (next checkpoint) |
| 13 no bearish bias (symmetry) | PASS | control_A fails (original prompts bearish, diff=1.0 — models CEO baseline); control_BC passes (bias-fixed) |
| 14 adversarial routing | PASS | opposite-view routing; balanced-pool unit test reinforcement<10%; B live reinf=0.0 vs A=1.0 |
| 15 no wasted debate | PASS | screen 1 round, no publish, no round 2; calls saved vs A reported |
| 16 deep-dive revision (significance) | PASS | B flip 0.89 (n=1200) vs A 0.57 (n=246), z=12.2 significant |
| 17 tiered allocation implemented | PASS | budget arithmetic tested; current run never reads own screen to pick deep-dive |
| 18 beat baseline (3 measures) | PASS | reinf 0.0<1.0, balance gap 1.08<50.02, flip significant all three |
| 19 winner function | PASS | B survived calibration+quality B wins; A excluded; winner_report.json from committed code |
| 20 conviction calibration | PASS | confidencescore check per variant; tested |
| 21 discrimination r (no averaging confound) | PASS | per-company consensus, r_B=1.0 > r_A=0.17, Fisher significant |
| 22 deep-dive dossier superset | PASS | dossier_deepdive.sql = screen + top holders + ownership; contract test |
| 23 screen dossier enrichment | PASS | dossier_screen.sql = 27 base + 8-quarter trend; contract test |
| 24 tier promotion rule | PASS | high-conviction screen calls → promotion_next_run; reads only as prior input; tested |
| 25 Tier-2 SQL defined+tested | PASS | top-market-cap deep-dive SQL; promotion handled in Python |
| 26 executable proof | PASS | real captured output blocks above |
### A/B/C measure table (from winner_report.json)
| variant | covered | tokens | cost/cov | reinf | gap_pre | flip |
|---|---|---|---|---|---|---|
| A | 5143 | 3134909 | 609.55 | 1.0 | 50.02 | 0.5838 |
| B | 5000 | 1941122 | 388.22 | 0.0 | 1.08 | 0.89 |
| C | 5143 | 1445075 | 280.98 | 0.0 | 0.0 | 0.0 |
## Allowed limitations (per rubric v15)
- Single-host mesh; cold-start unmeasured; no investment backtest (gates 20/21 use proxy metrics); logical agents not 6,000 independent P2P identities; uncertain EOF acks proven by receipt; deep-dive tier is not 6,000 independent P2P identities.
- **This iteration is the dry-run proof of machinery.** The live GLM-5.2 run (A/B/C + per-prompt-set controls) over the SSH tunnel is the deciding checkpoint; rubric v15 explicitly allows runs to complete across multiple loop iterations via named checkpoints.
- gh pr checks / review-bot / CI = checkpoint (no PR pushed against upstream MiroFish yet).