soup/examples/reward_hacking
Alpamys fa992bf381 fix(train): reward-hack mitigation review fixes (v0.71.26)
Fixes from 5 sequential ECC reviews (python/code/security/tdd/verification).

python-review (2 CRITICAL + HIGH/MED/LOW):
- signal/vote coherence: schema now requires the active detector in
  reward_hack_signals + rejects the inactive detector name (was silently
  dropping the primary signal from the vote).
- integral_clamp is its own field (was wrongly hard-wired to beta_ceil).
- task/backend gate runs before controller-config checks; EMA convention
  corrected; type hints; mutable-list default -> tuple + normalised compare.

code-review (4 HIGH + MED/LOW):
- _prune now trims _saved in sync with disk (rollback can't target a deleted
  checkpoint); bang-bang release_count resets after each relaxation (hysteretic
  descent); EMA formula uses standard convention; _escalate no longer burns a
  recovery attempt on a None target; max_recovery_attempts>=1 required with
  rollback; _action_history capped; on_step_end logs errors once; loud warning
  when the mitigation callback can't attach (was a silent safety-off).

security-review (HIGH + MED):
- restore_checkpoint / save_checkpoint refuse a SYMLINKED optimizer.pt
  (torch.load weights_only=False was an RCE via attacker-placed symlink);
  bool-before-int/float guards on all new numeric fields; reward_hack_signals
  max_length=4; empty-signals guard in the callback.

tdd-review: +13 coverage tests (dead-band hold, shim verbatim-on-error,
conservative boundary, read-only-beta dual-write, escalation postconditions,
both-restore, PID D exact, log cap + concurrency, no-top-level-import, fuzz
field-validity).

Test count 152 -> 180 (+2 POSIX-only symlink skips).
2026-07-01 16:38:00 +05:00
..
rewards.py fix(train): reward-hack mitigation review fixes (v0.71.26) 2026-07-01 16:38:00 +05:00