Commit Graph

8 Commits

Author SHA1 Message Date
Alpamys cba51e153b fix: close out the 3 deferred code-review items (EMA, hardware-fit, MoD)
These were documented as known limitations in the MEDIUM/LOW pass; now fixed.

1. reward_hack EMA smoothing window — smooth_signal("ema") folded only
   window[-1], so reward_hack_smoothing_window had no effect. Now a windowed
   EMA folds alpha over the whole retained window (oldest→newest) then the new
   sample, so a larger window incorporates more history; a 1-element window
   reduces to the old 2-tap form. Updated test_v07126 (0.3 → 0.275).

2. hardware_fit OOM gate wired into `soup train` — the analytical VRAM
   predictor was never called despite its docstring. Added
   _build_hardware_fit_input (SoupConfig → HardwareFitInput, best-effort;
   None when not statically predictable, e.g. batch_size="auto") and
   _hardware_fit_preflight, run after device detection. Refuses on predicted
   OOM (peak × 1.1 > available) unless the documented --allow-oom-attempt
   opt-out is passed; skips silently on CPU / unknown VRAM, and the flag is
   threaded through the --gpus re-exec.

3. MoD real token-dropping — mod_forward ran the full block on ALL tokens then
   masked (zero compute savings). Now the top-k tokens are gathered into a
   shorter sub-sequence, the block runs on ONLY those tokens (real saving),
   the gated result is scattered back, and unselected tokens pass through
   unchanged. Positional inputs (RoPE cos/sin, 4D-causal attention_mask,
   position_ids, cache_position) are gathered to the sub-sequence; any
   unsafe-to-gather case (positional forward args, KV cache, non-4D mask)
   falls back to the prior correct blend so attention can never be silently
   mis-computed. Validated on CPU (gather/scatter/passthrough/savings +
   fallback); the sub-sequence-attention numerics still warrant GPU validation
   at scale.

Adds tests/test_code_review_deferred.py (7 tests). ruff clean; full suite
14867 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:36:59 +05:00
Alpamys 0db7404a6a test(train): strip ANSI + widen terminal in reward-hack help assertions (v0.71.26)
The two TestRewardHackMitigationCli help-text assertions grepped the raw
`soup train --help` output for the new flags. CI runners emit color codes
that split long option names into non-contiguous characters (same failure
mode documented in test_eval_gate.py for --gate), so the raw substring check
failed on every OS/Python combo while passing locally on a no-color terminal.

Fix follows the established repo pattern: render with COLUMNS=200 (no option
wrapping) and strip ANSI escapes before the membership check. Verified under
FORCE_COLOR=1: both assertions pass. Test-only change, no version bump.
2026-07-01 18:16:53 +05:00
Alpamys 2ffc3743ae feat(train): --reward-hack-detector / --reward-hack-halt CLI flags (v0.71.26)
Folds in the doc-vs-reality cleanup the plan flagged: the docs referenced
soup train --reward-hack-detector / --reward-hack-halt but they were config-only.
Add both as CLI passthroughs mirroring --reward-hack-mitigation (validate value,
set cfg.training field, accelerate re-exec passthrough) so the docs are true and
the reward-hack CLI is consistent. +4 tests (test_v07126: 180 -> 184; full suite
-> 14788).
2026-07-01 17:02:37 +05:00
Alpamys fa992bf381 fix(train): reward-hack mitigation review fixes (v0.71.26)
Fixes from 5 sequential ECC reviews (python/code/security/tdd/verification).

python-review (2 CRITICAL + HIGH/MED/LOW):
- signal/vote coherence: schema now requires the active detector in
  reward_hack_signals + rejects the inactive detector name (was silently
  dropping the primary signal from the vote).
- integral_clamp is its own field (was wrongly hard-wired to beta_ceil).
- task/backend gate runs before controller-config checks; EMA convention
  corrected; type hints; mutable-list default -> tuple + normalised compare.

code-review (4 HIGH + MED/LOW):
- _prune now trims _saved in sync with disk (rollback can't target a deleted
  checkpoint); bang-bang release_count resets after each relaxation (hysteretic
  descent); EMA formula uses standard convention; _escalate no longer burns a
  recovery attempt on a None target; max_recovery_attempts>=1 required with
  rollback; _action_history capped; on_step_end logs errors once; loud warning
  when the mitigation callback can't attach (was a silent safety-off).

security-review (HIGH + MED):
- restore_checkpoint / save_checkpoint refuse a SYMLINKED optimizer.pt
  (torch.load weights_only=False was an RCE via attacker-placed symlink);
  bool-before-int/float guards on all new numeric fields; reward_hack_signals
  max_length=4; empty-signals guard in the callback.

tdd-review: +13 coverage tests (dead-band hold, shim verbatim-on-error,
conservative boundary, read-only-beta dual-write, escalation postconditions,
both-restore, PID D exact, log cap + concurrency, no-top-level-import, fuzz
field-validity).

Test count 152 -> 180 (+2 POSIX-only symlink skips).
2026-07-01 16:38:00 +05:00
Alpamys eb2edb1a51 feat(train): anti-gaming hardening + reward shaping + give-up explainer (v0.71.26 Part D)
Stage 3 of closed-loop reward-hacking mitigation: the controller itself must
not be gameable.

- schema: 6 Stage-3 tunables (signal_smoothing, smoothing_window,
  conservative_on_disagreement, reward_shaping, shaping_kind, shaping_strength)
  + bool guards; reward_shaping requires a control mode + strength>0.
- reward_hack_control: combine_conservative (disagreement -> MAX, stay cautious);
  detect_reward_distribution_drift (bimodal-collapse heuristic); shape_reward_fn
  + apply_reward_shaping (bounded length/repetition/sentinel penalty over the
  wrap_reward_funcs seam, inner called once, verbatim on shim error);
  explain_giveup (plain-English, mirrors why.py). Callback wires per-signal
  smoothing + conservative vote + opt-in drift guard (keep KL high) + logs the
  give-up explanation on early-stop.
- peft_wiring: thread smoothing/conservative params from tcfg.
- grpo.py / ppo.py: apply_reward_shaping BEFORE the buffer capture.

Includes an adversarial-fuzz suite (sawtooth/step/noisy/flip/out-of-range
traces): bounded output, no unbounded jump, no flap, anti-windup holds.

+35 tests (test_v07126: 117 -> 152).
2026-07-01 15:33:46 +05:00
Alpamys 21557b43c1 feat(train): PID-Lagrangian controller + rollback escalation ladder (v0.71.26 Part C)
Stage 2 of closed-loop reward-hacking mitigation: principled control + safety net.

- schema: 7 Stage-2 tunables (pid_kp/ki/kd, signal_target, rollback,
  rollback_patience, max_recovery_attempts) + bool guard; cross-validators —
  PID/rollback tunables require pid_lagrangian mode; reward_hack_rollback
  requires rl_checkpoint_save_every_steps (a cadence to roll back to).
- reward_hack_control: PIDLagrangianPolicy + pid_step (Stooke et al. — constraint
  'signal <= target' whose multiplier β is a PID law; integral anti-windup clamp,
  output clamp [floor,ceil], never crosses 0). Callback pid_lagrangian path drives
  β via PID + escalation ladder: raise -> rollback last-good checkpoint (after
  rollback_patience HACK steps) -> early-stop (after max_recovery_attempts),
  subsuming halt_on_hack.
- rl_checkpoint: RLCheckpointCallback.restore_checkpoint (reload PEFT adapter via
  set_peft_model_state_dict + optimizer.load_state_dict; best-effort, never raises).
- peft_wiring: build PIDLagrangianPolicy from tcfg; build the RL-checkpoint cb
  FIRST and hand its reference to the mitigation callback for rollback.

+26 tests (test_v07126: 91 -> 117; real-peft adapter round-trip included).
2026-07-01 15:18:07 +05:00
Alpamys 925e600702 feat(train): reversible bang-bang KL controller + hysteresis (v0.71.26 Part B)
Stage 1 of closed-loop reward-hacking mitigation: the first closed loop.

- schema: 8 Stage-1 tunables (beta_floor/ceil, trip/release band, dwell_steps,
  release_patience, kl_gain, signals allowlist) + bounds; extended
  _validate_reward_hack_compat (floor<ceil, release<trip, signal allowlist,
  control-mode XOR ref_model_ema_alpha, footgun-reject tunables while off).
- reward_hack_control: MitigationAction + BangBangPolicy + bang_bang_step (pure
  hysteresis: dwell before trip, release_patience before relax, beta geometric
  x/div kl_gain clamped [floor,ceil], never crosses 0; multi-signal vote).
  Callback kl_control path mutates BOTH trainer.beta and trainer.args.beta
  (GRPO stock + variant) / trainer.args.kl_coef (PPO); seeds beta from trainer;
  length-trend signal.
- peft_wiring: build BangBangPolicy from tcfg for kl_control.
- ppo.py: RLSignalBuffer parity + attach_rl_callbacks(task=ppo) (BETA — GRPO
  gets the on-GPU proof; PPO kl_coef mutation is unit-tested).
- train.py: --reward-hack-mitigation off|log_only|kl_control|pid_lagrangian
  flag + validation + accelerate re-exec passthrough.

+33 tests (test_v07126: 58 -> 91).
2026-07-01 15:05:32 +05:00
Alpamys bcb08ae0ab feat(train): reward-hack mitigation instrumentation + log_only telemetry (v0.71.26 Part A)
Stage 0 of closed-loop reward-hacking auto-mitigation: observe only, no
control action.

- schema: reward_hack_mitigation Literal[off/log_only/kl_control/pid_lagrangian]
  gated to grpo/ppo + non-mlx + requires reward_hack_detector; YAML-1.1 bool
  DWIM (off->"off", on/yes rejected with a quote hint).
- utils/reward_hack_control.py (no top-level torch): MitigationLogWriter
  (mirrors TraceLogWriter: thread-lock, rotation, redaction, cwd containment,
  symlink-reject), ControllerState (frozen), combine_signals, smooth_signal,
  telemetry helpers, RewardHackMitigationCallback (log_only path provably never
  mutates beta).
- reward_hacking.RewardHackCallback: last_drop_pct() accessor for the controller.
- peft_wiring: mitigation callback subsumes the plain detector when a mode is set.
- examples/reward_hacking/rewards.py: synthetic length-hack + sentinel proxies
  decoupled from a held-out true_score (the GPU-experiment fixture).
- trace_logger: public redact_value alias (DRY reuse).

+58 tests (tests/test_v07126.py).
2026-07-01 14:47:40 +05:00