Training correctness (silent wrong results):
- ppo: refuse a randomly-initialised reward head when only reward_fn is set
(trl 0.19.1 PPO can't use a reward_fn) instead of training against noise.
- edit_kernels (AlphaEdit): reject a non-finite key-norm (NaN <= 0.0 is False).
- preference_combine (ORPO): length-normalise log-probs so exp() doesn't
underflow and kill the odds-ratio correction (+_read_lens caller wiring).
- ipo: anneal the beta schedule from ipo_tau, not the DPO default dpo_beta.
- distill: mask padding + prompt tokens in the default KL term (labels!=-100).
- block_expansion: freeze all-but-the-ACTUAL-added blocks (clamp over-request).
- formats (KTO): map a -1 label to False (bool(-1) was silently True).
Features that silently did nothing:
- sft: actually install the LongLoRA S² attention override (defensive).
- train --gpus re-exec: pass through --gate/--push-as/--trust-remote-code/
--tracker/--diagnose-gate/--annex-xi/--repro-receipt/--profile/energy flags.
- eval gate-install hook: pass $GATE_SUITE to `soup eval against`, which now
validates the locked suite as a precondition (block on missing/tampered).
- deploy_measure: fold the candidate list into the cache key.
Security:
- sglang: loopback-only CORS (was wildcard).
- fetch: lstat the ORIGINAL path (realpath resolved the symlink -> S_ISLNK
never fired -> write followed the link).
- ui /api/data/inspect: is_under_cwd (commonpath) instead of str.startswith.
- registry lineage: unbounded cycle check (the depth-10 cap accepted a
far-away cycle-closing edge).
- gguf calib: read from the O_NOFOLLOW fd (no close+reopen TOCTOU window).
- namespace_pin: flag ANY created_at drift (repo-recreation moves it forward).
Robustness / cross-platform:
- bench: re-raise typer.Exit (RuntimeError subclass) instead of masking it.
- eval auto: catch typer.Exit so a benchmark failure falls through.
- data split: reject negative --val/--test (negative slice inverted the split).
- trace parser: read utf-8-sig so a BOM'd first record isn't dropped.
- terraform plan: tolerate batch_size="auto" in the runtime estimate.
- rl_checkpoint: only rank-0 writes; atomic optimizer save.
Adds tests/test_code_review_high.py (29 regression tests). ruff clean;
full suite 14844 passed / 120 skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>