Commit Graph

1 Commits

Author SHA1 Message Date
Alpamys b587deea94 fix: remediate 6 CRITICAL code-review findings
1. RLVR verifiable rewards (grpo.py, ppo.py): pass verifiable_domain=
   to load_reward_fn so `reward_fn: verifiable` + `verifiable_domain: math`
   no longer crashes at setup() — the headline RLVR feature was 100% broken.
2. serve: default --host to 127.0.0.1 (was 0.0.0.0) and add an opt-in
   --tool-auth-token wired to _create_app so the code-exec tool endpoints
   are not exposed unauthenticated on all interfaces; warn on non-loopback
   bind without a token.
3. eval benchmark: reject ','/'=' in the adapter's base_model_name_or_path
   (and --model path) before lm-eval model_args interpolation, closing the
   trust_remote_code=True injection (mirrors the ship.py guard).
4. webhooks: run the private/link-local SSRF check for BOTH http and https
   (was nested in the http-only branch, so https://169.254.169.254 and
   10.x/192.168.x sailed through). Adds allow_private_hosts= so the trusted
   loop_stages LAN-deploy caller (which does its own loopback/LAN tightening)
   keeps working.
5. cloud/modal: emit the output_dir via the already-repr'd _LOCAL_OUTPUT
   variable in a generated f-string instead of raw interpolation, closing
   the stub code-injection hole (also quote-safe, unlike bare {output_dir!r}).
6. eval gate: wire cfg.training.eval_gate through all 14 trainers to
   SoupTrainerCallback, and load the suite/baseline + build a live generator
   in on_train_begin so --gate actually halts training. Warn honestly when
   forgetting/checkpoint/early-stop knobs are set (they are not yet enforced).

Adds tests/test_code_review_critical.py (27 end-to-end regression tests).
ruff clean; full suite 14815 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 15:56:48 +05:00