Add `soup data doctor` and `soup data lint`, killing the top *silent*
fine-tune failures before a single training step: EOS-missing-from-labels
(the #1 "model never stops generating" bug), BOS duplication, no-system-role
templates, and preference-data length bias (the #1 silent DPO degradation) —
none of which any competitor (Unsloth/Axolotl/LlamaFactory) checks for.
- utils/data_doctor.py: 8-check chat-template compat report over a
tokenizer + sampled rows, OK/MINOR/MAJOR taxonomy mirroring diagnose;
--show-mask N renders per-token trained/masked colouring through the
SAME masking dispatch (_build_row_labels) the report itself uses, so
the two can never disagree about what's actually trained.
- utils/data_lint.py: preference-data linter (dpo/orpo/simpo/ipo/bco/kto)
— length bias (Cohen's d), label imbalance, near-duplicates (MinHash),
identical chosen==rejected pairs, prompt leakage.
- commands/data_doctor.py: Typer layer for both commands; strips C0
control bytes before untrusted dataset content reaches the terminal.
- commands/diagnose.py: hardens the --evidence loader against a TOCTOU
symlink swap (O_NOFOLLOW + fstat-on-open-fd), backporting the pattern
soup ship shipped in v0.71.25 (closes v0.71.25 known-limitation (4)).
Live smoke against the real HuggingFaceTB/SmolLM2-135M-Instruct tokenizer
(Windows + RTX 3050) found and fixed two genuine bugs beyond the synthetic
fixtures: the EOS check needed to span-search the whole trained region
(not just the last token), and two apply_chat_template call sites needed
a broad except Exception for jinja2.exceptions.TemplateError.
+173 tests (14788 -> 15042). 5 sequential ECC reviews, every finding fixed.
These were documented as known limitations in the MEDIUM/LOW pass; now fixed.
1. reward_hack EMA smoothing window — smooth_signal("ema") folded only
window[-1], so reward_hack_smoothing_window had no effect. Now a windowed
EMA folds alpha over the whole retained window (oldest→newest) then the new
sample, so a larger window incorporates more history; a 1-element window
reduces to the old 2-tap form. Updated test_v07126 (0.3 → 0.275).
2. hardware_fit OOM gate wired into `soup train` — the analytical VRAM
predictor was never called despite its docstring. Added
_build_hardware_fit_input (SoupConfig → HardwareFitInput, best-effort;
None when not statically predictable, e.g. batch_size="auto") and
_hardware_fit_preflight, run after device detection. Refuses on predicted
OOM (peak × 1.1 > available) unless the documented --allow-oom-attempt
opt-out is passed; skips silently on CPU / unknown VRAM, and the flag is
threaded through the --gpus re-exec.
3. MoD real token-dropping — mod_forward ran the full block on ALL tokens then
masked (zero compute savings). Now the top-k tokens are gathered into a
shorter sub-sequence, the block runs on ONLY those tokens (real saving),
the gated result is scattered back, and unselected tokens pass through
unchanged. Positional inputs (RoPE cos/sin, 4D-causal attention_mask,
position_ids, cache_position) are gathered to the sub-sequence; any
unsafe-to-gather case (positional forward args, KV cache, non-4D mask)
falls back to the prior correct blend so attention can never be silently
mis-computed. Validated on CPU (gather/scatter/passthrough/savings +
fallback); the sub-sequence-attention numerics still warrant GPU validation
at scale.
Adds tests/test_code_review_deferred.py (7 tests). ruff clean; full suite
14867 passed / 120 skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- license_matrix: permissive ↔ weak-copyleft is now symmetric (MIT + LGPL no
longer flagged incompatible regardless of order).
- formats.detect_format: check tool-calling before audio so an audio+tools row
keeps its tool_calls instead of being classified audio.
- formats: reject null message content in the alpaca / sharegpt / vision text
converters (routes the row to the drop path instead of literal None content).
DPO only rejects an explicit null (chosen/rejected may be message lists).
- eval/custom.tool_call_args_subset: hallucinated args on a no-arg expected
call now score 0.0 (was a dead `0.5 if ... else 0.5` ternary).
- monitoring/callback: SSE metric push uses `is not None` so a real 0.0 loss/lr
is not reported as None.
- cans/schema.DeployTarget: reject Windows drive-absolute paths (C:\..., C:/...).
- commands/diagnose: reject a non-numeric evidence score with a clear
BadParameter (was ValueError -> exit 1 with zero output).
- commands/generate: partial-save accumulated examples on a mid-run failure so
paid API spend is not discarded.
- __init__.py: fix the byte-corrupted em dash in the package docstring.
Two MEDIUM/LOW items reverted to documented known limitations after they broke
existing behaviour locked by tests: (1) the reward_hack EMA smoother is a
recursive 2-tap by design — smoothing_window only affects `median`; (2)
`soup train`'s MoD compute-savings and the hardware_fit OOM preflight wiring
are architectural, GPU-validation work left as follow-ups.
Adds tests/test_code_review_medium_low.py (10 tests). ruff clean; full suite
14861 passed / 120 skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Path.resolve()+relative_to() cwd-containment (breaks on Windows 8.3 short
names) -> utils.paths.is_under / is_under_cwd across all holdouts:
migrate/common.py, utils/ollama.py, commands/serve.py, commands/export.py,
commands/generate.py, data/loader.py, commands/data.py (×3), and
eval/checkpoint_intelligence.py (the pre-rmtree prune guard).
Unescaped Rich markup from external data: commands/_eval_v0550.py now
escapes the dataset-read exception message (×2). (adapters.py:315 already
escapes layer.name — that flagged instance was a false positive.)
Symlink-following writes -> project-safe helpers:
- utils/active_sampler.py: atomic_write_text (cwd-contained, symlink-reject).
- ui/app.py: tempfile.mkstemp (O_EXCL, unpredictable name) instead of a
fixed shared-temp path a local attacker could pre-symlink.
typer.Exit-vs-SystemExit audit: the two named instances (bench.py, eval.py)
were fixed in the HIGH tier; cli.py's top-level run() already handles both
SystemExit and typer.Exit correctly, so no further sites needed changes.
Adds tests/test_code_review_recurring.py (7 regression tests). ruff clean;
full suite 14851 passed / 120 skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Training correctness (silent wrong results):
- ppo: refuse a randomly-initialised reward head when only reward_fn is set
(trl 0.19.1 PPO can't use a reward_fn) instead of training against noise.
- edit_kernels (AlphaEdit): reject a non-finite key-norm (NaN <= 0.0 is False).
- preference_combine (ORPO): length-normalise log-probs so exp() doesn't
underflow and kill the odds-ratio correction (+_read_lens caller wiring).
- ipo: anneal the beta schedule from ipo_tau, not the DPO default dpo_beta.
- distill: mask padding + prompt tokens in the default KL term (labels!=-100).
- block_expansion: freeze all-but-the-ACTUAL-added blocks (clamp over-request).
- formats (KTO): map a -1 label to False (bool(-1) was silently True).
Features that silently did nothing:
- sft: actually install the LongLoRA S² attention override (defensive).
- train --gpus re-exec: pass through --gate/--push-as/--trust-remote-code/
--tracker/--diagnose-gate/--annex-xi/--repro-receipt/--profile/energy flags.
- eval gate-install hook: pass $GATE_SUITE to `soup eval against`, which now
validates the locked suite as a precondition (block on missing/tampered).
- deploy_measure: fold the candidate list into the cache key.
Security:
- sglang: loopback-only CORS (was wildcard).
- fetch: lstat the ORIGINAL path (realpath resolved the symlink -> S_ISLNK
never fired -> write followed the link).
- ui /api/data/inspect: is_under_cwd (commonpath) instead of str.startswith.
- registry lineage: unbounded cycle check (the depth-10 cap accepted a
far-away cycle-closing edge).
- gguf calib: read from the O_NOFOLLOW fd (no close+reopen TOCTOU window).
- namespace_pin: flag ANY created_at drift (repo-recreation moves it forward).
Robustness / cross-platform:
- bench: re-raise typer.Exit (RuntimeError subclass) instead of masking it.
- eval auto: catch typer.Exit so a benchmark failure falls through.
- data split: reject negative --val/--test (negative slice inverted the split).
- trace parser: read utf-8-sig so a BOM'd first record isn't dropped.
- terraform plan: tolerate batch_size="auto" in the runtime estimate.
- rl_checkpoint: only rank-0 writes; atomic optimizer save.
Adds tests/test_code_review_high.py (29 regression tests). ruff clean;
full suite 14844 passed / 120 skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1. RLVR verifiable rewards (grpo.py, ppo.py): pass verifiable_domain=
to load_reward_fn so `reward_fn: verifiable` + `verifiable_domain: math`
no longer crashes at setup() — the headline RLVR feature was 100% broken.
2. serve: default --host to 127.0.0.1 (was 0.0.0.0) and add an opt-in
--tool-auth-token wired to _create_app so the code-exec tool endpoints
are not exposed unauthenticated on all interfaces; warn on non-loopback
bind without a token.
3. eval benchmark: reject ','/'=' in the adapter's base_model_name_or_path
(and --model path) before lm-eval model_args interpolation, closing the
trust_remote_code=True injection (mirrors the ship.py guard).
4. webhooks: run the private/link-local SSRF check for BOTH http and https
(was nested in the http-only branch, so https://169.254.169.254 and
10.x/192.168.x sailed through). Adds allow_private_hosts= so the trusted
loop_stages LAN-deploy caller (which does its own loopback/LAN tightening)
keeps working.
5. cloud/modal: emit the output_dir via the already-repr'd _LOCAL_OUTPUT
variable in a generated f-string instead of raw interpolation, closing
the stub code-injection hole (also quote-safe, unlike bare {output_dir!r}).
6. eval gate: wire cfg.training.eval_gate through all 14 trainers to
SoupTrainerCallback, and load the suite/baseline + build a live generator
in on_train_begin so --gate actually halts training. Warn honestly when
forgetting/checkpoint/early-stop knobs are set (they are not yet enforced).
Adds tests/test_code_review_critical.py (27 end-to-end regression tests).
ruff clean; full suite 14815 passed / 120 skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The two TestRewardHackMitigationCli help-text assertions grepped the raw
`soup train --help` output for the new flags. CI runners emit color codes
that split long option names into non-contiguous characters (same failure
mode documented in test_eval_gate.py for --gate), so the raw substring check
failed on every OS/Python combo while passing locally on a no-color terminal.
Fix follows the established repo pattern: render with COLUMNS=200 (no option
wrapping) and strip ANSI escapes before the membership check. Verified under
FORCE_COLOR=1: both assertions pass. Test-only change, no version bump.
Folds in the doc-vs-reality cleanup the plan flagged: the docs referenced
soup train --reward-hack-detector / --reward-hack-halt but they were config-only.
Add both as CLI passthroughs mirroring --reward-hack-mitigation (validate value,
set cfg.training field, accelerate re-exec passthrough) so the docs are true and
the reward-hack CLI is consistent. +4 tests (test_v07126: 180 -> 184; full suite
-> 14788).
Fixes from 5 sequential ECC reviews (python/code/security/tdd/verification).
python-review (2 CRITICAL + HIGH/MED/LOW):
- signal/vote coherence: schema now requires the active detector in
reward_hack_signals + rejects the inactive detector name (was silently
dropping the primary signal from the vote).
- integral_clamp is its own field (was wrongly hard-wired to beta_ceil).
- task/backend gate runs before controller-config checks; EMA convention
corrected; type hints; mutable-list default -> tuple + normalised compare.
code-review (4 HIGH + MED/LOW):
- _prune now trims _saved in sync with disk (rollback can't target a deleted
checkpoint); bang-bang release_count resets after each relaxation (hysteretic
descent); EMA formula uses standard convention; _escalate no longer burns a
recovery attempt on a None target; max_recovery_attempts>=1 required with
rollback; _action_history capped; on_step_end logs errors once; loud warning
when the mitigation callback can't attach (was a silent safety-off).
security-review (HIGH + MED):
- restore_checkpoint / save_checkpoint refuse a SYMLINKED optimizer.pt
(torch.load weights_only=False was an RCE via attacker-placed symlink);
bool-before-int/float guards on all new numeric fields; reward_hack_signals
max_length=4; empty-signals guard in the callback.
tdd-review: +13 coverage tests (dead-band hold, shim verbatim-on-error,
conservative boundary, read-only-beta dual-write, escalation postconditions,
both-restore, PID D exact, log cap + concurrency, no-top-level-import, fuzz
field-validity).
Test count 152 -> 180 (+2 POSIX-only symlink skips).
#81 Quant Menu for vision/audio modality
- config/schema: drop the `modality != "text"` rejection in
_validate_quant_menu_supported_tasks (mlx-backend gate retained) so the full
Quant Menu (gptq/awq/hqq:Nbit/aqlm/eetq/mxfp4/fp8) applies to vision+audio.
- trainer/sft: _setup_vision_transformers + _setup_audio_transformers call
build_quantization_config_for_loader (strict superset of the inline 4bit/8bit
BNB blocks they replaced); drop BitsAndBytesConfig import; retain the
prepare_model_for_kbit_training gate on (4bit,8bit,mxfp4).
#80 multipack DataLoader sharding under FSDP/DeepSpeed/DDP
- utils/multipack_trainer: get_train_dataloader routes the multipack DataLoader
through accelerator.prepare when num_processes > 1 so accelerate's
BatchSamplerShard shards whole FFD bins across ranks (even_batches=True avoids
the epoch-boundary collective hang; seed identical across ranks, no `+ rank`).
Single-process path unchanged. Defence-in-depth guard against an unconfigured
MagicMock num_processes.
Tests: +37 in tests/test_v07119.py (13770 -> 13807). ruff clean.
Supersedes the v0.40.5 vision/audio Quant-Menu and v0.40.4 multipack-FSDP
known-limitations. Full multi-GPU validation remains an INFRA-BLOCKED QA item.
TestTrainCliMinillmOnPolicy::test_flag_in_help removed newlines + spaces but
not ANSI codes, so under CI FORCE_COLOR the Rich-rendered long flag (ANSI codes
between the dashes) failed the substring check. Strip ANSI + remove all
whitespace before the check, matching the cloud/sandbox help tests + the
v0.71.17 precedent. Verified under FORCE_COLOR=1 (114 passed).
Typer/Rich injects ANSI codes under CI FORCE_COLOR that split flag names
(--mole, --citation-style) and wrap panel text, breaking raw substring
asserts. Normalise output (strip ANSI + collapse whitespace) before matching.
Same pattern as the v0.71.1 --record-thumbs fix.
Pre-existing v0.71.13 test flaked on windows-latest CI: record_thumb stamps
time.time() and count_new_thumbs_since uses strict `>`, so on Windows' ~15.6ms
clock resolution the two mid-train thumbs could land in the same tick as
run_started (ts == run_started, dropped). Sleep one clock tick at the start of
slow_train so the thumbs are strictly later. Prod semantics unchanged; tests-only.
#261 iterative_dpo._default_train_fn rendered output as a {dir: ...} mapping
that SoupConfig rejected; render it flat so the spawned soup train succeeds.
#246 CMA-ES merge now loads the base model once and reuses it across the
candidate population (_CachedBaseScorer) instead of reloading per candidate.
#245 soup loop estimate_cost wires run_cost.estimate_run_cost_usd off the last
completed run instead of a 0.0 placeholder; never crashes the daemon.
#244 soup train --track-energy --energy-out persists the measurement JSON so
soup bom emit --energy can attach it to an ML-BOM.
#170 --diagnose-gate is RANK-aware: gate once per cluster (RANK==0), not per node.
Tests: 13476 -> 13511 (+35 in tests/test_v07115.py). Validated end-to-end on
SmolLM2-135M / RTX 3050.
PR #256 attached energy only to CycloneDX; #244's contract is "both
outputs". Add _energy_annotations() so the SPDX model package carries
energy as OTHER annotations (same soup:<field>=<value> naming). Strengthen
the happy-path test to assert energy actually lands in the BOM (not just
that a file is written) and add a --format both case asserting energy in
BOTH cdx + spdx.
Closes#244
Adds --energy flag to soup bom emit: validates the measurement JSON (cwd containment + symlink rejection + JSON + EnergyMeasurement shape) and attaches energy properties to both CycloneDX and SPDX outputs.
Co-authored-by: gittihub-jpg <gittihub-jpg@users.noreply.github.com>
CI failed on the HF-rate-limited runners: test_anchor_term_with_file did a
live from_pretrained that 429'd, so it failed AND its unique MiniLLM-anchor
lines went uncovered, tipping the 77% gate to 76.77% on exactly those jobs
(macos + 3.11 stayed green where the cache warmed).
- Skip test_anchor_term_with_file on OSError (offline / rate-limited) instead
of failing.
- Add test_anchor_term_with_fake_model: a fake tokenizer + tiny nn.Module
exercise the identical _load_anchor + anchor_term lines with no network, so
coverage no longer depends on HF availability.
- Add TestReachableInternals cushion (prompt_compile._resolve_metric,
prompt_distill._build_provider_fn + default-provider wiring) so the gate
sits comfortably above 77% (the DSPy/TextGrad/GEPA optimiser bodies are
uncoverable without the [compile] extra).
Tests 13424 -> 13430.
Lift the v0.68.0 deferred-stub family to live (closes#225, #226, #227, #229):
- #229 local-rl train --once: harvest thumbs -> DPO/KTO/ORPO train via a
soup train subprocess (argv list, no shell); state table tracks last_train_at
(skip-on-no-new-thumbs + skip-on-insufficient-pairs); no --once renders a
systemd/launchd nightly scheduler scaffold. New local_rl_scheduler.py.
- #226 distill-prompt: call the teacher once per trace (Ollama/Anthropic/vLLM)
and write a real dataset (sft/kl -> messages; preference -> chosen/rejected).
- #225 compile / #227 compile-tools: live DSPy/GEPA/TextGrad dispatch behind the
new [compile] extra with a friendly ImportError when absent; injectable seams.
Security: reject \n/\r in the model id + shell-quote ExecStart args (systemd
injection defence). Fix: render train output as a plain string (schema-valid),
with a regression test against SoupConfig.
Tests 13329 -> 13424. Smoked end-to-end: real DPO train on SmolLM2-135M (RTX 3050)
+ real Ollama teacher distillation.
CI renders Typer help with ANSI colour codes under FORCE_COLOR that split the
leading `--` from the flag name, so raw-substring assertions on `--steer`
(serve) and `--output`/`--top-k` (steer train) passed locally but failed in CI.
Strip ANSI before the membership check (same fix as the v0.71.1 --record-thumbs
help assert). Local + FORCE_COLOR=1: 142/142 pass. tests-only, no version bump.