125 KiB
Changelog
All notable changes to Soup CLI are documented here.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Detailed, per-release notes for every published version live on the GitHub Releases page. This file tracks unreleased changes and links out for historical detail rather than reproducing 70+ versions of notes.
Unreleased
Fixed — the trl bounds shipped in v0.72.4 were wrong at both ends.
- The ceiling was over-tight by five releases. v0.72.4 capped
trl<0.25from a table claimingmax_prompt_lengthwas removed forbcoat 0.25 and forkto/orpo/simpoat 0.26. It was not. What happened at 0.25 and 0.26 is that those config classes moved intotrl/experimental/while staying publicly re-exported fromtrlwith the field intact — a module relocation read as a field removal. The real stages arektoat 0.27,bco/orpo/simpoat 0.28,dpo/ipoat 0.29, so the cap is now<0.27andtrl0.25.0–0.26.2 are usable again. Settled by construction, not by reading source: all six configs build on 0.26.2 with the exact keyword arguments the wrappers pass, and 0.27.0 raisesKTOConfig.__init__() got an unexpected keyword argument 'max_prompt_length'whiledpoandorpostill build. - The floor was impossible.
>=0.7.0could never have worked:setup()importsGRPOTrainerunconditionally andtrlfirst exports it at 0.14.0 (OnlineDPOTrainer/KTOTrainer/BCOTrainer/BasePairwiseJudgearrive at 0.11.0; 0.7.0 has none of them). The floor is now>=0.14.0. Resolvers pick the newest allowed version, so this only bit under a constraints file or anyone reading the metadata as a statement of support. - One detail in the v0.72.4 note was also imprecise:
ORPOConfigandCPOConfigare not deleted at 0.29 — the modules survive undertrl/experimental/. They are removed from the publictrlnamespace, which is what Soup imports, so the 0.29 break is anImportErrorrather than a rejected keyword argument. Worse, not milder, than described.
Fixed — encoding corruption in pyproject.toml. Fourteen em-dashes had been
round-tripped through cp1251. Thirteen were in comments; one was the unit pytest
marker description, which pytest --markers prints to users.
Fixed — docs/commands.md. Missing newlines collapsed five commands onto two
lines, hiding three of them from readers, and the page promised "the full command
list" while omitting seven commands that are documented on the topic pages.
[0.72.4] - 2026-08-03
Added — preference losses over layer streaming: DPO, ORPO, SimPO and KTO.
Layer streaming kept the frozen base in CPU RAM and fed it to the GPU one decoder
layer at a time, but only for task: sft. This release opens it to the four
preference losses. The whole risk was one thing: DPO needs a reference model, and a
second model instance would double memory and defeat the feature entirely.
- The reference is the same streamed base with its adapters disabled — one set of weights, one stream, no second pass. Measured on an RTX 3050 4 GB with a 730 MB model: streamed DPO peaked at 0.914x the SFT peak, with a byte-identical RAM store and buffer pool. Forcing a real second instance in the same harness cost +730.44 MB against 730.44 MB of weights — exactly one copy. That control is what makes the first number mean something.
- KTO is not reference-free, contrary to how it is usually described: it selects its reference exactly the way DPO does, so it gets the same treatment and the same memory assertion. ORPO and SimPO genuinely are reference-free.
- Bit-exact against a resident run of the same loss —
0.0difference for all four, the standard every slot in this series inherits. - The pre-flight now knows that a paired loss is twice the rows. DPO, ORPO and SimPO concatenate chosen and rejected into one tensor, so a VRAM budget computed at one row per example would have under-predicted by half — and on Windows the consequence is not an error but a silent spill to host memory that makes the run an order of magnitude slower.
- KTO requires
batch_size >= 2(its KL term is degenerate at 1). Soup now says so when your config is read, rather than minutes later after sharding the checkpoint. KTO is streamable at all only because v0.72.3 lifted the batch-1 restriction. grpoandpporemain excluded permanently, not "not yet": generation rollouts re-read every layer once per generated token, which destroys the amortisation streaming depends on. The refusal says so and deliberately names no release.
The streaming setup now lives in one shared place instead of being copied per trainer, so the NF4 pre-flight, the RAM/disk tier choice and the VRAM fit refusal cannot drift between SFT and the preference losses.
Fixed — trl is now capped, and that is a real bug fix, not a CI tweak.
Six trainers (bco, dpo, ipo, kto, orpo, simpo) pass max_prompt_length to
their trl config, and trl removed it in stages. So anyone who ran
pip install 'soup-cli[train]' and resolved to a recent trl had
soup train --task orpo fail on import.
Correction (see Unreleased). This release shipped the cap as
<0.25on the strength of a staged-removal table that was itself wrong. The real stages arektoat 0.27,bco/orpo/simpoat 0.28 anddpo/ipoat 0.29; the cap is now<0.27.
That was already true before this release and nothing caught it: the trl imports
live inside setup(), which no test had ever called on those wrappers, so CI stayed
green while the code only worked on older trl. This release's end-to-end preference
tests are what surfaced it.
The boundary was read off the published wheels per config rather than inferred from a version number — the removal being staged is exactly why a single spot-check gives the wrong answer, and see the correction above for how that method can still land on the wrong answer. Supporting the newer API is its own piece of work; declaring a dependency the code actually works with comes first.
Honest costs: streaming makes the reference free in memory, not in time — DPO traverses the layer stack three times per step against SFT's two, measured at 1.52x the layer reads. And the VRAM pre-flight is a sound upper bound for preference losses rather than a tight estimate; see Known Limitations in the release notes.
[0.72.3] - 2026-07-28
Added — layer streaming breadth: more architectures, bigger batches, resume, and a disk tier.
Layer streaming (v0.72.0–.2) was deliberately narrow: Llama/Qwen only, batch 1, no gradient accumulation, no resume, RAM only. This release lifts all of it, and each capability was gated against a streamed-vs-resident bit-exactness reference before any of it was written.
- Six more model families.
mistral,gemma,gemma2,gemma3_text,phiandphi3join the allowlist, each verified bit-exact against the same checkpoint loaded resident, under both bf16 and NF4. Phi-3 is the notable one: it fuses Q/K/V into a singleqkv_proj, so there is noq_projto find, and it is bit-exact anyway. Multimodalgemma3is deliberately not accepted — onlygemma3_text. batch_sizeabove 1, and gradient accumulation. Both previously refused.- A pre-flight VRAM budget that accounts for batch and vocabulary. Streaming bounds
the weights; activations and the logits tensor are untouched by it and both scale
with
batch × seq. On a 152k-vocab model at batch 8 the logits term alone measured 8.71 GB — 146× the entire layer-buffer pool.soup trainnow predicts peak VRAM and refuses a run that will not fit, naming the two knobs that scale it. - A throughput forecast, quoted as a range from a GEMM ceiling measured on your own card in that session, alongside the SM clock — never a compiled-in per-card constant.
--resumeand--hf-resumework with streaming.- A disk overflow tier.
stream_source: auto(the default) uses RAM when the base fits and falls back to an NVMe disk tier when it does not, holding nothing resident.stream_source: ramrefuses instead of falling back. Non-NVMe disks are still refused outright. soup doctor --diskreports the detected media type.
Fixed
estimate_logits_bytescharged 6 bytes per logit element; the measured peak is 14 (transformersholds the bf16 logits, the fp32 upcast, log-softmax's fp32 output and the fp32 gradient live at once). The old figure under-predicted that term by 2.33×.- Adapters could not be loaded into a streamed model:
load_state_dictnarrows keys by child name, so a canonical checkpoint matched 0 of N tensors and PEFT reported only a warning — a resumed run reproduced the from-scratch loss curve exactly. The streaming layer now redirects canonical keys at load time, mirroring the v0.72.1 save-side fix. - The NVMe-only tier guard was wired to a hardcoded constant and could never fire.
- Streaming weight sources are now released when training ends or raises; the disk tier holds one open shard handle per decoder layer.
- Subprocess helpers resolve tools to absolute paths (on Windows,
CreateProcesssearches the current directory beforePATH). - The
[mcp]extra is now capped atmcp<2. The SDK's 2.0.0 release removedmcp.shared.memory.create_connected_server_and_client_sessionand droppedServer.list_tools, breakingsoup mcp serve's round-trip tests for anyone installing fresh. Support for the 2.x API is tracked separately.
Known limitations
- The RAM-vs-disk performance gap is unmeasured on the development hardware and no number is claimed for it. safetensors memory-maps the shards, so the OS page cache keeps them resident between steps on a machine with spare RAM, and at ~5 effective TFLOPS the NVMe read hides under compute. The disk tier's correctness is verified bit-exact against the RAM tier; its speed relative to RAM is not characterised.
- End-to-end
soup train --resumecould not be demonstrated on the development box:transformersrefusestorch.loadbelow torch 2.6 (CVE-2025-32434), which blocks every resume there, streaming or not. The streaming-specific half — the adapter round-trip and loss continuity — is verified on the production CUDA path. - Loading into a streamed model works;
named_parameters()andstate_dict()still disagree in memory, which is the deliberate cost of a serialisation-only design. - Layer streaming remains BETA.
[0.72.2] - 2026-07-28
Added — NF4 layer streaming: fine-tune Llama-3.1-8B on a 4 GB laptop GPU.
Layer streaming (v0.72.0) keeps the frozen base in CPU RAM and streams it to the GPU one decoder layer at a time, so peak VRAM is bounded by one layer instead of the whole model. It was bf16-only, which capped it at about 3B on a small card. Quantising the streamed base to NF4 shrinks it ~4×, and that is what brings 8B within reach.
Add one line to a streaming config:
training:
stream_layers: true
quantization: 4bit # NF4
batch_size: 1
Measured on a 4 GB RTX 3050 Laptop (Windows, batch 1, S=512, gradient
checkpointing, 50 steps after 10 warm-up, PagedAdamW8bit):
| Model | tok/s | Peak VRAM | RAM store | GPU util |
|---|---|---|---|---|
| Llama-3.1-8B-Instruct | 119.6 | 3.32 GB | 3.60 GB (page-locked) | 100% |
| Qwen2.5-3B | 264.2 | 1.76 GB | 1.43 GB (page-locked) | 100% |
For scale: 1M training tokens is about 2.3 h at 8B on that card (arithmetic from the measured rate, not a separate measurement).
Qwen2.5-3B also got 1.85× faster than the bf16 streaming path (264.2 vs 143.1 tok/s) — and the reason is not arithmetic. A 1.43 GB store fits under the machine's page-locked memory ceiling where a 5.55 GB one did not, which restores asynchronous host-to-device copies and lifts GPU utilisation from 79.3% to 100%.
Correctness. A streamed NF4 run is bit-exact against a resident NF4 run: the same quantised bytes through the same bitsandbytes kernels. Logit equality, non-zero gradients at layer 0, and a matching multi-step loss curve are all regression tests, not one-off measurements.
The base is quantised once, offline, and cached under ~/.soup/layer-stream/.
The cache is keyed to the quantisation and the source checkpoint, so switching
between none and 4bit, or retraining a base in place, re-shards instead of
silently streaming the wrong bytes.
Fixed — a streamed 4-bit run reported its parameter count ~6.5× too high. SmolLM2-135M printed "878,154,048 total" (true: 134,515,008). Display only — training was unaffected — but at 8B it would have read ~52 B.
Scope is unchanged and still BETA: RAM tier, task: sft, Llama/Qwen,
batch_size: 1, no gradient accumulation, no --resume. Every rejected config
names the release that lifts it. quantization values other than none and
4bit are refused.
Fixed — soup --help was 5x slower than it should be. Since v0.72.0 the CLI
imported PyTorch on startup, taking 6.0 s where it now takes 1.15 s.
soup reward stress (v0.71.41) put utils/reward_stress on the light CLI path.
That module imported utils/reward_hack_control to reuse a single string
constant — and reward_hack_control resolves its TrainerCallback base class at
module scope, which pulls in transformers and torch. Importing a ~4.4 s
dependency for one constant made every soup invocation pay for the training
stack, including commands that never touch a model.
Nothing produced wrong results; this was purely startup latency. The light core
(pip install soup-cli without the [train] extra) was never broken — it fell
back cleanly when torch was absent, just slowly when it was present.
Added tests/test_cli_startup_is_light.py, which asserts the invariant at
runtime (import soup_cli.cli must not put torch in sys.modules) instead of
inspecting source text for import torch, which is what the previous guards did
and why this went unnoticed.
[0.72.1] - 2026-07-27
Fixed — layer-streaming adapters were saved in an unloadable form. If you
trained with stream_layers: true on v0.72.0, the adapter that run wrote is
inert: every tensor was saved under a key containing an extra .inner.
segment, so soup merge, soup serve, soup chat and
PeftModel.from_pretrained all loaded zero adapter tensors and silently
returned the untuned base model. PEFT emitted only a UserWarning, so nothing
failed and nothing looked wrong.
The training itself was correct — the streamed run's numerics are unaffected, and v0.72.0's bit-exactness results still stand. Only the saved file was affected.
If you have a v0.72.0 streamed adapter: re-save or re-run it on v0.72.1.
There is no way to recover the original file's association with the base model
beyond renaming its keys; re-running is the reliable path. A quick check —
if adapter_model.safetensors contains keys with .inner. in them, it is
affected:
python -c "from safetensors.torch import load_file; \
print([k for k in load_file('adapter_model.safetensors') if '.inner.' in k][:3])"
Streamed adapters now save byte-for-byte in the same layout as an ordinary LoRA run, and are portable to any tool that has never heard of layer streaming.
Also fixed — --hf-resume bypassed the streaming resume refusal. The guard
only tested --resume, but --hf-resume reaches resume_from through a
different branch. That combination previously appeared to work by accident
(checkpoint and live model shared the same key shape); once adapters are saved
canonically it would instead have matched nothing and continued with a
freshly initialised adapter, silently. Both flags are now refused for streaming
runs, naming v0.72.3.
Also in this release: every "this lands in vX.Y.Z" refusal message was corrected after the v0.72.x roadmap was renumbered (NF4 streaming is now v0.72.2; the disk tier, wider architectures, larger batches, gradient accumulation and checkpoint/resume are v0.72.3; preference losses are v0.72.4).
Known limitation: in memory the streamed model's named_parameters() still
carries the wrapper segment, so loading into a streaming run (--resume)
remains unsupported and is refused with a message naming v0.72.3.
[0.72.0] - 2026-07-26
Superseded by v0.72.1 — adapters saved by this version load as zero tensors. The entry below is left as published; the defect and the fix are described under [0.72.1]. Version numbers named as "upcoming" below were also renumbered there (NF4 is v0.72.2, not v0.72.1).
Layer streaming (BETA) — fine-tune models that don't fit in your card. The frozen base lives in CPU RAM and is streamed into two pre-allocated VRAM buffers one decoder layer at a time, so peak VRAM is bounded by the size of one layer instead of the whole model. Only the LoRA adapters, their gradients and optimizer state stay resident. Slower than resident training — but these models did not run on the card at all.
Measured on the development box (RTX 3050 Laptop 4 GB, Windows 11, 16.9 GB RAM), batch 1, gradient checkpointing on, 50 steps after 10 warm-up:
| Model | S | tok/s | GPU util | Peak VRAM |
|---|---|---|---|---|
| Qwen2.5-0.5B | 512 | 978.6 | 91.4% | 1.47 GB |
| Qwen2.5-1.5B | 512 | 525.0 | 96.8% | 1.82 GB |
| Qwen2.5-1.5B | 1024 | 487.6 | 96.7% | 2.96 GB |
| Qwen2.5-3B | 512 | 143.1 | 79.3% | 2.15 GB |
Qwen2.5-3B trains in 2.15 GB on a 4 GB card where a resident run OOMs. The honest cost: 1.43× slower than resident, measured at 0.5B — the only apples-to-apples comparison available on this box, because 1.5B and above cannot run resident here at all.
Added
training.stream_layers: true— stream the frozen base layer-by-layer from CPU RAM.training.stream_source(auto/ram/disk) andtraining.stream_buffers(2–8, default 2 = double buffering) tune it.soup_cli/utils/layer_stream.py— tier choice, pinned-vs-pageable decision, architecture allowlist, VRAM/throughput arithmetic (no torch import).soup_cli/utils/layer_shard.py— rewrites an HF checkpoint into one safetensors shard per decoder layer, one tensor at a time, so sharding a model that does not fit never needs it to fit.soup_cli/utils/layer_stream_runtime.py— buffer pool, CPU-RAM weight source, prefetch scheduler on a dedicated CUDA stream, and the streamed layer wrapper.- Shards are cached under
~/.soup/layer-stream/(override withSOUP_LAYER_STREAM_CACHE_DIR) and keyed to the source checkpoint's fingerprint, so a base retrained in place re-shards instead of silently training against stale weights. - When the base cannot be page-locked, the RAM store falls back to pageable memory and says so, including the measured cost (GPU utilisation ~97% → ~79%).
Changed
- The pre-flight hardware-fit gate is skipped for streaming runs: it models a resident run and would otherwise refuse exactly the runs streaming enables.
- Gradient checkpointing is handled per-layer by the streamer; the HF Trainer's own is left off so layers are not recomputed twice.
Known limitations
- BETA, and proof-of-mechanism at 3B. Nothing above 3B was measured. No 8B/14B claim is supported.
- Scope: RAM tier, bf16,
task: sft, Llama/Qwen, batch size 1, no gradient accumulation, no--resume. Every refusal names the release that lifts it. - 4-bit (NF4) streaming is v0.72.1 — NF4 weights carry a quantisation state and cannot be byte-copied into a plain buffer.
- The disk overflow tier, larger batches, gradient accumulation and checkpoint/resume are v0.72.2.
- The 3B number used a pageable store (this box cannot page-lock 5.55 GB), so it is a lower bound.
expandable_segments:Trueis silently ignored on Windows; Soup detects this and does not claim it is active.- Numbers are Windows/WDDM and therefore systematically pessimistic vs Linux.
[0.71.41] - 2026-07-19
soup reward stress: is your reward verifier gameable? Turn the reward-hacking
detector on the verifier itself. soup reward synth (v0.71.40) proves a verifier
separates your references from friendly perturbations; stress asks the adversarial
question a reward-hacking model asks at train time — does the verifier pay out for
degenerate junk? It feeds empty, length-padded, repetition, and sentinel-spam
completions and flags any the verifier accepts. Pure, offline, exit 0 = robust /
2 = gameable / 1 = error. Nothing in TRL / Unsloth / Axolotl / OpenRLHF tests a
verifier for gameability.
Added
soup reward stress <reward.py|builtin> [--references golds.jsonl]— adversarial verifier probe. Attacks (--attacks empty,length,repetition,sentinel,--sentinel) are scored against the real gold, so numeric / tool_call / json_schema verifiers get a valid target and still must reject the junk. Reports a per-attack accept-rate table + an overall gameability verdict (--max-gameable,--threshold,--output-report). Loads the target through the existing reward loader, so it probes a synthesized.pyand a builtin (accuracy/format/verifiable). A gold-requiring verifier probed with no--referencesis a hard error, never a false "robust".
Fixed
- Corrected the Telemetry section in the ops docs: Soup's telemetry primitives exist but are not wired to any command — no data is ever sent today (the previous wording implied a live opt-in sender). Wiring is deferred until a public privacy policy ships.
[0.71.40] - 2026-07-19
soup reward synth: auto-generate a deterministic reward verifier from your data.
Point it at a JSONL of reference (gold) outputs and it infers a verifier, emits a
readable / committable .py reward function, and — the moat — refuses to emit one
that can't tell your references from auto-generated bad answers. Nothing in
TRL / Unsloth / Axolotl / OpenRLHF synthesizes a reward; every reward today is
hand-written, hand-picked, or a trained-weights artifact.
Added
soup reward synth <references.jsonl> -o reward.py— deterministic verifier synthesis. Four families, auto-detected (or pick with--kind):numeric(last-number /\boxed{}/####extraction, exact or--tolerance),json_schema(induced keys + types + required),regex(positional char-classes over equal-length golds),tool_call(per-toolrequired/allowedargument binding). The emitted file is self-contained and ridesload_reward_fn's existing.pypath — no new trusted-exec surface; you read, edit, commit, and diff it.- Mandatory calibration report — the synthesized verifier is loaded back and run
against its own references (must accept ≥90%) and auto-perturbed negatives (must
reject). A degenerate always-accept verifier is refused (
exit 2), never silently emitted.--plan-onlyreports the induced spec without writing;--output-reportpersists the calibration JSON. - Comma-separated
reward_fn("accuracy,format") now trains — it resolves to a reward ensemble (GRPOTrainer(reward_funcs=[...]), and unlocks therm_ensemblereward-hack detector which needs ≥2 rewards). GRPO-only, validated at config-parse time. Fixes a recipe (deepseek-v3-reasoning) that shipped exactly this and previously crashed withUnknown reward function(#311).
Changed / Fixed
training.reward_fngains a field validator (null-byte / blank / oversize / empty-comma-segment rejection) — the oldest arbitrary-code field was the least guarded. Comma +verifiablewithout averifiable_domainnow fails at parse time like the bareverifiableform.envs/calculator.py/envs/guess_number.pydocstrings corrected: the reward isreward_fn: verifiable+verifiable_domain: math(the barereward_fn='math'they showed was never valid).
[0.71.39] - 2026-07-19
"CI for weights, not prompts": close the evidence loop. soup ship's verdict is
now something Soup can emit, commit, review, and bind to the exact model that
produced it — turning the soup ci init gate from "edit two numbers in a JSON file"
into a reproducible, provenance-bound check that renders on every PR.
Added
soup ship --emit-evidence <path>— re-serialises the verdict into the--evidenceINPUT schema, so a run's output is replayable as input: feeding it back through--evidence(same--forgetting-threshold) reproduces an identical verdict. Output is finally input.ShipConfigundereval.shipinsoup.yaml+soup ship --config soup.yaml— commit the gate policy (task_eval/task_mode/general_suite/forgetting_threshold/judge_model/baseline) so the verdict is reviewable in a PR diff and reproducible. An explicit CLI flag always wins (CLI > config > default).soup ship --push owner/repo#N— post the verdict as a GitHub PR comment (reuses thesoup adapters pr --pushgh apiplumbing). Best-effort: a missing token /ghfailure warns but never flips the SHIP / DON'T-SHIP exit code.- Evidence provenance + staleness gate. With
--emit-evidence,--configSTAMPS aprovenanceblock (config_sha— a semantic, order-insensitive recipe hash — plusbase_modeland a best-effortdata_sha) onto the evidence. With--evidencealone,--configGATES: it refuses (exit 3) evidence whoseconfig_shadrifted from the committed config, so a PR that changedsoup.yamlbut forgot to recompute its evidence is caught. The gate policy (eval.ship) is EXCLUDED from the hash, so tuningforgetting_thresholdnever falsely invalidates evidence about an unchanged model. soup ci init --config <soup.yaml>— binds the generated gate'ssoup shipstep to the committed config (provenance/staleness enforcement in CI).
Changed
- Exit-code note:
soup ship --configusage / staleness errors are exit3(usage), preserving0 = SHIP,2 = DON'T SHIP,1 = runtimefrom v0.71.38.
Security
provenance.config_sharead from untrusted evidence is shape-validated as a hex digest before being echoed (a raw value could smuggle terminal ESC bytes pastrich.markup.escape).provenance.data_shahashes the training file through anO_NOFOLLOWfd with a symlink-rejecting containment check and an 8 GiB cap (was an unguardedhash_file).soup ci init's path validation now rejects#and the YAML 1.1 line breaks (NEL / LS / PS), closing a plain-scalar comment-truncation of the generatedrun:step (also hardens the pre-existing--data/--suite/--evidenceargs).
[0.71.38] - 2026-07-17
soup ship's regression leg now has teeth. Leg 2 (the catastrophic-forgetting
/ regression gate that carries the whole SHIP / DON'T-SHIP claim) was 15
hand-written trivia prompts scored by case-insensitive substring containment
— it scored "B" for "Berlin", "ok" for "look", "3" for "13", and
had zero items for tool-calling, safety, or JSON validity. This release makes
the gate real: a fixed, extraction-based scorer + bundled, offline, zero-dep eval
suites that catch a regression the old gate waved through.
Changed
- Fixed answer scorer (breaking — verdicts can change).
soup ship's leg-2 MCQ / instruction / arithmetic answers are now scored by answer-extraction- a boundary-aware match, replacing the raw substring test. A spurious substring inside another word ("Berlin", "look", "13") no longer scores a correct answer, so an existing run's verdict may flip — intentionally, because the old gate was reporting false negatives.
- Bundled, offline general suite. The default
--general-suiteis now seven hand-authored suites shipped in the wheel:mini_mmlu/mini_common_sense/mini_instruction(expanded), a newmini_arithmetic, and three behavioural suites the old gate had no coverage for —mini_tool_call(function-calling),mini_format_json(JSON validity), andmini_safety(refusal-rate). Each is scored to a per-model absolute score by the pure scorers Soup already ships (eval/custom,utils/diagnose); no lm-eval, no network, no download. Every suite is large enough that a single-item flip (1/N < 0.05) trips the default threshold instead of being rounded away. soup shipexit codes: usage errors moved 2 → 3. Exit2now means only DON'T-SHIP; a typo'd flag or bad--general-suiteexits3(mirroringsoup plan/soup env check), so CI can tell a config error from a caught regression. Offline--evidenceread/parse errors stay1. Breaking for anyone parsing exit2as "usage error".
Fixed
soup shiphelp + docstrings no longer describe--task-mode pairwiseas "reserved for a later release" (it shipped in v0.71.31); the deadSUPPORTED_TASK_MODESgate is removed.soup diagnose's package docstring said "Six" probes (there are seven — citation) and pointed live loading at an unshipped version; it now re-exports all sevenscore_*probe functions so callers need not reach into submodules.
[0.71.37] - 2026-07-17
Every pip install soup-cli[extra] command now works on Windows cmd.exe, and
eval-gate benchmark tasks run instead of always failing.
Fixed
-
Install hints are now quoted so they work in every shell. Soup printed
pip install 'soup-cli[ui]'— bash / zsh / PowerShell syntax.cmd.exehas no single-quote quoting, so it hands the quotes to pip verbatim and pip refuses:ERROR: Invalid requirement: "'soup-cli[train]'": Expected package name at the start of dependency specifierEvery hint, README command, and docs example now uses
pip install "soup-cli[extra]", which works in cmd, PowerShell, bash, and zsh alike — the same spelling the repo already used forpip install -e ".[dev]". Measured on Windows: single quotes fail only oncmd.exe; double quotes pass everywhere; dropping the quotes passes on Windows but breaks zsh, which globs the bracket.Nothing in Soup can rescue the command after it is typed — pip and the shell own it, and Soup is not installed yet when the README command runs — so the fix is the spelling we print. A regression test now scans the package and every docs code block for the single-quoted form.
If you followed an older tutorial and hit
Invalid requirement, swap the'for"; nothing is wrong with the package. -
Eval-gate
type: benchmarktasks now actually run.eval/gate.pyprobed for aforgetting.run_mini_benchmarkhelper that never existed, so everytype: benchmarktask in a gate suite failed 100% of the time — while advising an[eval]extras install that could not fix it. The gate now callsForgettingDetectordirectly (the same waysoup shipalready did), and an unknown benchmark name fails with the list of valid names. Thanks @Sanjays2402! (#315, closes #310)
[0.71.36] - 2026-07-16
Data Moat II — a semantic layer over your training data, plus two tools for what a fine-tune forgets and leaks.
Added
soup data dedup --semantic— near-duplicate removal over embedding cosine instead of MinHash shingling. Catches reworded duplicates that MinHash misses (measured: 0.88–0.91 cosine on rewordings MinHash scored as distinct) while correctly keeping distinct-but-similar instructions. Zero new dependencies: usestransformersfrom the[train]extra. Read the known-limitation below before lowering--threshold.soup data topics <data>— cluster a dataset and label each cluster with c-TF-IDF terms, plus a coverage table (82% code · 6% math) and a warning for thin topics. Labels are emergent term clusters, not a fixed taxonomy.soup data canary insert|check— Secret-Sharer memorization probe. Insert K high-entropy secrets, then check any model/adapter: each secret's loss is ranked against never-inserted controls drawn from the same space. Exit 2 on MAJOR so CI can gate. Measured on SmolLM2-135M: a memorized set lands at percentile 0.0 (loss 1.7–2.5) against a clean model's 4.1–6.2.soup train --replay old.jsonl --replay-ratio 0.1— continual-learning rehearsal. Interleaves a seeded sample of an old dataset into training so a new task does not erase the old one.ris the fraction of the final mixed set (n_replay = round(r/(1-r) · n_new)), rows are interleaved rather than appended, and an undersized pool reports the shortfall instead of repeating rows. Mixed intotrainonly — validation stays pure new-task.
Fixed
- The hardware-fit gate refused to train any local checkpoint. A merged
model (
soup merge -o ./mymodel) has no size marker in its name, so the size guesser returned its 7B default, predicted ~16 GB of VRAM and refused. Local checkpoints are now measured from their safetensors header (0.135B actual vs 7.0B guessed) — this had blockedsoup merge→ train-from-merged entirely. Third instance of this class after the v0.71.32 (Whisper) and v0.71.33 (Msuffix) fixes. pip install 'soup-cli[extra]'hints printed without the extra. Rich ate the bracket, so every "install the missing dependency" message across 17 sites told users to runpip install 'soup-cli'— which succeeds and still leaves the feature broken. Affected[eval],[data],[serve],[ui],[tui],[compile],[mcp],[carbon]and others, including Typer help text.- Replay rows bypassed the image/audio path-traversal validation that the primary dataset receives.
Known limitations
- Semantic dedup is not a paraphrase detector. Measured with
all-MiniLM-L6-v2, paraphrase cosines (0.49–0.76) overlap with
genuinely-distinct rows (0.54–0.76): "Add two numbers" vs "Multiply two
numbers" scores 0.759, higher than the true paraphrase "reverse a string" /
"invert the order of characters" at 0.491. No threshold separates them, so
lowering
--thresholdto chase paraphrases deletes real training rows. The default (0.8) is deliberately conservative. - Replay is validated at proof-of-mechanism scale. On SmolLM2-135M + LoRA, replay retained the old task 7% better than a no-replay control — the correct direction — but forgetting without it was only +4%, i.e. mild. The effect size at full fine-tuning or 7B+ is unproven on a 4 GB box.
- Canary exposure is the sampled-control approximation, not full-space rank enumeration. "No exposure" is not proof of no memorization.
data topics/dedup --semanticrequire[train](torch) and download an embedding model. Plain MinHashdedupstays on the light core.- Replay v1 is
sft/pretrainonly and is incompatible withpacking/multipack.
[0.71.35] - 2026-07-15
Added
- Compliance templates —
soup init --template hipaa|soc2|eu-ai-act|sr-11-7. Four regulation-shaped starting configs. Soup's compliance controls are CLI flags/commands rather than config keys, so each template is a valid training config plus header comments naming the exact commands for that regime (PHI scrubbing + air-gap for HIPAA, BOM/attest/sign for SOC 2, Annex XI + energy tracking for the EU AI Act, repro-receipt + diagnose/ship for SR 11-7). Templates default to a license-clean Apache-2.0 base. soup card <registry-id> -o MODELCARD.md— model-card autogen. Turns a Local Model Registry entry into a publishable, provenance-carrying HF model card: base model, training config, eval scorecard, config/data hashes, lineage (ancestors) and a table of every registered artifact. Adapter vs full-model is inferred from registered artifacts, falling back to the training config (LoRA rank, with Spectrum/LISA full-FT correctly treated as dense), so the card sets the rightlibrary_nameand never misreports the model type.soup push --card <registry-id>— render that registry-driven card and upload it asREADME.md, overriding the auto-generated one. A bad ref fails fast before any network call; HF hub only.soup ci init— fine-tuning CI. Writes.github/workflows/soup-gate.yml, a PR gate chainingsoup data validate→soup expect→soup ship --evidence(exit 2 blocks the merge). Every interpolated path is validated to stay under the repo root and shell-quoted; the branch and Python version are regex-gated; the write is atomic, symlink-rejecting, and refuses to clobber an existing workflow without--force.- Compliance quickstart — a new docs/compliance.md walkthrough: template → PII scrub → train with receipt/Annex XI/energy → registry → BOM + attestation → scan/sign/verify → air-gap → model card → CI gate.
Fixed
- GGUF export now actually works on Windows (validated end-to-end against a
locally-built llama.cpp: SmolLM2-135M → q4_0 / q4_k_m / q8_0 / f16 →
soup deploy ollama→ live inference). Four real bugs, each of which independently broke the path:soup export --format ggufcloned llama.cpp into your current directory.SOUP_DIRis the bare name.soup, but the lookup used it relatively rather than anchoring to~like the rest of the codebase — so the canonical~/.soup/llama.cppwas never found and a fresh ~200 MB checkout was dropped into whatever directory you ran from.- The first GGUF export downgraded your PyTorch and broke CUDA. The auto-clone
ran
pip install -r <llama.cpp>/requirements.txtinto your interpreter, and llama.cpp pinstorch~=2.2.1against the CPU wheel index (observed: torch 2.5.1+cu → 2.2.2+cpu, transformers 4.57 → 4.46). Soup now installs only the convert script's extra dependencies, unpinned, and never touches torch. - A correctly-built llama.cpp was not found on Windows. MSVC (like Xcode) is a
multi-config generator and emits
build/bin/Release/llama-quantize.exe; only the flat single-config layout was searched. soup deploy ollamafailed on a relative GGUF path with "pull model manifest: file does not exist" — Ollama resolvesFROMagainst the Modelfile's directory, and Soup writes the Modelfile to a temp dir. The Modelfile now emits an absolute path.
- Model-card injection hardening (affects the pre-existing
soup pushcard too). The## Trainingsection interpolatedbase/task/scheduler/recipeunescaped. SinceSoupConfig.baseandschedulerhave no charset validator, a crafted-but-valid config could smuggle raw HTML — or a backtick breaking out of the surrounding code span — into a card published to the Hub. All values now go through the markdown escaper, which additionally neutralises backticks and strips C0/ESC control bytes.
[0.71.34] - 2026-07-15
Added
soup adapters arithmetic— task-vector algebra over LoRA adapters (add / scale / negate). Apply task arithmetic (arXiv:2212.04089) to LoRA deltas via an expression such as"coder + 0.5*math - toxic", mapping names to adapter dirs with repeatable--adapter name=path. Produces one merged adapter you can serve or merge.- Signed, un-normalized element-wise combine over same-rank adapters; the effective
delta
ΔW = B @ Ascales linearly with each coefficient (negation flips the delta,0.5·halves it) via a √|c| factor split — not thec²a naive sum gives. Mixed-rank inputs are refused with a clear "harmonize rank" message. - Reuses the backdoor-scan gate (refuses a FAIL-scanned input unless
--allow-unscanned) and a same-base-model check (--allow-cross-baseto override). Hand-written expression parser (noeval), cwd-contained/symlink-rejecting paths, exit 0 = ok / 1 = refusal.
- Signed, un-normalized element-wise combine over same-rank adapters; the effective
delta
- LISA — Layerwise Importance Sampled AdamW (arXiv:2403.17919). Full-fine-tuning
quality at LoRA-like memory: every N steps LISA freezes all decoder layers except a
small random set (embeddings + head always trainable). Enable with
training.lisa_enabled: true(+lisa_num_layers,lisa_interval_steps) on atask: sft, transformers, text,quantization: nonerun; mutually exclusive with LoRA features and the other freeze mechanisms. Live on a 4 GB GPU for small models.
[0.71.33] - 2026-07-13
Added
-
soup draft— train and, above all, MEASURE a speculative-decoding draft model.soup draft measure --target <m> --draft <d> --prompts p.jsonlreports a draft's acceptance rate (the fraction of the target's own greedy tokens the draft would have proposed correctly) plus real plain-vs-assisted throughput. This is the honest gate: it tells you whether speculative decoding is worth enabling before you ship it. Exit 0 / 2 (below--min-acceptance, for CI) / 1.soup draft distill --target <tuned> --draft-base <tiny> --data d.jsonl -o draft/distils your target into the tiny base (logit KD via the existingtask: distilltrainer) and emits a dense draft model, ready to load as anassistant_model.soup draft list, plus a local draft registry (~/.soup/drafts.json) thatsoup serve --auto-specconsults before the built-in pairing table — so a draft you trained yourself is picked up automatically.- Draft and target must share a tokenizer; a mismatched pair is refused up front (speculative decoding proposes draft token ids into the target's vocabulary, so a mismatch silently produces garbage rather than failing).
Read the measured results before you use this — see Known limitations below. On the validated pair, distillation did not improve acceptance, and speculative decoding was a net slowdown.
soup draft measureis what tells you that.
Known limitations
- Distilling a draft did not improve its acceptance rate on the validated pair.
Measured on
SmolLM2-360M-Instruct(target) with aSmolLM2-135M-Instructdraft: the stock draft already scored 69.3%, and distilling it moved that to 69.7% after 2 epochs and back to 69.3% after 10 epochs — i.e. no gain beyond noise. A small same-family draft is already near its capacity ceiling for agreeing with the target, and logit KD cannot buy capacity it does not have. Whether distillation materially raises acceptance for a genuinely diverged fine-tune, or a larger target/draft pair, is unproven on a 4 GB box — tracked as a scale issue. - Speculative decoding was a net slowdown on that pair (measured 0.55–0.64×): the
draft's forward pass costs more than the tokens it saves at this size.
soup draft measurereports this truthfully rather than assuming a speedup — which is precisely the point of shipping the measurement. - Acceptance is teacher-forced greedy agreement, the metric the Medusa/EAGLE papers report. It is exact and deterministic, and it is the right number for comparing drafts — but it is not the accepted-token count of a sampling run, which also depends on the rejection-resample cascade.
- Same-tokenizer only. Cross-tokenizer drafts (ULD-aligned + universal assisted decoding) are deferred.
Changed
- PRM-guided GRPO scores completions in a single batched forward.
PRMScorer.__call__now right-pads all completions into one[B, T]tensor (+ attention mask) and runs oneoutput_hidden_statespass instead of one forward per completion, cutting per-step reward latency on non-tiny models. Numerically identical to the per-completion path (parity + mixed-length tests). Closes #298 (#301 by @Ekaanksh-dev).
Fixed
- The hardware-fit gate no longer refuses to train small models.
model_size_from_namedid not understand anM(millions) suffix, soSmolLM2-135Mfell through to the 7B default, was predicted to need ~14 GB of weights, and was blocked. Every draft-sized model hit this.1.7Bwas also being read as7B(the7bmarker matched inside1.7b), whileQwen2.5-7B-Instruct-1Mcorrectly stays 7B (1M is the context, not the parameter count). soup serve --backend vllmno longer force-enablestrust_remote_code. The vLLM path now goes through the same--trust-remote-codedefault-deny gate (and warning panel) as the transformers backend, so serving an untrusted repo never executes its code silently.- Multi-adapter serving (
soup serve --adapters name=path) now actually switches adapters. The named adapters are loaded into the model and selected per request (viaPOST /v1/adapters/activate/{name}or the requestadapterfield); previously every request silently ran the startup model. The base model is served when no adapter is selected. - Vision datasets reject out-of-directory image paths.
llava/sharegpt4vrows are containment-checked againstimage_dir(mirroring the audio loader), so a crafted{"image": "/etc/passwd"}row can no longer read arbitrary local files. soup train --dry-run --gpus Nno longer launches a real multi-GPU run. The accelerate re-exec is skipped under--dry-run. The re-exec also now forwards--minillm-on-policy,--capture-activations, and--capture-prompts(previously dropped on multi-GPU runs).- MLX SFT now builds a real optimizer (
AdamWfrom the configured LR) instead of passingoptimizer=None, which left the model untrained. soup data inspect/preview/searchescape dataset- and Hub-derived text so a stray[/]no longer crashes the command and a crafted[link=…]tag can't render a phishing hyperlink.soup runslist/show escape config-derived fields too.soup infer --task asrhardening — an oversized reference no longer crashes the whole batch after transcription (that row's metric is skipped); an all-skipped run exits non-zero instead of reporting success;--asr-taskis validated upfront; and dataset-derived filenames are control-stripped before printing. ASR training now caps transcript labels to Whisper's decoder limit, warns on >30 s audio, and picks fp16 on pre-Ampere GPUs instead of hardcoding bf16.- Knowledge-distillation KD term aligns with the CE term. The token-level divergence is now computed over causal-shifted positions, so the distillation signal covers exactly the trained tokens (previously off by one).
- Miscellaneous robustness —
load_config_from_stringraisesValueError(notTypeError) on a non-mapping YAML document; the Web UI Bearer-token check is constant-time; andsoup doctor --vscode/ the LR-finder report use the centralised atomic, symlink-rejecting writer.
[0.71.32] - 2026-07-07
Added
- ASR fine-tuning (
task='asr', Whisper) — fine-tune Whisper on your accent or domain, locally. whisper-tiny (39M) / base (74M) train on a 4 GB GPU.- New
AsrTrainerWrapper(HFSeq2SeqTrainer+WhisperProcessor); data rows are{"audio": <path>, "text": <transcript>}under the newdata.format='asr'. Audio decodes via the hardenedload_audio_mono(16 kHz mono, soundfile pre-probe +O_NOFOLLOW+ symlink/size guards); transcripts become decoder labels (pad → −100, decoder-start token stripped). - Optional LoRA on q/v attention projections via
training.asr_lora: true(default full fine-tune);training.asr_language/training.asr_task(transcribe|translate) set the decoder prefix and are persisted to anasr_generation.jsonsidecar so inference restores them. soup infer --task asr— transcribe an{"audio": path[, "text": ref]}JSONL; reports per-row and corpus WER/CER when references are present. Loads a full model or a PEFT/LoRA adapter dir. Flags:--asr-language,--asr-task,--audio-dir(audio paths are cwd/dir-contained; UNC/traversal rejected).- New pure-python
utils/asr_metrics.py— WER / CER /word_accuracy(= 1 − WER, for a higher-is-better ship metric leg) /corpus_wer, with a light Whisper-style text normalizer (no new dependency). - Recipes:
whisper-tiny-asr,whisper-base-asr(live-trainable),whisper-large-v3-asr(parse-only, needs a larger GPU), plussmolvlm-256m-sft(vision). Catalog 138 → 142.
- New
Fixed
model_size_from_namenow knows Whisper checkpoint sizes (tiny…large), so the hardware-fit gate no longer mistakes a 39M whisper-tiny for the 7B default and blocks ASR training on consumer GPUs.
[0.71.31] - 2026-07-06
Added
- Judge-in-the-loop suite — put an LLM judge in the loop across the workflow:
task='online_dpo'— Online DPO training (wraps TRLOnlineDPOTrainer): the model generates two completions per prompt on-policy each step and a judge (a pairwise LLM judge over the existing ollama/openai-compatible backend) OR areward_modelpicks the winner. Config:training.online_dpo_judge: "ollama://model"(or setreward_model— exactly one),online_dpo_loss_type: sigmoid|ipo,online_dpo_max_new_tokens;betareusesdpo_beta. Transformers + text only. Recipe:online-dpo-smollm2-135m. Adapts to the installed TRL: on trl 0.19.x the judge is a swap-debiased pairwise comparison; on trl 1.x (which removed pairwise judges) the sameJudgeEvaluatoris used as a pointwise reward function — a documented per-version behaviour difference.soup data best-of-n— Best-of-N rejection sampling (BOND-lite): sample N completions from--baselocally, a--judgescores each pointwise, and the winner is written as an SFT chat row (with provenance).--emit-pairsalso writes winner-vs-loser DPO pairs.soup data evolve— Evol-Instruct instruction evolution (WizardLM depth / breadth) over an ollama/vllm provider, completing the synthetic-data suite (Magpie / Forge / Persona / evolve).soup ship --task-mode pairwise— a true pairwise judge win-rate as the ship leg-1 task-win (base = 0.5 coin-flip, tuned = its win-rate; swap-debiased), fusing with the catastrophic-forgetting guard into one SHIP / DON'T-SHIP verdict.
Security
soup data best-of-n/evolvewrite outputs via atomicmkstemp+os.replace(re-validated cwd containment), closing the TOCTOU symlink-swap window between the containment check and the write. All judge/provider URLs are SSRF-validated; model loads probetrust_remote_code.
[0.71.30] - 2026-07-05
Added
- PRM-guided GRPO — use a trained Process Reward Model as the per-step
reward inside GRPO (the o1-era process-supervision signal). Set
training.prm_reward: <PRM dir|id>(a model produced bysoup traintask=prm) andtraining.prm_aggregate: min|prod|last; the PRM splits each generated completion into reasoning steps, scores every step with its reward head, and folds the per-step scores into one scalar reward that GRPO optimises. The PRM reward replacesreward_fnand rides the existing reward-shaping + reward-hack-mitigation seam, so the v0.71.26 controller observes it. Cross-validators gatetask='grpo'+backend='transformers'+modality='text'. Default aggregation ismin(weakest-link);prodassumes calibrated[0,1]step scores. - Bundled rollout environments — three pure-Python toy environments
(
soup_cli.envs.calculator/retrieval_qa/guess_number) exposing arollout(prompts)entry point so the liveopenenvGRPO rollout path runs out-of-the-box:training.rollout_backend=openenv+training.rollout_func=soup_cli.envs.calculator:rollout. Three ready-made recipes added (grpo-env-calculator/grpo-env-retrieval-qa/grpo-env-guess-number); catalog 134 → 137.
Fixed
soup train task=prmproducer conformance (surfaced by the v0.71.30 live smoke): the PRM trainer now casts its reward head to the base-model dtype (bf16 CUDA runs previously crashed on the firstcompute_loss), saves the tokenizer alongside the model (so a PRM checkpoint is loadable standalone), and returns the standard trainer-result shape (previously the CLI crashed with aKeyError: 'initial_loss'right after saving).
Notes
- Proof-of-mechanism only: validated on a tiny model (SmolLM2-135M) with a tiny
synthetic PRM and synthetic reward — not a production reward-model claim
(scale ask tracked in #286). The bundled environments are deterministic
single-shot seeders, not interactive multi-turn model-in-the-loop episodes
(the live
openenvcontract passes only prompts). Step split is a newline heuristic; PRM completions are scored one forward pass each.
[0.71.29] - 2026-07-05
Added
soup shrink— depth-prune a model + optional distill-heal ("The Unreasonable Ineffectiveness of the Deeper Layers", arXiv:2403.17887). Ranks decoder layers by the angular distance of the residual stream across a contiguous block over a calibration set, drops the least-important block (first and last layer always protected), optionally heals by distilling the original model into the pruned student, and emits a single dense smaller model plus a one-screen SHIP / DON'T SHIP perplexity verdict.soup shrink --model <id|path> [--drop-ratio 0.25 | --drop-layers N] --calib <calib.jsonl> [--heal <heal.jsonl> --heal-steps N] [--tolerance 0.10] [-o <dir>] [--device cpu] [--attach-to-registry <id>] [--plan-only]. Exit codes: 0 = SHIP, 2 = DON'T SHIP, 1 = error. The heal runs as an isolatedsoup trainsubprocess (LoRA-student logit distillation) and the adapter is fused back so the shipped artifact stays a single dense model. Arch allowlist v1: Llama / Qwen / SmolLM. Validated live on SmolLM2-135M (drop 25 %: 30 -> 22 layers, ppl x2.98 unhealed; drop 4 + heal: ppl x1.35, recovered).
Security
soup shrinkcontains--calib/--heal/--output-dir(and every derived write path:<out>/model,<out>/heal_adapter, the fuse staging dir) under cwd withos.path.realpath+commonpath+ O_NOFOLLOW + symlink rejection, re-validating derived paths right before each write (TOCTOU). The heal subprocess uses an argv list (no shell) with a timeout; its config is schema-validated before spawn; subprocess output is C0/ESC-stripped before it reaches the terminal.--modeldefaultstrust_remote_code=Falsewith a probe + warn.
[0.71.28] - 2026-07-04
Added
- MCP server (
soup mcp serve) — drive Soup from any Model Context Protocol client (Claude Code / Cursor / Cline / Continue) over stdio. No fine-tuning CLI ships an MCP server. Exposes 14 read-only tools as JSON —advise,data_inspect/data_validate/data_score/data_doctor,recipes_search/recipes_show,runs_list/runs_show,registry_list/registry_show,profile,diagnose_evidence,ship_evidence— plus two plan-only mutating tools (train_start,export) gated behind--allow-mutating(they render the exact command that would run; they never execute). The officialmcpSDK is behind a new[mcp]extra (pip install 'soup-cli[mcp]'), lazy-imported so the core CLI stays light. Security: stdio-only (no network listener); every path argument re-enters cwd-containment + symlink rejection; output is control-char sanitized; errors are path-free; string / size / int bounds enforced.
Fixed
- The DPO / IPO / KTO / BCO trainers now apply configured vocabulary expansion
(
data.add_new_tokens/data.new_special_tokens) via the sharedapply_vocab_expansion()helper during setup, consistent with the SFT path — they previously ignored it. Closes #292 (#293 by @CODING-DARSH). - The ORPO / SimPO / GRPO trainers now also apply configured vocabulary
expansion via the shared
apply_vocab_expansion()helper — completing consistent vocab-expansion behavior across every SFT/preference/RL trainer. Closes #294 (#295 by @CODING-DARSH).
[0.71.27] - 2026-07-03
Added
- Fine-tune Doctor — kill the top silent fine-tune failures before a
single training step; no competitor (Unsloth/Axolotl/LlamaFactory) ships
any of these three:
soup data doctor <data> --model <id|path>— chat-template compatibility report over 8 checks:chat_templatepresent,template_renders cleanly, has{% generation %}markers,eos_in_labels(the #1 "model never stops generating" bug — every assistant turn's trained span must actually contain an EOS/EOT token; checks every turn, not just the last),bos_duplication(template + tokenizer both prepending BOS),system_rolesupport (Mistral-style templates reject a leading system turn),unknown_roles, andtruncation_risk(p95 rendered length vsmax_length). Same OK / MINOR / MAJOR taxonomy assoup diagnose; exit 0 = OK/MINOR, exit 2 = MAJOR.--train-on-responses-only/--train-on-messages-with-train-fieldselect the same masking strategysoup trainwould use, so the report and--show-masknever disagree about what's actually trained.soup data doctor ... --show-mask N— render N sample rows with per-token trained/masked colouring through the REAL collator path (answer-only / per-message-train-field / RAFT span-mask) — not a reimplementation — so an assistant-mask bug is visible instantly.soup data lint <data>— preference-data linter for dpo/orpo/simpo/ipo/bco/kto:length_bias(chosen systematically longer than rejected — the #1 silent DPO degradation, reported as a Cohen's d effect size),label_imbalance(KTO desirable:undesirable ratio),near_duplicates(MinHash/LSH, reuses thesoup data dedupkernel),identical_pairs(chosen == rejected — zero preference signal), andprompt_leak(the prompt echoed verbatim inside the completion — a common synthetic-data pipeline bug). Optional--modelfor exact token-length bias (default: word count).- Validated live against the real
HuggingFaceTB/SmolLM2-135M-Instructtokenizer on Windows + RTX 3050 — this smoke pass found and fixed two genuine bugs beyond what synthetic fixtures alone caught: an EOS check that required the EOS token to be the literal last trained token (real templates often have a trailing formatting token after the closing tag that stays inside the trained span), and two call sites that only caught(ValueError, TypeError)around a tokenizer'sapply_chat_templatewhen a real Jinjaraise_exception()(Mistral-style no-system-role guard) raisesjinja2.exceptions.TemplateError.
Fixed
- Harden
commands/diagnose.py's--evidenceloader against a TOCTOU symlink swap: opens withO_NOFOLLOWand size-checks the open fd viaos.fstatinstead ofos.path.getsizeon the path before the open — backports the hardened loader shipped forsoup shipin v0.71.25 (closes v0.71.25 known-limitation (4)). - Harden judge-model URL validation against a hostname prefix bypass
(
http://localhost.attacker.com) —GateTask._valid_judge_url/_parse_judge_urlnow useurllib.parse.urlparse+ hostname checks instead ofstartswith. Closes #283 (#288 by @CODING-DARSH). SFTTrainerWrappernow applies configured vocabulary expansion (data.add_new_tokens/data.new_special_tokens) and resizes the model embeddings during initialization — previously these fields were accepted by the schema but silently ignored. Closes #289 (#287 by @CODING-DARSH).- Vision and audio SFT paths now apply that same configured vocabulary
expansion (
data.add_new_tokens/data.new_special_tokens) via the sharedapply_vocab_expansion()helper, consistent with the text SFT path — they previously ignored it. Closes #290 (#291 by @CODING-DARSH).
Security
soup data doctorstrips C0 control characters (keeping tab/newline/CR) from dataset-derived content before it reaches the terminal — Rich'smarkup.escape()only neutralises[...]tag syntax, not raw escape sequences, so an untrusted training row (e.g. an unknownrolefield, or--show-mask's decoded token text on a byte-level BPE tokenizer) could otherwise carry a literal ESC byte through to the terminal (title-bar / OSC-8 link spoofing, or obscuring a MAJOR verdict via cursor tricks).--outputJSON is unaffected (json.dumpsalready escapes control characters).
[0.71.26] - 2026-07-01
Added
- Closed-loop reward-hacking auto-mitigation. The trainer now detects
reward hacking mid-run and self-corrects — instead of only halting. Set
training.reward_hack_mitigation(orsoup train --reward-hack-mitigation) to one of four modes on a GRPO/PPO run (requiresreward_hack_detector):log_only— instrument only: append a per-stepmitigation_log.jsonl(drop_pct, verdict, reward mean/std, completion-length trend, repetition) and never touch training.kl_control— a reversible bang-bang + hysteresis controller: when the hacking signal trips, raise the KL coefficient β (geometric, clamped to[floor, ceil], never crossing 0); relax it when the signal recovers. Dwell + release-patience prevent flapping; a multi-signal vote combines the detector drop with a length-trend and repetition signal.pid_lagrangian— a PID-Lagrangian controller (Stooke et al.) that holds the hacking signal at a target, plus an escalation ladder: raise β → roll back to the last-good RL checkpoint → early-stop.- Anti-gaming hardening: per-signal EMA/median smoothing, conservative-on-disagreement voting, a reward-distribution-drift guard, and optional bounded reward shaping on the gamed proxy (length / repetition / sentinel). A plain-English give-up explanation is logged on early-stop.
- Proof-of-mechanism only (see Known Limitations): validated on SmolLM2-135M + a synthetic length-hacking task on a single RTX 3050 — all four stages pass live, including a real mid-run rollback. PPO ships BETA (mechanism unit-tested; the on-GPU proof is GRPO-only).
- Ready-made
qwen2.5-coder-7b-sftrecipe forQwen/Qwen2.5-Coder-7B-Instruct(catalog 133 → 134) (#285 by @Deadpool2000).
Security
RLCheckpointCallback.restore_checkpoint/save_checkpointrefuse a symlinkedoptimizer.pt—torch.load(weights_only=False)on an attacker-placed symlink in a shared checkpoint dir was an RCE vector.- Bool-before-int/float guards on every new
reward_hack_*numeric field;reward_hack_signalsbounded (max_length=4); the mitigation log writer is cwd-contained with symlink-reject-on-rotate and secret redaction.
[0.71.25] - 2026-06-27
Added
soup ship— the SHIP / DON'T-SHIP verdict. After fine-tuning, answer one question: did the model get better, or did I break it?soup shipfuses two checks into a single binary decision — leg 1: the task metric strictly improved (base → tuned); AND leg 2: no general benchmark regressed past a forgetting threshold (default 0.05 absolute points). It SHIPs only when both hold — otherwise DON'T SHIP, even if the task metric looks great. The output is a one-screen verdict + the reason, with CI-gateable exit codes (0 = SHIP, 2 = DON'T SHIP, 1 = runtime error).- Leg-1 modes:
--task-mode metric(eval accuracy) orjudge_score(LLM-as-a-judge); pairwise win-rate is planned for a later release. - Leg-2 suite: built-in mini benchmarks by default (offline, CPU), or
--general-suite <names>to route lm-eval benchmarks;--baseline registry://… | file.jsonsupplies recorded base scores. --evidence ev.jsondecides offline from pre-computed scores (no model load);--output verdict.jsonpersists the machine-readable verdict.
- Leg-1 modes:
- Friendlier error messages: the CUDA-OOM hint now also suggests
gradient_checkpointingand4bitquantization, plus new mappings for Hugging Face gated repos (huggingface-cli login/HF_TOKEN) andtrust_remote_codeerrors. Closes #272 (#282 by @Akshaya-reddy18).
Security
soup shipinput hardening:--evidenceis opened withO_NOFOLLOW+ an fstat size cap (16 MiB) under cwd containment;--task-evalis cwd-contained and symlink-rejected;--judge-modelis validated by scheme/host viaurlparse(blocks thehttp://localhost.attacker.comprefix bypass); lm-eval model ids reject,/=injection;--general-suiteis bounded (≤ 50 names, ≤ 256 chars each).
[0.71.24] - 2026-06-21
Added
- 2026 model-family recipe expansion (catalog 116 → 133). 17 new ready-made
SFT recipes for the open-weight models released Feb–Jun 2026, each with its
Hugging Face repo-ID verified to resolve:
- Qwen 3.5 (Apache-2.0):
qwen3.5-0.8b-sft,qwen3.5-2b-sft,qwen3.5-4b-sft,qwen3.5-9b-sft,qwen3.5-27b-sft, and theqwen3.5-35b-a3b-sft/qwen3.5-122b-a10b-sft/qwen3.5-397b-a17b-sftMoE sizes. - Qwen 3.6 (Apache-2.0):
qwen3.6-27b-sft,qwen3.6-35b-a3b-sft. - DeepSeek-V4 (MIT):
deepseek-v4-flash-sft,deepseek-v4-pro-sft. - GLM (MIT):
glm-5.1-sft. - Kimi (Modified MIT):
kimi-k2.5-sft,kimi-k2.6-sft. - MiniMax (MiniMax Community License — commercial use needs a separate
agreement):
minimax-m3-sft. - Mistral Large 3 (Apache-2.0, 675B/41B-active multimodal MoE):
mistral-large-3-sft.
- Qwen 3.5 (Apache-2.0):
- Unit-test coverage for the
warmup.pyauto-warmup-steps helper (#274 by @shatakshi-1404).
Fixed
- Stale recipe repo-ID:
glm-5-sftnow points atzai-org/GLM-5(the org migrated fromTHUDM).
[0.71.23] - 2026-06-12
Added
- Native Spectrum targeted training (closes #266). A new
soup spectrum scanreads a model's.safetensorsshards one tensor at a time (no model load — peak RAM is the largest single weight matrix), computes a singular-value SNR per weight matrix with a Marchenko-Pastur noise threshold (arXiv:2406.06623), ranks layers within each module-type group and prints the top--top-percentas a ready-to-pastetraining.unfrozen_parametersYAML block. This lets you scan even a very large model's layer SNR on a CPU box and then full-fine-tune only the high-signal layers.soup spectrum scan --model <id|path> --top-percent 50 [--modules mlp,attn|all] [--output patch.yaml]— SNR table + the YAML patch; results cache at~/.soup/spectrum/<slug>.json(override viaSOUP_SPECTRUM_CACHE_DIR).- New schema field
training.unfrozen_parameters: list[str]— regex patterns of parameter names to keep trainable; the SFT trainer freezes every parameter then unfreezes the matched set (full fine-tuning, LoRA off). Mutually exclusive with LoRA features /freeze_layers/freeze_ratio/train_router_only/expand_layers; requirestask=sft,backend=transformers,modality=text, andquantization=none. - The SNR kernel is pure-numpy and transpose-invariant (singular values are
identical for
WandW.T); GPT-2Conv1Dnaming (c_attn/c_fc/c_proj) is recognised alongside Llama-style names. - The existing
spectrumtrainer-plugin wrapper is untouched (back-compat). LISA (per-step layer sampling) is tracked separately in #267.
Security
soup spectrum scanvalidatesunfrozen_parameterspatterns at parse time: rejects nested-unbounded-quantifier regexes (ReDoS), null bytes, empties, and caps count (50k) and length (512). Hub downloads route through the SSRF-hardened, namespace-pinnedhubs.snapshot_download; symlinked shards and matrices above a 2^31-element SVD cap are skipped;--outputstays under cwd.
[0.71.22] - 2026-06-10
Added
- Perf & measure polish — a 4-issue patch tightening four live paths from
the recent BETA lifts. Pure code, validated on Windows + RTX 3050.
- MiniLLM on-policy KV-cache (closes #263). The on-policy distillation
rollout (
soup trainwithtraining.minillm_on_policy: true) now threadspast_key_valuesso each step forwards only the new token instead of re-feeding the whole prefix — resolving the O(L²) per-step cost from v0.71.18. A LoRA student (the common distill case) activates the cache too: the new_supports_kv_cacheprobe unwraps the PEFT model viaget_base_model()before deciding. The teacher is always cached; the student cache respects the retained autograd graph and degrades gracefully if a model returns no cache mid-loop. soup serve --moleKV-cache (closes #262). Each of the N task adapters in a served MoLE now keeps its own KV cache in lockstep, created fresh pergenerate()call (never stored on the instance, so there is no cross-request leak). Top-k zero-weight adapters are still skipped, and the output is byte-identical to the no-cache path on a real MoLE.- Deploy-autopilot live measure factories (closes #143).
soup deploy autopilot --measureships a first-party transformers loader factory (lazy import, per-candidate quant config via the Quant Menu loader;before= base,after= quantised) replacing the inject-only test hooks. The baseline is now scored once and the whole candidate list is pre-validated up front, so a typo in--measure-candidatesraises before any model load instead of burning N live loads or doubling peak VRAM. - Live-codec TTS via SNAC, partial (#265-partial). The live-codec
encode path (
data.format='audio') is validated for Orpheus:load_audio_mononow probessoundfile.info(duration + byte cap) beforesoundfile.read(no multi-GB decode into RAM) and reads through anO_NOFOLLOWfile descriptor; a real SNAC-backed encode of a 24 kHz wav produced 42 Orpheus codec tokens.
- MiniLLM on-policy KV-cache (closes #263). The on-policy distillation
rollout (
Fixed
- MiniLLM on-policy KV-cache was silently disabled for LoRA students (the
PEFT wrapper hid the base model's
past_key_valuessupport) — now probed viaget_base_model(). - Deploy-measure no longer re-scores the baseline once per candidate or burns live model loads on a bad candidate (per-candidate validation moved up front).
load_audio_monocapped audio duration only after decoding into RAM — the cap is now checked fromsoundfile.infobefore reading.
Known limitations
- KV-cache correctness is validated (cache == no-cache equality on real tiny artifacts) but large-model throughput gains were not measured on the 4 GB dev box.
- #265 stays open — the live-codec
data.format='audio'SNAC encode path is validated for Orpheus only; the other four TTS families keep their per-family codec dependency gate. - The deploy-measure first-party factory's real quantized (bitsandbytes 4-bit) load is CUDA + bitsandbytes-gated; on Windows / no-bnb the injected test seams are the validated path.
- The MoLE serve KV-cache assumes single-sequence (
B == 1) decode.
[0.71.21] - 2026-06-10
Added
- Precision & rollout lift (BETA, hw-gated) — lifts five deferred
NotImplementedErrorstubs to live code.- FP8 attention + NVFP4 (closes #141).
training.fp8_attention: truenow converts the model's attention projections (q/k/v/o + fused qkv variants) to FP8 training modules via torchao'sconvert_to_float8_trainingwith an attention-onlymodule_filter_fn(Hopper SM ≥ 9.0 gate);training.nvfp4: truequantises via torchao'sNVFP4Config(Blackwell SM ≥ 10.0 gate). Both are wired into the v0.28 speed/memory pipeline and degrade to a visible yellow advisory when the gate fires — a conversion failing partway raises an honest "model may be PARTIALLY converted" error rather than silently training on a half-converted model. - vLLM sleep mode (closes #124).
training.vllm_sleep_mode: trueis live:create_vllm_engine(sleep_mode=True)setsAsyncEngineArgs.enable_sleep_mode(vLLM ≥ 0.7 gate with a friendly upgrade message), the newvllm_sleep_cycle(engine, level=1|2)context manager wraps the optimisation step (wake infinally), and the GRPO trainer threads the flag into TRL'sGRPOConfigwhen the installed TRL exposes the hook (advisory otherwise). - Multi-turn agent rollout launchers (closes #125).
soup trainwithtask: grpo+training.rollout_backend: openenv+training.rollout_func: my_module:fnnow runs a LIVE rollout: the resolver imports the operator's callable (same trusted-code policy asdata.prompt_strategy), feeds it the dataset prompts as seeds, and the returned{prompt, answer?}rows replace the prompt dataset. Rows are normalised (extra keys stripped, message-list prompts deep-copied, non-string answers rejected loudly).art/ruler/nemo_gymraise a friendly ImportError when the backend package is missing and an honest BETA gate when present (injectable_EXTERNAL_ROLLOUT_RUNNERSseam). Validated by a real GRPO + openenv rollout train on SmolLM2-135M. - Apple-adapter conversion (closes #228).
soup apple-adapteris live forhf-to-mlx/mlx-to-hf: PEFT LoRA safetensors ↔ mlx-lm adapters with both matrices transposed (lora_A [r,in]↔lora_a [in,r]), bf16 sources upcast via the torch loader,adapters.safetensors+num_layersemitted for mlx-lm'sload_adapters, rank/alpha/dropout carried through, legacyadapters.npzstill read, optional v0.60 Merkle-root signing. The*-to-appledirections stay upstream-gated (no published FoundationModels adapter spec). Validated by a real bf16 PEFT adapter round-tripping with numeric equality. - Llama-4 expert delinearization (closes #97).
soup delinearize-llama4now runs a live torch runtime: fused 2-D expert tensors[E*dim_in, dim_out]reshape to 3-D[E, dim_in, dim_out](expert count fromconfig.jsonor--num-experts), other tensors pass through, JSON sidecars are copied, writes are atomic.--plan-onlykeeps the old render-and-exit flow.
- FP8 attention + NVFP4 (closes #141).
Fixed
safetensors.numpy.savesilently mangles non-contiguous (transposed) arrays — the apple-adapter writer now makes every array C-contiguous first (caught by the new round-trip assertions).
Known limitations
- fp8_attention / nvfp4 / vllm_sleep_mode are BETA hardware-gated — the
converters and gates ship validated via capability probes and fake-module
dispatch tests, but end-to-end runs need a Hopper/Blackwell GPU + torchao
(or vLLM ≥ 0.7), none of which exist on the maintainer's RTX 3050 /
Windows box. The
art/ruler/nemo_gymrollout adapters are honestly BETA-gated until validated against the upstream packages.
[0.71.20] - 2026-06-09
Added
- Modality II trainers — TTS / BitNet / MoE expert quant (BETA, hw-gated)
— lifts three v0.52.0 schema-only
NotImplementedErrorstubs to real code.- TTS fine-tuning (closes #131).
soup trainwithtask='tts'+modality='audio_out'now routes to a liveTTSTrainerWrapper. TTS families (Orpheus / Sesame-CSM / Llasa / Spark / Oute) are decoder language models, so a TTS fine-tune is next-token cross-entropy over interleaved[text][audio-codec-token]chat sequences — the wrapper reuses the SFT path and adds per-family emotion-control templating (Orpheus / Oute) and registration of operator-supplied codec special tokens (data.new_special_tokens) with an embedding resize. The pre-encoded chat workflow (codec tokens produced offline, then trained withdata.format=chat) is the live, validated path; the live-codec workflow (data.format='audio', encode raw audio at train time) needs the family's heavyweight codec dependency (SNAC / BiCodec / XCodec2 / …) and is hardware/dependency-gated with a friendly per-familyRuntimeError. Verified end-to-end on SmolLM2-135M-Instruct. - BitNet 1.58-bit (closes #134).
build_bitnet_trainerreturns a liveBitNetTrainerWrapperthat gates on the upstreamonebitllmspackage (absent → friendlyRuntimeErrornaming it).soup export --format bitnet | tq1_0now runs a real llama.cpp TQ1_0 ternary export (reuses the v0.53.1 gguf convert→quantize pipeline) instead of the deferred panel; it requires a built llama.cpp toolchain (friendlyFileNotFoundErrorwhen absent). - MoE expert quant + router-only training (closes #136).
apply_moe_expert_quantdetects fused-MoE expertnn.Linearblocks and replaces them with bitsandbytesLinear4bit(nf4) /Linear8bitLt(int8_rowwise), leaving attention + the router in full precision; it runs beforeget_peft_model(QLoRA-on-experts) so PEFT attaches to the quantized base.train_router_onlyfreezes every expert and keeps the gating router trainable, applied after LoRA. CUDA-gated (friendlyRuntimeErrorwhen bitsandbytes/CUDA absent). Validated live on an RTX 3050: 8 expert Linears → 8Linear4bitwith dequant error 0.0155 vs source (weights genuinely carried), router-only freeze, and device-aware placement.
- TTS fine-tuning (closes #131).
Known limitations
- The TTS live-codec workflow, BitNet 1.58 training (
onebitllms), and BitNet GGUF export (llama.cpp) are hardware/dependency-gated — the friendly gates ship and the plumbing is validated, but the end-to-end runs against real TTS models + audio codecs / a BitNet base + onebitllms / a built llama.cpp toolchain stay open infra-blocked items on the maintainer's RTX 3050 / Windows box.
[0.71.19] - 2026-06-09
Added
- Quant Menu for vision / audio modality (closes #81). The Quant Menu
(
gptq/awq/hqq:Nbit/aqlm/eetq/mxfp4/fp8) was rejected by the config modality gate formodality in {vision, audio}— those paths carried inlineBitsAndBytesConfigblocks that handled only4bit/8bit. v0.71.19 drops the gate (the mlx-backend gate is retained) and threads the unifiedbuild_quantization_config_for_loaderthrough_setup_vision_transformers/_setup_audio_transformers, so multi-modal SFT can train a LoRA on top of any pre-quantized base. The4bit/8bitconfig shapes are byte-for-byte the same as the old inline blocks;mxfp4still routes throughprepare_model_for_kbit_training. Verified: the unified loader returns the right config object for every format on both modalities, and_setup_vision_transformersthreads aGPTQConfigintoAutoModelForVision2Seq.from_pretrained.
Fixed
- Multipack DataLoader sharding under FSDP / DeepSpeed ZeRO / DDP (closes
#80). The multipack
get_train_dataloaderoverride built a rawDataLoaderand returned it directly, so under distribution every rank trained on the same packed bins (no data sharding). It now routes the loader throughaccelerator.prepare(...)whennum_processes > 1— exactly what HF Trainer's ownget_train_dataloaderdoes — so accelerate'sBatchSamplerShardround-robins whole bins across ranks (preserving the FFD packing) and equalises per-rank batch counts. The single-process path is unchanged (byte-for-byte the validated v0.40.4 raw-DataLoader behaviour). Verified live: a single-GPU multipack SFT on SmolLM2-135M trains end-to-end (RTX 3050). Full multi-GPU validation remains a QA item (no multi-GPU box); the distributed routing is mocked-tested.
[0.71.18] - 2026-06-08
Added
- MiniLLM true on-policy rollout (closes #257).
training.minillm_on_policy: true(withminillm_enabled: true) replaces the offline distribution blend with the real on-policy procedure of Gu et al. 2024 §3.1: each step samples a fresh autoregressive rollout from the per-token mixtureratio·teacher + (1-ratio)·student, then accumulates the length-normalised reverse-KLKL(student || teacher)on the full distributions (differentiable w.r.t. the student only; sampled tokens are detached). Newtraining.minillm_rollout_lengthknob ([1, 512]; auto-derivesmin(max_length, 32)when unset — the loop re-forwards the full prefix each step, so keep it small). Verified live: on-policy distill on tiny-gpt2 (student + frozen teacher), finite loss, end-to-end train. - Cross-tokenizer ULD with token-sequence alignment (closes #258). New
training.uld_strategy: wasserstein_alignedhandles fully-disjoint tokenizers (not just a vocab-size mismatch): per batch element the student and teacher token sequences are aligned over their decoded character spans (offset-overlap when both decode to the same text, difflib Ratcliff-Obershelp char matching otherwise), the teacher logits are mean-pooled onto the student positions, and the existing sorted-Wasserstein-1 surrogate is applied. Verified live: aligned distill with a GPT-2 BPE student + a Llama SentencePiece teacher, finite loss, end-to-end train. soup agent eval --sandbox(closes #110). Each heuristic-passing tool-call prediction is now executed against a generated mock of the endpoint in the v0.25.0 RLVRcode_execsandbox and classified into ok / tool_error / timeout / arg_error. The endpoint path, its required path params, and the predicted arguments are base64-embedded as data (no code interpolation). Strong isolation (RLIMIT / namespaces / sandbox-exec) is POSIX-only; on Windows the subprocess + 5 s timeout + 10 KB output cap + network guard still apply (a friendly reduced-isolation advisory is printed). Verified live on Windows: 4-prediction scorecard (ok=1 / tool_error=1 / arg_error=2 / timeout=0).soup train --cloud modal(closes #16). Render a self-contained Modal.com app fromsoup.yamlfor serverless GPU training when you have no local GPU. The config YAML is base64-embedded as data (no interpolation, no secrets); the--gputype (t4 / l4 / a10g / a100 / a100-80gb / l40s / h100) is validated against a closed allowlist. Default is plan-only: write the stub + print themodal runcommand.--cloud-submitattempts a live submit gated on a Modal token (modal setup/MODAL_TOKEN_ID+MODAL_TOKEN_SECRET). New[modal]extra (pip install 'soup-cli[modal]'; only needed for live submit — plan-only render needs no dependency). Verified live: real stub rendered, exit 0.
[0.71.17] - 2026-06-08
Added
- Serve-time MoLE (closes #259). A
task='moe_lora_routing'run now writes a self-describingmole_manifest.jsonnext tomole_gate.pt, andsoup serve --mole <dir>loads the base + N frozen task LoRAs + the trained gate and blends them per token at decode time (custom blend loop — non-streaming + streaming).--molerequires--backend transformersand is mutually exclusive with--bank/--steer/--adapters/--speculative-decoding. The base model comes from--base(or the manifest when unset). Verified live on SmolLM2-135M (2 task adapters, real generation + SSE streaming). - Per-request multi-tenant vector banks (closes #260).
soup serve --banknow resolves the active VeRA/VB-LoRA user per request via acontextvars.ContextVar, so concurrent requests on a threaded server never race on shared instance state. The streaming path re-selects the user inside the generator's own context. Verified live: twoX-User-Idheaders produce distinct steered outputs, an absent / unknown id self-clears to the clean baseline (no cross-request leak), and a repeated user is deterministic. - Epoch-aware RAFT document shuffle (closes #253).
data.raft_epoch_shuffle: truere-permutes the golden + distractor documents each training epoch (per-epoch salt) so the model can't latch onto one fixed citation slot.epoch=0reproduces the legacy single-permutation order exactly. Verified live on a 2-epoch SmolLM2-135M RAFT run. soup diagnose --citation-style/--shuffle-seed(closes #254). The live citation failure-mode probe now accepts the citation style (bracket / inline / footnote) and the RAFT shuffle seed so the golden[doc-N]ids line up with what the model saw at train time. Verified live (rows=6, mean_recall=1.000).
Fixed
- MoLE
train()now returns theinitial_loss/final_loss/total_steps/duration_secs/durationkeys the generic train handler reads, sosoup train task=moe_lora_routingcompletes cleanly (previously raisedKeyError: 'initial_loss'after writing the gate). Surfaced by the #259 smoke.
[0.71.16] - 2026-06-07
Added
- Covariance-preconditioned ROME via
--cov-corpus(closes #250).soup edit set --method rome --cov-corpus <jsonl|txt>now estimates the key covarianceC = E[k kᵀ] + λIover a stats corpus and uses the preconditioned updateu = C⁻¹ k*instead of the covariance-freeC = Ipath — the genuine ROME closed form, which spreads the rank-1 update mass to reduce collateral interference with other facts. Falls back toC = Iwhen no corpus is given. The exact post-conditiondown(k*) += deltais preserved either way. The corpus loader is cwd-contained, symlink-rejected (O_NOFOLLOW + raw-path lstat), and size/line-capped;--cov-corpusis rejected (fail-loud) for any method other thanrome. Verified on realgpt2(prob 0.005 → 0.9997) and SmolLM2-135M. - GPT-2 (
transformer.h/mlp.c_proj) support in the edit kernels (closes #251). ROME / MEMIT / AlphaEdit now edit GPT-2-family models, not just Llama-family. TheConv1Dweight layout ([in, out], transposed relative tonn.Linear's[out, in]) gets a transpose-aware rank-1 update, AlphaEdit null-space projection, and MEMIT band dim-check. PEFT-wrapped GPT-2 / Llama models are unwrapped viaget_base_model. Verified end-to-end on realgpt2. - Mixtral joins the LongLoRA architecture allowlist (closes #147). A bare
mistraltoken does not appear inmixtral(m-i-x vs m-i-s), so the existingis_mistral_modeldetector excluded the MoE variant. A dedicatedis_mixtral_modelhelper +MixtralAttentionentry in the S² forward-override regex +_SEPARATE_QKV_FAMILIESnow cover Mixtral-8x7B / 8x22B (the attention is the standard separate-QKV shell; the MoE lives in the MLP).
Fixed
- Atomic
EditGovernoredit-count increment (closes #252). Two concurrentsoup edit setruns on the same base model could lose an increment: each read the persisted count, added locally, and the last writer clobbered the first.save_statenow re-reads the persisted count INSIDE the cross-process lock and merges this run's delta (edit_count − persisted_baseline), mirroring the v0.60.0namespace_pinpattern. Verified: two governors recording 3 + 2 edits from the same baseline persist a merged 5 (not a clobbered 2 or a naive +1).
Notes
- Test count: 13511 → 13595 (+84 net; +81 in
tests/test_v07116.py).
[0.71.15] - 2026-06-07
Fixed
- Iterative-DPO config render bug (closes #261).
soup iterative-dpo's default per-round trainer renderedoutput: {dir: ...}(a mapping), whichSoupConfig.output(a plain string) rejected — so the spawnedsoup trainsubprocess failed at config validation. Now rendersoutput: <str>, mirroring the v0.71.13 #229local-rlfix. A regression test captures the rendered YAML and validates it viaload_config_from_string; verified end-to-end with a realsoup trainround on SmolLM2-135M.
Changed
- CMA-ES merge loads the base model once (closes #246).
soup adapters merge --strategy cmaespreviously reloaded the (multi-GB) base model into a fresh PEFT wrapper on every candidate in the population. The default scorer now loads the base once and reuses it across the wholepopulation × generationsloop — each candidate only loads its small merged LoRA, applies it, generates, and unloads it. Verified on SmolLM2-135M: the base loads exactly once across N candidates. soup loopbudget gate now estimates real cost (closes #245). The pre-wired loop's per-iteration cost estimate was a hard0.0placeholder, so the dollar budget gate never tripped. It now wires v0.34run_cost. estimate_run_cost_usdoff the most-recent completed run's GPU + duration (the best forward signal for a repeating loop). Falls back to0.0on the first iteration / a CPU / unpriced GPU; never crashes the daemon.--diagnose-gateis multi-node aware (closes #170). The post-training diagnose gate (and the--annex-xi/--repro-receipt/ capture hooks) fired onLOCAL_RANK==0, so a shared-filesystem multi-node run ran them once per node. They now gate on the global chief (RANK==0whenRANKis set, elseLOCAL_RANK==0) — once per cluster.
Added
soup train --track-energy --energy-out <path>(closes #244) persists the measured energy/CO2 reading as JSON sosoup bom emit --energy <path>(the v0.71.3 #256 consumer) can attach it to an ML-BOM. Atomic + cwd-contained + symlink-rejected. Completes the train → BOM energy hand-off.
[0.71.14] - 2026-06-05
Added
- Live FSDP shard consolidation (closes #96).
soup merge-sharded-fsdp-weightslifts the v0.44.0 plan-only stub: it now streams eachpytorch_model_fsdp_*.binshard viatorch.load(weights_only=True)(no arbitrary pickle exec), unions the per-rank parameter fragments into one state-dict, and writes a single.safetensorsatomically. Memory-friendly (one shard loaded at a time). New--plan-onlyflag prints the plan without writing. Single-process — no multi-GPU needed to MERGE. (Per-rank disjoint-parameter / FULL_STATE_DICT shards; DCP sharded-tensor reconstruction is out of scope — useaccelerate merge-weightsfor those.) - Live
kv_cache_typewiring on the transformers serve backend (closes #140).soup serve --kv-cache-type q8_0 | bf16 | f16 | fp8lifts the v0.53.1apply_kv_cache_typeNotImplementedErrorstub:bf16/f16load the model in that dtype (the KV cache inherits it);q8_0routes an 8-bit HQQ quantized KV cache throughmodel.generate(needspip install hqq);fp8raises a friendly runtime error (vLLM + Hopper-only — the transformers backend has no fp8 KV path). vLLM / SGLang KV-cache-dtype routing stays in the infra-blocked tail. - ONNX export QA verified (closes #71) —
soup export --format onnxexercised end-to-end on a tiny model: export exits 0,model.onnxloads in ONNX Runtime withinput_idspresent, and a forward pass produces a real output. Recorded intests/qa/v07114_qa.md.
Notes
- GGUF export (#70), AWQ/GPTQ export (#72), the CUDA + llama.cpp QA doc (#144),
HF Hub push/Spaces deploy (#74), and the Community-QA tracking meta-issue (#79)
remain open with
infra-blockedlabels — they need a built llama.cpp toolchain,autoawq/auto-gptqWindows wheels, or HF credentials the QA box lacks. Seetests/qa/v07114_qa.md.
[0.71.13] - 2026-06-04
Added
- Prompt-compile family — live wiring (closes #225, #226, #227, #229). Four
soupcommands that shipped as deferred-stubNotImplementedErrorin v0.68.0 are now real, validated end-to-end (real DPO train on SmolLM2-135M + real Ollama teacher distillation on RTX 3050). soup local-rl trainruns a real nightly DPO/KTO/ORPO train (#229).--onceharvests the latest thumbs-up/down DPO pairs from the local-RL SQLite and trains them via asoup trainsubprocess (argv list, no shell); astatetable trackslast_train_atso a re-run with no new feedback skips, and a run with fewer than--min-pairs(default 10) skips. Without--onceit renders a systemd.service/.timer+ launchd.plistscheduler scaffold into--scheduler-dirfor the user to install. New flags:--once,--min-pairs,--output/-o,--scheduler-dir,--hour,--minute.soup distill-promptprepares a real distillation dataset (#226). For each prompt in the traces JSONL the teacher is called once via the v0.20 provider helpers (Ollama / Anthropic / vLLM);sft/klemit{messages:[user, assistant=teacher]}andpreferenceemits{prompt, chosen=teacher, rejected=student}. New flags:--provider,--base-url,--temperature,--max-rows.soup compileruns DSPy / GEPA / TextGrad prompt-program optimisation (#225) andsoup compile-toolsruns the TextGrad / GEPA tool-schema optimiser (#227), both lazy-importing the optimiser libraries behind the new[compile]extra (pip install 'soup-cli[compile]') with a friendlyImportErrornaming the extra when absent.--plan-onlystill renders the plan and exits 0.
Security
- systemd / launchd injection defence (#229).
local-rland the scheduler renderers reject\n/\rin the model id and shell-quote everyExecStartargument, so a crafted model id cannot inject extra unit directives.
Fixed
local-rltrain config renderedoutputas a mapping (#229). The nightlysoup trainYAML now emitsoutput: <dir>(a plain string the schema accepts) instead ofoutput: {dir: <dir>}; a regression test validates the rendered config againstSoupConfig.
[0.71.12] - 2026-06-04
Added
- Architecture + distillation + adapter-training — live wiring (closes #145, #146, #148, #158, #84, #221, #222). Seven surfaces that shipped schema-only in earlier releases are now real, validated end-to-end on tiny models (SmolLM2-135M / a locally-built tiny Llama).
- Sequence-level knowledge distillation is live (#145).
task: distillnow acceptsdistill_mode: token|sequence; sequence mode trains the student on the teacher's generated continuations (cross-tokenizer-friendly hard-label KD) instead of per-token logit matching.sequencemode is mutually exclusive with the v0.70 cross-tokenizer ULD logit path. - Classifier LoRA is live (#146).
task: classifier|reranker|cross_encodernow attaches a LoRA adapter to the sequence-classification head whenlorais configured, so a frozen encoder + small adapter can be trained instead of the full model. - LLaMA Pro block expansion is per-architecture (#148).
expand_layersnow interleaves zero-initialised identity blocks for Llama / Qwen / Mistral decoder stacks (was Llama-shaped only), withfreeze_trainable_layersfreezing the original blocks so only the new ones train. - LongLoRA S² shifted-sparse attention is live (#158).
use_longlora: truenow installs the shifted-sparse-attention forward override on the Q/K projections (Llama / Mistral / Qwen / Phi), restoring the patched forwards on context exit. - Mixture-of-Depths is live (#84).
use_mod: trueattaches a per-layer top-k token router (mod_capacity_factor) so only a subset of tokens receive each block's residual update. Architecture allowlist: Llama / Qwen / Mistral; unsupported bases warn and skip. - VeRA / VB-LoRA multi-tenant serving is live (#221).
soup serve --bank <bank.json> [--bank-strength S]reconstructs the shared projection + per-user scaling vectors and installs a decode-time forward hook; the active user is selected per request via theX-User-Idheader (an unknown/absent id is a zero-delta no-op, so there is no cross-request leak). Serves N personas at ~KB-per-user instead of a full LoRA each. - MoLE per-token adapter routing is live (#222).
task: moe_lora_routingwithmole_task_adapters: [...]trains a per-token gating network that blends N frozen task LoRAs (mole_top_k/mole_temperature); only the router trains. The gate is saved asmole_gate.ptalongside the run.
Changed
apply_bank_to_serve(#221) andbuild_gating_kernel(#222) now return live objects (aLoadedVectorBankand atorch.nn.Modulerouter) instead of the v0.67.0 deferred-stubNotImplementedError.
[0.71.11] - 2026-06-04
Added
- GRPO / RL callbacks — live wiring (closes #235, #236, #237, #238, #239, #240, #159, #160). The reward-hacking, cross-tokenizer distillation, MiniLLM, mid-epoch RL checkpoint, iterative-DPO and echo-trap surfaces that shipped schema-only in v0.70.0 are now real, validated end-to-end on SmolLM2-135M.
- Reward-hacking detector is live (#235).
--reward-hack-detector info_rm|rm_ensemblenow installs a GRPOTrainerCallbackthat reads the per-step rewards (via a shared, thread-safe reward-fn capture buffer), computes an InfoRM cluster-separation drop (info_rm) or RM-ensemble divergence (rm_ensemble), classifies OK/WARN/HACK, logs the verdict tostate.log_history, and halts training on HACK when--reward-hack-haltis set.rm_ensemblerequires ≥2 reward functions. - Cross-tokenizer ULD distillation is live (#236).
task: distillwith--uld-strategy wasserstein|topk_alignnow computes a real Wasserstein-1 (sorted-CDF) or top-k-aligned distillation loss inside the distill trainer, handling student/teacher vocab-size mismatch by clamping teacher ids to the teacher vocab. - MiniLLM reverse-KL distillation is live (#237).
--minillm-enabledadds a teacher-mixed, length-normalised reverse-KL term plus an optional pretrain-anchor SFT term (--minillm-pretrain-anchor-path/--minillm-pretrain-anchor-weight) that keeps the student near coherent language. The anchor corpus reader is cwd-contained + symlink-rejecting with a per-line byte cap. - Mid-epoch RL checkpoint is live (#238).
--rl-checkpoint-save-every-steps Nwrites a real adapter + optimizer state + JSON manifest every N steps during PPO/GRPO and prunes to--rl-checkpoint-keep-last, so a long RL run survives a crash without losing the optimizer momentum. soup iterative-dpoorchestrator is live (#239). Runs the full sample → reward-score → build-pairs → DPO-train loop across rounds: each round samples completions from the previous round's adapter, the next round trains a fresh LoRA from the base on that round's harvested pairs.--plan-onlystill renders the plan without running.- Echo-trap detector is live (#240).
--echo-trap-enabledinstalls a GRPO callback that scores per-trajectory n-gram repetition, classifies OK/WARN/TRAP against--echo-trap-threshold, logs the verdict, and halts on TRAP when--echo-trap-haltis set (catches RAGEN-style degenerate repetition in multi-turn agent RL). - GRPO variant fallback now warns once (#159). When a
--grpo-variantcustomcompute_lossfalls back to the base trainer (because the installed TRL renamed the loss inputs), the trainer logs a one-shot WARNING instead of silently degrading to the default objective.
Changed
- GRPO reference-model EMA no longer materialises full state dicts (#160).
--ref-model-ema-alphanow updates the reference model in place by iteratingnamed_parameters()(ref = (1-α)·ref + α·policy), eliminating the three model-sized allocations per step the v0.53.11 path made. A total name/shape-mismatch (0 shared parameters) logs a one-shot WARNING so a misconfigured EMA can't silently no-op.
[0.71.10] - 2026-06-03
Added
- RAG family — live wiring (closes #199, #200, #201, #202). The four retrieval / steering surfaces that shipped schema-only in v0.62.0 are now real, validated on SmolLM2-135M.
- RAFT span-mask training is live (#199).
data.format: raftrows ({query, golden_doc, distractor_docs, answer}) now train answer-only: the prompt span is masked to-100and each document is labelled[doc-N]so the model learns to cite the supporting document. Documents are shuffled reproducibly (data.raft_shuffle_seed). Rows whose prompt fillsmax_length(answer fully truncated) are dropped with a warning rather than silently shrinking the effective dataset. soup ra-dit— one-shot two-stage orchestrator (#200). Trains the retriever (stage 1, embedding/contrastive) then the generator (stage 2, RAFT-SFT) in a single command, recording the trained retriever as the generator's paired retriever. Asoup trainof a generator-stage config with no retriever model set now auto-links the most-recent RA-DIT retriever run from the Registry.--plan-onlyvalidates both configs without training;--retriever-modeloverrides the auto-link.soup steer train/apply+soup serve --steerare live (#201). Fit a CAA (contrastive activation addition), ITI (inference-time intervention) or RepE (representation-engineering PCA) control vector from{positive, negative}contrastive pairs, persist it as a safetensors + config artifact, and apply it at decode time via a forward hook (soup serve --steer <name> --steer-strength <s>).soup eval citation+ citation-span loss boost are live (#202). Score citation precision / recall / F1 over{predicted, expected_ids}or RAFT rows (--shuffle-seedaligns the golden[doc-N]id with what the model saw at train time). Whencitation_faithful: true, bracketed[doc-id]spans in the answer get a boosted per-token loss weight. A newcitationfailure mode is available insoup diagnose.
[0.71.9] - 2026-06-03
Added
- Knowledge edit + unlearn — live wiring (closes #193, #194, #196, #197, #203). The v0.61.0 / v0.62.0 schema-only stubs are now live, validated on SmolLM2-135M.
soup edit set(ROME / MEMIT / AlphaEdit) is live (#194). Newsoup_cli/utils/edit_kernels.pyships covariance-free rank-1 weight-edit kernels: ROME (single-layerW += δ·kᵀ/‖k‖²), MEMIT (residual distributed across a layer band), AlphaEdit (ROME update projected orthogonal to the down-proj's top singular direction).apply_editloads the model, optimises the target residual, applies the rank-1 update, and optionally saves with cwd-containment + symlink rejection.--output,--device,--governor/ --no-governorflags added. On a tiny model a ROME edit movedP("Lyon" | "The capital of France is")from 0.0016 → 0.96.soup edit difflive before/after generation (#194). Pass--before-model+--after-model(+--probes) to generate completions through both models and surface the probes whose output changed.- EditGovernor SQLite persistence + cross-process locking (#196). New
EditGovernorStore(mirrorsnamespace_pin.NamespacePinStore— $HOME/$CWD/$TMPDIR containment, TOCTOU symlink rejection, WAL + busy_timeout,fcntl/msvcrtsidecar lock, POSIX 0600).save_governor/load_governor/default_governor_db_path(env overrideSOUP_EDIT_GOVERNOR_DB) persist per-base-model edit-count + verdict across separatesoup edit setruns. apply_editconsults the EditGovernor automatically (#197). When a governor is supplied,check_can_edit()runs BEFORE the model load (refusing on norm blowup / edit cap) andrecord_edit()runs AFTER with the measured Frobenius delta.- Live GRACE codebook (#203).
GraceCodebook(epsilon-ball nearest-key lookup),apply_grace_edit(captures a residual key + optimises a value + appends to a codebook sidecar),save_codebook/load_codebook(atomic, cwd-contained, symlink-rejected),install_grace_hook(decode-time residual substitution). Newedited_model/grace_codebookRegistry artifact kinds. soup train --task unlearnis live (NPO / SimNPO / RMU) (#193). Newsoup_cli/utils/unlearn_kernels.py(NPO(2/β)·mean(-logσ(-β(πlp-reflp))), length-normalised SimNPO, RMU representation steering) + a self-containedUnlearnTrainerWrapperloop loading a LoRA policy, a frozen reference (NPO/RMU), and forget/retain JSONL datasets. NPO/SimNPO forget loss decreased on the tiny-model smoke. Warns when run without a retain set.
Security
_save_edited_model/UnlearnTrainerWrapperoutput dirs +save_codebook/load_codebook+_load_unlearn_rowsenforce cwd-containment, raw-path symlink rejection (TOCTOU), null-byte rejection, and file-size / per-line caps.apply_grace_edithonours the governor for direct callers.
[0.71.8] - 2026-06-03
Added
- Probes & SAE — real weights + live downloads (closes #215, #216, #217,
#218, #219). A new shared
soup_cli/utils/probe_kernel.pyprovides the linear-probe math (contrast-pair derivation, apply, flag-rate, verdict bands, operator-supplied weight loading, deterministic synthetic fallback); every heavy import (numpy/torch/safetensors) is lazy. soup probe sleeper --weights <w.npz|.npy|.safetensors>(#215) — load a real calibrated probe direction instead of the synthetic fallback. Weights are cwd-contained, symlink-rejected,O_NOFOLLOW-opened,allow_pickle=False, and size-capped.compute_contrast_probe(positive, negative)derives a probe from contrast-pair activations.soup probe sae-diff <repo> --auto-download(#216) — fetch an allowlisted SAE from the HF Hub into~/.soup/sae-cache/(validated againstHF_HUB_ALLOWLISTBEFORE any network call) via a new SSRF-hardenedsoup_cli.utils.hubs.snapshot_download(repo-id shape + home/cwd/tmp cache containment + namespace-pin TOFU gate).soup probe truth/soup probe harm(#217) — TruthfulQA-style honesty and HarmBench-style misuse activation probes (6 bundled bases each, 5% / 20% verdict bands,--weightsto skip the allowlist with a real probe). The probe pack now ships truth + harm entries per base.soup probe interference --measure <eval_suite> --base-model <m> --adapter name=path ...(#218) — auto-measure the N×N interference matrix by actually loading the base + each LoRA adapter (PEFT multi-adapter), measuring loss for each adapter alone (diagonal) and each co-loaded pair (add_weighted_adapter(combination_type="cat"), off-diagonal). Exit 2 on a MAJOR worst-pair.soup train --capture-activations <layer> --capture-prompts <jsonl>(#219) — a post-training hook writes an SAE-diff-ready per-token activation snapshot to<output>/activations/activations.json.resolve_layer_moduleresolves the samemodel.layers.Npath whether or not a LoRA adapter is loaded (PEFT-wrapper fallback).
Security
- Probe / SAE / capture file I/O is cwd-contained +
O_NOFOLLOW(TOCTOU close)- size-capped; SAE weight loads use
allow_pickle=False. SAE auto-download validates the allowlist before any network call and rejects a glob result that resolves outside the snapshot dir (symlink-escape guard).
- size-capped; SAE weight loads use
Notes
- #215 is partial: the operator-supplied / contrast-pair / synthetic paths ship now, but the 6 large-base Anthropic-calibrated probe vectors remain upstream-gated (no public calibrated artifact exists). Documented as a known limitation.
[0.71.7] - 2026-06-02
Added
- Eval live runners — six probe surfaces that previously emitted heuristic
/ neutral stubs now load a real model and run live (closes #161, #162, #208,
#211, #212, #165). New shared
soup_cli/utils/live_eval.pyprovides the model-loading primitives (generator / multi-generator closures, masked cross-entropy eval-loss, a short-LoRA probe, and held-out logit agreement); every heavy import (torch/transformers/peft/lm_eval) is lazy. soup advise --probe-model <id>— runs a LIVE ROI probe: zero/few-shot token-F1 baselines, a short LoRA probe (relative held-out-loss improvement + real wall-clock), and base-model proximity (held-out logit agreement) folded into the dataset profile. Without--probe-model,--probestays the offline heuristic.soup tunability --live— replaces the offline heuristic with a real per-candidate LoRA probe (loads eachrepo_id, trains--probe-stepson a held-out-excluded slice, reports the held-out-loss drop).soup eval capability --live --model <id>— invokes lm-eval-harness per resolved task (or a--tasksoverride) with--limit/--device, isolating per-task failures and surfacing a no-metric result as an explicit error.soup eval behavior --base-model <id> [--adapter <path>]— generates pre/post responses on the bundled behaviour battery and scores the live diff.soup diagnose --base-model <id> [--adapter <path>] [--dataset <jsonl>] [--tokenizer <id>]— runs all six failure-mode probes (forgetting / refusal / format / mode_collapse / memorization / contamination) live viasoup_cli.utils.diagnose.live.run_live_diagnose; falls back to neutral OK or--evidenceJSON when no model is supplied.
Security
- The two new JSONL dataset readers (
diagnose.live._load_dataset_rows,tunability._load_jsonl_rows) open withO_NOFOLLOWafter the cwd-containment check, closing the check→open TOCTOU window (matches the v0.65 / v0.67 reader policy).
[0.71.6] - 2026-06-02
Added
soup buildlive runner — the dbt-for-SFT DAG (soup build <manifest>) now materialises datasets instead of only dry-running the plan. Five built-in transforms ship live (identity,drop_empty,lowercase,strip,dedup_exact);tablerebuilds from scratch,viewre-derives on every run, andincrementalre-transforms only the rows whose content hash changed (tracked in a SQLite state store, keyed by row hash and the model's transform+config fingerprint so a transform change re-runs everything). Custom transforms are passed per-run via the Python API'stransforms=map. Outputs are written atomically; the--output-diris symlink-checked before any directory is created.soup data gen-magpielive generator — the Magpie synthetic generator (Xu et al. 2024) now actually generates. It feeds an aligned model its chat-template prefix (chatml / llama3 / gemma / mistral families auto-detected) and harvests the self-generated user instruction + assistant response via raw completion. Live providers:ollama(/api/generateraw) andvllm(/v1/completions) — both SSRF-hardened (loopback-only HTTP);anthropicis rejected (no raw-completion endpoint). Optional--quality-filterdrops low-quality rows via the v0.47 toxicity/educational scorers; exact-duplicate instructions are de-duplicated.soup eval irt-subset --model {1pl,2pl,3pl}— the IRT eval-cost optimiser gained 2PL (per-item discrimination) and 3PL (+guessing floor) joint coordinate-ascent MLE fits alongside the existing 1PL Rasch.1plkeeps the closed-form path for back-compat;2pl/3plroute through the newfit_irt.- Tokenizer-aware memorization probe —
score_memorization(..., tokenizer=...)andsplit_prefix(..., tokenizer=...)(used bysoup diagnose) now split the prefix/suffix on real token-id boundaries and measure echo-overlap over sub-word tokens when a tokenizer (HF id / path / duck-typed object) is supplied, catching BPE-level memorization that whitespace tokenisation misses. Default (no tokenizer) keeps the whitespace behaviour.
Fixed
soup data augment --provider ollama|vllmno longer crashes — the command imported a non-existentOllamaProvidersymbol and raisedImportErroron every non-OpenAI provider. It now routes through the shared, SSRF-hardened provider factory;--model/--base-urlare honoured, the output path is containment- and symlink-checked, and the write is atomic.
Security
- Ollama / vLLM provider URLs reject
0.0.0.0—validate_ollama_url/validate_vllm_urldropped the bind-any wildcard from their loopback allow-set (nowlocalhost/127.0.0.1/::1only), matching the newervalidate_hub_endpoint/validate_webhook_urlSSRF validators. Reachable now that Magpie threads a user-supplied--base-urlthrough these providers.
[0.71.5] - 2026-06-02
Added
soup eval againstnow reads eval metrics —ExperimentTracker.get_metric_seriesfalls back to theeval_resultstable when the metric is not a per-step training column (loss/lr/grad_norm/speed/gpu_mem). Sosoup eval against <base> --candidate <run> --metric task_accuracyreturns a real score series (benchmark scores live ineval_results, notmetrics) instead of "Empty series". Per-step columns still read frommetrics— no regression for existing callers.soup adviselearns from past project outcomes —soup advisenow reads this project's accepted-verdict history (~/.soup/advise_history.jsonl) and biases the rubric: 3+ successful SFT precedents flip a marginal RAG call to SFT; 3+ negative GRPO outcomes suppress GRPO in favour of SFT-on-traces; an encouraged choice gets a small confidence nudge. Scoped per-project (one project's record never biases another). No history → identical to before.- Slack/Discord webhooks on four more commands —
--slack-url/--discord-url(SSRF-hardened, loopback-only HTTP, RFC1918 rejected, never crashes the command) now ship onsoup ingest,soup prune-prompt,soup ab(fires only on areject_h0/accept_h0decision, notcontinue), andsoup data active-sample— not justsoup drift-alarm. The validator + sender moved to a sharedsoup_cli/utils/webhooks.py. - Tokenizer-aware
soup prune-prompt—--tokenizer <model_or_path>detects and strips the shared system-prompt prefix on token boundaries instead of characters, so a multi-byte UTF-8 prefix can never be split mid-code-point. Default (no--tokenizer) keeps the whitespace-character behaviour. - Curriculum bucketing by loss percentile —
DynamicCurriculumCallbacknow buckets samples by the percentile rank of the live loss (or perplexity) signal within a rolling window whendata.curriculum_metricisloss/perplexity, so a consistently-hard sample is routed to the same difficulty bucket across recomputes.lengthand warm-up still use round-robin. --hubonsoup data pushandsoup data forge—soup data push --hub modelscope|modelersuploads a dataset via the matching SDK (repo_type=dataset, commit message sanitised);soup data forge --hub <non-hf> --teacher owner/namepre-fetches the teacher model from that hub (and warns when the teacher is not a repo id so--hubis never silently ignored). HF stays the default.
Notes
- Live SaaS pull adapters for
soup ingest(Langfuse / LangSmith / Helicone / OpenPipe / OpenAI SDKs, issue #204) remain deferred: they need credentialed vendor accounts with populated trace data to validate honestly. Tracked as an open,infra-blocked(external-account) item.soup ingestcontinues to parse the JSONL export you pull from your dashboard.
[0.71.4] - 2026-06-02
Added
- Live canary verdict for
soup adapters merge—--canary <suite.json>scores the merged adapter against the first input and classifies OK / MINOR / MAJOR using the Quant-Lobotomy taxonomy (drop <2% OK, <5% MINOR, else MAJOR).--strict-verdictexits 2 on MAJOR. Pre-scored{"baseline_scores","candidate_scores"}suites run with no model load; a{"tasks":[...]}suite uses an injectable scorer. Replaces the v0.57UNKNOWNstub. - Live evolutionary merge —
soup adapters merge --strategy cmaes --eval <suite> --budget <t>now runs the full CMA-ES loop: each candidate is merged, materialised, scored against the eval suite, and the best-weighted merge is written to--output. Replaces the v0.67 plan-only stub. - Publish an adapter PR to GitHub —
soup adapters pr <title> --base-sha <hex> --adapter <path> --push owner/repo#Nposts the rendered PR Markdown as a GitHub PR comment viagh api(argv-list, body over JSON stdin; no shell). Token resolves fromGITHUB_TOKEN/GH_TOKEN. - Pre-wired
soup loopproduction stages —soup loop init --pre-wired(orsoup loop watch --pre-wired) swaps the v0.58 no-op stage stubs for real harvest (traces → preference pairs) → DPO train → eval-gate → canary-deploy callables.soup loop statusnow shows thepre_wiredflag. - Loop iterations as Soup Cans + Registry lineage —
soup loop watch --pack-canspacks each successful iteration as a v0.26 Soup Can and appends a Registry entry (tagloop-iter), chaining a real lineage DAG across iterations visible throughsoup history. `soup loop replay --extract ` unpacks a recorded iteration. - Branch pointers into the Registry —
soup adapters branch <name> --attach-to-registry <id>links a branch snapshot to a Registry entry (shown as abranchesnode insoup history);soup adapters branch <name> --from-registry <id>derives a fresh snapshot's config + base from an entry.
Security
- The backdoor-scan gate (v0.71.2 #192) and license-conflict gate (v0.60 Part E)
now run for all merge strategies, including
--strategy cmaes(previously bypassed because cmaes returned before the gates). soup loopcanary deploy restrictsSOUP_LOOP_SERVE_ENDPOINTto loopback / RFC1918-private hosts (a serve endpoint is the operator's own box/LAN), beyond the general webhook SSRF policy which permits any HTTPS host.soup adapters pr --pushbuilds theghchild environment from an allowlist so unrelated secrets (HF_TOKEN/OPENAI_API_KEY/ …) never reach the subprocess.- The canary-suite JSON read uses
O_NOFOLLOW+os.fstat(size cap enforced on the same fd) to close the symlink/size-cap TOCTOU window.
[0.71.3] - 2026-06-01
Added
- Energy & CO2 measurement for training —
soup train --track-energywraps the training window in a codecarbon offline tracker (no IP-geolocation network call) and reports kWh / CO2 / grid intensity, feeding those numbers into--annex-xi. NewEnergyTrackercontext manager; graceful no-op when codecarbon is absent (pip install soup-cli[carbon]).--energy-countrypicks the ISO-3166 alpha-3 grid for the CO2 estimate (defaultUSA). - PDF Annex XI/XII documents —
soup train --annex-xi report.pdfnow renders a reportlab PDF (a.mdpath still renders markdown).pip install soup-cli[pdf]. - Auto-populated training-corpus domains in Annex XI/XII — the top crawled domains (with shares) are now extracted from the training JSONL and listed in the EU AI Act docs, replacing the previous empty placeholder.
- Soup Can manifest v3 with embedded attestations —
soup can pack --attest <statement.json>(repeatable) embeds in-toto Statements into a v3 can manifest; v1/v2 cans still load. Each statement is shape- and size-validated. - Local audit log auto-instrumentation — every
soupcommand now appends one HIPAA/SOC2-shaped record to~/.soup/audit.jsonl(secrets redacted, args capped). Opt out per-invocation with--no-audit-logor globally withSOUP_NO_AUDIT_LOG=1. Tail/rotate withsoup audit-log. - Reproducibility receipt in airgap bundles —
soup airgap-bundle --repro-receipt <receipt.json>embeds an SR 11-7 receipt asrepro-receipt.json; auto-detected from<model>/repro-receipt.jsonwhen not supplied.
Security
soup can pack --attestnow rejects oversize attestation files by their raw size before parsing them into memory (defence against memory-exhaustion).- The new file-loading paths (attestation JSON, airgap receipt, training-corpus
scan, PDF write) are all cwd-contained + TOCTOU symlink-rejected and
size-capped; the audit auto-log redacts
hf_/sk-/Bearertokens and never crashes the CLI on a broken log.
[0.71.2] - 2026-06-01
Added
- ed25519 signing for
soup adapters sign/soup attest— real detached signatures (over the adapter Merkle root / the in-toto statement) via a new[sign]extra (pip install soup-cli[sign], pullingcryptography).soup adapters sign --backend ed25519 --key <priv.pem>(or--generate-key <out.pem>, orSOUP_SIGNING_KEY);soup adapters verify [--public-key <trusted.pem>]does a cryptographic verify and, with a trusted key, genuine authentication.soup attest emit --sign ed25519 --key <priv.pem>writes a<output>.sigsidecar; newsoup attest verify <statement> --signature <sig>verifies it (canonical-JSON, so it's platform/newline-independent). Sigstore keyless signing stays infra-blocked (needs an OIDC provider + Fulcio/Rekor network — can't be honestly validated offline). - Anti-AI-Jacking namespace pin on Hub downloads — HF model fetches now
consult a trust-on-first-use pin store: a repo whose author changes (or whose
created_atjumps backward) is refused unless the namespace shift is explicitly allowed. Fails open when repo metadata is unavailable. - License auto-detection at
soup adapters merge— when--licenseisn't given, the license is read from each adapter'sadapter_config.json/config.json/ model-card frontmatter (HFllama3.1-style ids mapped to canonical) and the conflict gate runs automatically. - Backdoor-scan gate at
soup adapters merge— refuses to merge any input whosesoup adapters scanreturns FAIL (or can't be scanned) unless--allow-unscannedis passed; WARN is advisory.
Changed
- License-conflict overrides (
--license-override <reason>) are now recorded to the audit log for legal review. - The namespace-pin store now uses SQLite WAL + busy-timeout and a cross-process file lock around its get+insert, so concurrent writers don't lose the trust anchor.
Security
- ed25519 verification fails closed (any tamper / wrong key / missing key ⇒
invalid). Signing keys + trusted public keys are symlink-rejected and
size-capped via a shared reader (no cwd-containment — keys live outside the
project).
--generate-keyrefuses to overwrite any existing path.
[0.71.1] - 2026-06-01
Added
soup env fix— render a reproducible install plan fromsoup-env.lock. Emits copy/pasteuv pip installcommands (--format uv-pip, default) or arequirements.txtbody (--format requirements);--outputoptionally writes arequirements.txtunder cwd. Print-only by design — never shells out to a package manager.soup lock write --env-lock <path>— auto-derive--env-hashfrom asoup-env.lockso operators who ransoup env lockdon't copy the hash by hand.--env-hashstill wins when passed explicitly.soup serve --record-thumbs <db>— capture thumbs-up/down feedback into a local-RL SQLite at startup, plus a newPOST /v1/thumbsendpoint (transformers backend). Returns 404 when the flag isn't set.- Judge-calibration persistence:
JudgeCalibrationReport.to_dict,write_judge_calibration, andload_judge_calibration, backed by a newjudge_calibrationregistry artifact kind. Loading re-validates the report so a corrupt on-disk field is rejected. - Bundled MUSE and WMDP unlearning eval fixtures so
soup eval unlearning --benchmark muse|wmdpruns out of the box. WMDP forget-set probes ship redacted (placeholder prompts +REFUSEDresponses) — Soup never ships verbatim hazardous content.
Changed
soup completionsnow introspects a cached base model's actual LoRA target modules (config-onlyAutoConfigload,local_files_only=True, never networks or raises) and falls back to the canonical default shape when the base isn't cached locally.build_dagexposes avalidate_build_sourcehelper (cwd-containment + symlink rejection) for build-manifest source paths.
0.71.0 - 2026-06-01
Changed
- Breaking — install split. The heavy training stack (
torch,transformers,peft,trl,datasets,bitsandbytes,accelerate) moved out of the core install into a new[train]extra.pip install soup-cliis now a light CLI + data-tools install with no PyTorch; runpip install 'soup-cli[train]'(or[all]) to fine-tune. Existing users who train must reinstall with[train]. Version pins are unchanged. - Trimmed
README.mdto a ~238-line front door; the full feature reference now lives underdocs/(one topic page per area, indexed from the README). - Raised the pytest coverage gate from 50% to 77% (
--cov-fail-under=77). - Migrated to a
src/layout (src/soup_cli/) for cleaner packaging and to stop tests accidentally importing the in-tree package.
Added
[train]and[all]optional-dependency extras ([all]pullstrain,serve,ui,data).[dev]self-references[train]so CI and contributors still get the full stack frompip install -e ".[dev]".- Friendly error mapping: a missing heavy dependency (
torch,transformers,peft,trl,datasets,bitsandbytes,accelerate) now surfaces "Training needs the [train] extra. Run: pip install 'soup-cli[train]'". py.typedmarker (PEP 561) so downstream type checkers pick up Soup's inline type hints..pre-commit-config.yamlwith ruff (lint + format) and standard file-hygiene hooks.- Lenient
mypyconfiguration and a non-blockingtype-checkCI job. - This
CHANGELOG.md.
Removed
- The historical, per-version security-fix log that had grown inside
SECURITY.md(~220 KB).SECURITY.mdis now a concise security policy; the detailed hardening notes remain in git history and the GitHub Releases notes.