soup/docs
Alpamys bcbf72e586 feat(train): layer streaming breadth — 6 more archs, bigger batches, resume, disk tier (v0.72.3)
Lifts the v0.72.0-.2 scope freeze. Every capability was gated against a
streamed-vs-resident bit-exactness reference before it was written.

- Six more families (mistral/gemma/gemma2/gemma3_text/phi/phi3), each
  bit-exact vs the same checkpoint loaded resident, under bf16 AND NF4.
  Multimodal gemma3 stays refused — only gemma3_text.
- batch_size > 1 and gradient_accumulation_steps > 1 now work.
- A batch- and vocab-aware VRAM pre-flight that refuses a run predicted
  not to fit. Fitted to 10 real runs: worst error 0.85%, never
  under-predicts. On Windows an over-budget step does not OOM; WDDM
  spills silently, so the estimator is the only guard.
- A throughput bracket from a GEMM ceiling measured on the user's own
  card in the same session, printed with the SM clock.
- --resume / --hf-resume: load_state_dict narrows keys by child name, so
  a canonical checkpoint matched 0 of N tensors and PEFT warned only.
  Keys are now redirected at load time, mirroring the v0.72.1 save fix.
- An NVMe disk overflow tier (stream_source: auto|ram|disk), bit-exact
  against the RAM tier. Its speed relative to RAM is UNMEASURED here and
  no figure is claimed.
- soup doctor --disk reports the detected media type.

Fixes: estimate_logits_bytes charged 6 bytes/element where the measured
peak is 14; the NVMe tier guard was wired to a hardcoded constant;
streaming sources leaked handles when training raised; subprocess
helpers resolved tools by bare name (CWE-427 on Windows).

112 tests in tests/test_v07203.py; 16867 -> 16977.
2026-07-28 23:15:20 +05:00
..
README.md feat(train): layer streaming — fine-tune models larger than VRAM (v0.72.0 BETA) 2026-07-26 23:58:06 +05:00
adapters-and-governance.md docs: v0.71.34 adapter algebra + LISA (version bump + CHANGELOG + docs) 2026-07-15 13:39:46 +05:00
backends-and-ops.md feat(reward): soup reward stress — adversarial verifier gameability probe (v0.71.41) 2026-07-19 20:53:37 +05:00
commands.md feat(train): layer streaming breadth — 6 more archs, bigger batches, resume, disk tier (v0.72.3) 2026-07-28 23:15:20 +05:00
compliance.md feat(compliance): init templates + soup card + soup ci init + GGUF-on-Windows (v0.71.35) 2026-07-15 20:09:10 +05:00
data.md fix(cli): quote install hints so `pip install soup-cli[extra]` works on cmd.exe (v0.71.37) 2026-07-17 20:40:54 +05:00
evaluation.md feat(ship): close the evidence loop — emit-evidence + config + provenance + PR comment (v0.71.39) 2026-07-19 12:30:35 +05:00
models.md fix(cli): quote install hints so `pip install soup-cli[extra]` works on cmd.exe (v0.71.37) 2026-07-17 20:40:54 +05:00
peft-and-efficiency.md docs: close the v0.71.35 checklist gaps (index row, toolchain, stale counts) 2026-07-15 21:04:01 +05:00
performance-and-quantization.md feat(train): layer streaming breadth — 6 more archs, bigger batches, resume, disk tier (v0.72.3) 2026-07-28 23:15:20 +05:00
serving-and-export.md fix(cli): quote install hints so `pip install soup-cli[extra]` works on cmd.exe (v0.71.37) 2026-07-17 20:40:54 +05:00
training.md feat(train): layer streaming breadth — 6 more archs, bigger batches, resume, disk tier (v0.72.3) 2026-07-28 23:15:20 +05:00

README.md

Soup Documentation

← Back to the main README

The main README is the 5-minute front door. This directory holds the full feature reference — every soup capability, grouped by area.

Guide Covers
Training tasks & methods SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, tool-calling, PRM, pre-training, distillation, classification, vision/audio/TTS, unlearning, RAFT/RA-DIT, loop-hardening detectors, reward-verifier synthesis
PEFT, long context & efficiency DoRA, LoRA+, rsLoRA, VeRA, OLoRA, NEFTune, PiSSA, ReLoRA, optimizer & PEFT zoo, LLaMA Pro, GaLore, YaRN/LongLoRA, packing, curriculum, auto-tuning, depth pruning + distill-heal (soup shrink)
Performance & quantization QAT, FP8, Quant Menu (I + II), KV-cache, NVFP4, save formats, Cut Cross-Entropy, gradient checkpointing, kernels, activation offloading, layer streaming, multi-GPU / DeepSpeed / FSDP
Data engineering Formats, the Axolotl/LF-parity pipeline, data tools, synthetic generation & forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs
Evaluation & probes Eval design/gate, eval-gated training, benchmarks, NLG metrics, calibration, Elo arena, diagnose, soup ship verdict, post-train X-ray probes, A/B, drift, tunability, soup advise
Serving & export OpenAI-compatible server, batch inference, benchmarking, merge/export, Anthropic Messages endpoint, speculative decoding (train + measure your own draft), deploy autopilot, Web UI, Agent Forge
Adapters, registry & governance Adapter lifecycle/management, model registry, Soup Cans, the data flywheel (soup loop), knowledge editing, steering, supply-chain controls
Compliance & governance quickstart HIPAA/SOC2/EU-AI-Act/SR-11-7 init templates, provenance (BOM/attest/repro-receipt), audit log, air-gap, model-card autogen (soup card), CI gate (soup ci init)
Backends, platform & ops MLX/Unsloth backends, Modal cloud GPU training, alternative hubs, HF Hub integration, autopilot, experiment tracking, plan/apply, env lockfiles, hardware-fit, completions, plugins, utility commands
Command reference The full soup command list
Supported models & extras Recommended model families, the VRAM size guide, the pip extras matrix

Per-release notes live on the GitHub Releases page; see also the repo-root CHANGELOG.md.