Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
Go to file
Alpamys 4dae34a6a4 docs(v0.72.0): list layer streaming in the README docs index and commands reference
Two spots the release missed: the README's own docs-index row for
Performance & quantization (docs/README.md's equivalent row was already
updated), and docs/commands.md, which lists config-driven training
features in the same style as LISA and Spectrum.

Docs-only — no version bump.
2026-07-27 01:06:53 +05:00
.github ci: warm HF Hub cache before tests to fix tiny-model 429 flake 2026-06-04 19:53:33 +05:00
docs docs(v0.72.0): list layer streaming in the README docs index and commands reference 2026-07-27 01:06:53 +05:00
examples fix(cli): quote install hints so `pip install soup-cli[extra]` works on cmd.exe (v0.71.37) 2026-07-17 20:40:54 +05:00
src/soup_cli feat(train): layer streaming — fine-tune models larger than VRAM (v0.72.0 BETA) 2026-07-26 23:58:06 +05:00
templates Initial project setup: CLI skeleton + config + trainer + data pipeline 2026-02-20 16:14:56 +05:00
tests fix(tests): two v0.72.0 tests were platform-dependent, not portable 2026-07-27 00:42:42 +05:00
.dockerignore Enhance #14 - Add official Docker support for easier onboarding (#20) 2026-04-10 23:04:00 +05:00
.gitignore chore: stop tracking internal planning docs 2026-07-05 20:52:50 +05:00
.mailmap docs: add CONTRIBUTORS.md + .mailmap; Recognition section; fix test-count drift 2026-06-02 15:25:19 +05:00
.pre-commit-config.yaml chore: project hygiene — py.typed, pre-commit, mypy CI, CHANGELOG, slim SECURITY.md 2026-05-31 19:44:47 +05:00
AGENTS.md docs: drop all public references to the gitignored CLAUDE.md / plan.md 2026-06-01 12:05:36 +05:00
CHANGELOG.md feat(train): layer streaming — fine-tune models larger than VRAM (v0.72.0 BETA) 2026-07-26 23:58:06 +05:00
CODEOWNERS chore: migrate to src-layout 2026-05-31 12:40:06 +05:00
CODE_OF_CONDUCT.md Fix Phase 6.1 community files: real emails, DPO data format, correct file names 2026-03-24 11:19:56 +05:00
CONTRIBUTING.md feat(train): layer streaming — fine-tune models larger than VRAM (v0.72.0 BETA) 2026-07-26 23:58:06 +05:00
CONTRIBUTORS.md docs: credit @Sanjays2402 for the benchmark gate-task fix (#315) 2026-07-17 16:31:43 +05:00
Dockerfile chore: release v0.71.0 — split heavy deps into [train] extra 2026-06-01 12:27:08 +05:00
LICENSE chore(license): migrate from MIT to Apache-2.0 2026-04-21 22:32:42 +05:00
NOTICE chore(license): migrate from MIT to Apache-2.0 2026-04-21 22:32:42 +05:00
README.md docs(v0.72.0): list layer streaming in the README docs index and commands reference 2026-07-27 01:06:53 +05:00
SECURITY.md feat(train): layer streaming — fine-tune models larger than VRAM (v0.72.0 BETA) 2026-07-26 23:58:06 +05:00
docker-compose.yml fix(docker): correct image name and modernize compose file 2026-04-10 23:06:04 +05:00
pyproject.toml feat(train): layer streaming — fine-tune models larger than VRAM (v0.72.0 BETA) 2026-07-26 23:58:06 +05:00
soup.png fix: replace SVG logo with PNG (GitHub doesn't render SVG in README) 2026-04-01 18:22:06 +05:00
soup_logo_svg.svg fix: use relative logo path in Web UI, add SVG logo to repo 2026-04-01 18:23:57 +05:00

README.md

Soup

Soup

Fine-tune and post-train LLMs in one command. No SSH, no config hell.

Website · Quick Start · Config · Docs · Commands · Models

PyPI Downloads Python 3.10+ Apache-2.0 License Tests CI Website


Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.

pip install "soup-cli[train]"   # add [train] to fine-tune; bare `soup-cli` is the light CLI
soup init --template chat
soup train

Why Soup?

Training LLMs is still painful. Even experienced teams spend 30-50% of their time fighting infrastructure instead of improving models. Soup fixes that.

  • Zero SSH. Never SSH into a broken GPU box again.
  • One config. A simple YAML file is all you need.
  • Auto everything. Batch size, GPU detection, quantization — handled.
  • Works locally. Train on your own GPU with QLoRA. No cloud required.

What's New

v0.72.0 — Layer streaming (BETA). Fine-tune models that don't fit in your card. The frozen base streams from CPU RAM one decoder layer at a time into a small pool of VRAM buffers while the LoRA adapters stay resident, so peak VRAM is bounded by one layer rather than the whole model. Qwen2.5-3B trains in 2.15 GB on a 4 GB card, where a resident run OOMs.

  • training.stream_layers: true — a config key, not a CLI flag. Tune with stream_source (ram in v0.72.0; disk is v0.72.2) and stream_buffers (28, default 2).
  • Measured on a 4 GB RTX 3050 Laptop (batch 1, gradient checkpointing on): 0.5B at 978.6 tok/s / 1.47 GB, 1.5B at 525.0 tok/s / 1.82 GB, 3B at 143.1 tok/s / 2.15 GB.
  • Honest cost: 1.43× slower than resident, measured at 0.5B — the only apples-to-apples comparison available on that box, because 1.5B and above cannot run resident there at all.
  • Proof-of-mechanism at 3B. Nothing above 3B was measured; no 8B/14B claim is supported. Scope: RAM tier, bf16, task: sft, Llama/Qwen, batch size 1, no gradient accumulation, no --resume. 4-bit (NF4) streaming is v0.72.1 and is refused with a clear message today.
# soup.yaml — then just `soup train --config soup.yaml`
training:
  stream_layers: true      # base streams from RAM; only the adapter trains
  batch_size: 1
  quantization: none       # NF4 streaming lands in v0.72.1
Previous release — v0.71.40, soup reward synth (generate a reward verifier from your data)

Point soup reward synth at a JSONL of reference outputs and it infers a deterministic verifier, writes a readable / committable .py reward function, and — the part nobody else does — refuses to emit one that can't tell your references from bad answers (four families: numeric / json_schema / regex / tool_call; a mandatory calibration report is the moat). Reward ensembles (reward_fn: "accuracy,format") also train now. (#311)

soup reward synth references.jsonl -o reward.py --output-report calib.json
Previous release — v0.71.39, CI for weights not prompts (emit + provenance-bind the ship verdict)

soup ship's verdict became emittable, committable, and provenance-bound: --emit-evidence makes a run replay into an identical verdict, eval.ship in soup.yaml + --config makes the gate policy reviewable, and --config binds evidence to the exact recipe that produced it (stale evidence → exit 3). soup ship --push owner/repo#N posts the SHIP / DON'T-SHIP card on the PR.

Previous release — v0.71.38, The gate grows teeth (real leg-2 regression gate)

soup ship's regression leg became real: a fixed, extraction-based scorer over seven bundled, offline suites (MCQ · arithmetic · tool-calling · JSON validity · safety/refusal). A tune that wins your task but quietly breaks tool-calling now gets a DON'T SHIP. Zero new deps.

soup ship --base ./base --adapter ./my-lora --task-eval my_task.jsonl
#   exit 0 = SHIP · 2 = DON'T SHIP · 3 = bad flags · 1 = runtime error
Previous release — v0.71.33, soup draft (measure speculative decoding)

soup draft measure reports a draft model's acceptance rate + real plain-vs-assisted tok/s (exit 0/2/1 for CI); soup draft distill distils your target into a dense tiny draft, auto-wired into soup serve --auto-spec. The honest result on a small same-family pair: distillation didn't move acceptance (69.3% → 69.3%) and assisted decoding was a net slowdown — which is exactly the number you want before shipping speculative decoding.

soup draft measure --target ./my-tuned-model --draft HuggingFaceTB/SmolLM2-135M-Instruct \
  --prompts prod-prompts.jsonl        # -> acceptance %, real tok/s, ship-or-not

Full history: CHANGELOG.md · GitHub Releases.

Quick Start

1. Install

# Light core: CLI + config + data tools, no PyTorch
pip install soup-cli

# Add the training stack (torch, transformers, peft, trl, datasets, …)
pip install "soup-cli[train]"

# Everything (train + serve + ui + data) in one shot
pip install "soup-cli[all]"

# Or from GitHub (latest dev)
pip install git+https://github.com/MakazhanAlpamys/Soup.git

The full extras table (fast, mlx, serve, eval, ui, vision, audio, …) lives in docs/models.md.

Use double quotes around the extra. They are the only spelling that works in every shell — cmd.exe, PowerShell, bash, and zsh.

Older tutorials and videos (including some of ours) show the single-quoted pip install 'soup-cli[train]'. That is bash / zsh / PowerShell syntax, and it fails on Windows cmd.exe, which has no single-quote quoting and hands the quotes straight to pip:

ERROR: Invalid requirement: "'soup-cli[train]'": Expected package name at the start of dependency specifier

If you hit that, swap the ' for " — pip is rejecting a literal quote character, nothing is wrong with the package. (Dropping the quotes entirely works on Windows too, but zsh then reads [train] as a glob and fails.)

soup init, soup data …, and the other data/inspection commands work on the light install. Fine-tuning (soup train) needs the [train] extra.

2. Create a config

soup init                       # interactive wizard
soup init --template chat       # or start from a template

Templates: chat, code, tool-calling, medical, reasoning, vision, kto, orpo, simpo, ipo, bco, rlhf, pretrain, moe, longcontext, embedding, audio.

3. Train, test, ship

soup train --config soup.yaml                 # LoRA, quantization, batching — all handled
soup chat  --model ./output                    # talk to your model
soup push  --model ./output --repo you/my-model

soup merge  --adapter ./output                              # merge LoRA into the base
soup export --model ./output --format gguf --quant q4_k_m   # GGUF for Ollama / llama.cpp

More export targets (ONNX, TensorRT, AWQ, GPTQ, BitNet) and deployment options live in docs/serving-and-export.md.

Configuration

A complete soup.yaml:

base: meta-llama/Llama-3.1-8B-Instruct
task: sft
# backend: unsloth  # 2-5x faster, pip install "soup-cli[fast]"

data:
  train: ./data/train.jsonl
  format: alpaca
  val_split: 0.1

training:
  epochs: 3
  lr: 2e-5
  batch_size: auto
  lora:
    r: 64
    alpha: 16
  quantization: 4bit

output: ./output

config/schema.py is the single source of truth for every field. Advanced data, training, and PEFT options are documented under Documentation.

Documentation

The full feature reference lives in docs/. Start here:

Guide Covers
Training tasks & methods SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, tool-calling, PRM, pre-training, distillation, classification, vision/audio/TTS, unlearning, RAFT/RA-DIT, loop-hardening detectors
PEFT, long context & efficiency DoRA, LoRA+, rsLoRA, VeRA, OLoRA, NEFTune, PiSSA, ReLoRA, optimizer & PEFT zoo, LLaMA Pro, GaLore, YaRN/LongLoRA, packing, curriculum, auto-tuning
Performance & quantization QAT, FP8, Quant Menu (I + II), KV-cache, NVFP4, save formats, Cut Cross-Entropy, gradient checkpointing, kernels, activation offloading, layer streaming, multi-GPU / DeepSpeed / FSDP
Data engineering Formats, the Axolotl/LF-parity pipeline, data tools, synthetic generation & forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs
Evaluation & probes Eval design/gate, eval-gated training, benchmarks, NLG metrics, calibration, Elo arena, diagnose, post-train X-ray probes, A/B, drift, tunability, soup advise
Serving & export OpenAI-compatible server, batch inference, benchmarking, merge/export, Anthropic Messages endpoint, speculative decoding (train + measure your own draft), deploy autopilot, Web UI, Agent Forge
Adapters, registry & governance Adapter lifecycle/management, model registry, Soup Cans, the data flywheel (soup loop), knowledge editing, steering, supply-chain controls (scan/sign/BOM/attest/audit/airgap)
Compliance & governance quickstart HIPAA/SOC2/EU-AI-Act/SR-11-7 init templates, provenance (BOM/attest/repro-receipt), audit log, air-gap, model-card autogen (soup card), CI gate (soup ci init)
Backends, platform & ops MLX/Unsloth backends, alternative hubs, HF Hub integration, autopilot, experiment tracking, plan/apply, env lockfiles, hardware-fit, completions, plugins, utility commands
Command reference The full soup command list
Supported models & extras Recommended model families, the VRAM size guide, the pip extras matrix

Data Formats

All formats are auto-detected from JSONL, JSON, CSV, Parquet, or TXT:

  • alpaca{"instruction": ..., "input": ..., "output": ...}
  • sharegpt{"conversations": [{"from": "human", "value": ...}, ...]}
  • chatml{"messages": [{"role": "user", "content": ...}, ...]}
  • dpo / orpo / simpo / ipo{"prompt": ..., "chosen": ..., "rejected": ...}
  • kto{"prompt": ..., "completion": ..., "label": true}
  • llava / sharegpt4v (vision), audio, plaintext (pre-training), embedding, prm, pre_tokenized, video, multimodal

Full schemas and the Axolotl/LlamaFactory-parity data pipeline (remote URIs, streaming, sharding, interleaving, vocab expansion, document ingestion) are in docs/data.md.

Common Commands

soup train  --config soup.yaml        # train (SFT/DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/...)
soup infer  --model ./output --input prompts.jsonl   # batch inference
soup chat   --model ./output          # interactive chat
soup serve  --model ./output          # OpenAI-compatible API server
soup merge  --adapter ./output        # merge LoRA into the base model
soup export --model ./output --format gguf           # export for deployment
soup eval   benchmark --model ./output               # evaluate
soup data   inspect ./data/train.jsonl               # dataset stats
soup recipes list                     # 100+ ready-made model recipes
soup autopilot --model <id> --data d.jsonl --goal chat  # zero-config
soup doctor                           # check GPU / deps / environment

The complete command list is in docs/commands.md.

Supported Models

Soup works with any text-generation model on the HuggingFace Hub — if it loads with AutoModelForCausalLM, it works, zero config changes. Llama 3.x/4, Qwen 2.5/3, Gemma 3, Mistral, Mixtral, DeepSeek R1/V3, Phi-4, and 100+ others ship as ready-made recipes (soup recipes list).

VRAM Max model (QLoRA 4-bit) Example
8 GB ~7B Llama-3.1-8B, Mistral-7B
16 GB ~14B Phi-4-14B, Qwen2.5-14B
24 GB ~34B CodeLlama-34B, Yi-1.5-34B
48 GB ~70B Llama-3.3-70B
80 GB+ 70B+ (full) or MoE Mixtral-8x22B, DeepSeek-V3

Full model + vision tables and the optional-extras matrix are in docs/models.md.

Docker

Run Soup without installing CUDA or PyTorch locally (image published to GHCR on every release):

docker pull ghcr.io/makazhanalpamys/soup:latest
docker run --gpus all -v $(pwd):/workspace ghcr.io/makazhanalpamys/soup train --config soup.yaml
docker compose up   # or build locally

Requirements

  • Python 3.10+
  • GPU with CUDA (recommended), Apple Silicon (MPS), or CPU (experimental — very slow)
  • 8 GB+ VRAM for 7B models with QLoRA

All training tasks run on CPU for testing (quantization auto-disabled). Optional extras (train, all, fast, vision, qat, serve, serve-fast, ui, eval, deepspeed, liger, mlx, onnx, tensorrt, …) are listed in docs/models.md.

Troubleshooting

soup doctor    # GPU, system resources, dependencies, and version in one place
  • ImportError: DLL load failed while importing _C (Windows) — reinstall PyTorch for your CUDA version: pip install torch --index-url https://download.pytorch.org/whl/cu121.
  • soup versionpip show soup-cli — multiple Python installs; use a virtualenv.

Development

git clone https://github.com/MakazhanAlpamys/Soup.git
cd Soup
pip install -e ".[dev]"

ruff check src/soup_cli/ tests/    # lint
pytest tests/ -v                   # unit tests (fast, no GPU)
pytest tests/ -m smoke -v          # smoke tests (downloads a tiny model, trains)

pre-commit install                 # optional: ruff lint+format on commit

See CONTRIBUTING.md for the full workflow and SECURITY.md to report a vulnerability.

Contributors

Built by the community ❤️ — thank you to everyone who has contributed. See CONTRIBUTORS.md.

Contributors

License

Apache-2.0. Copyright © the Soup contributors.