soup/docs
Alpamys 937abb9e0d feat(data): soup data doctor + soup data lint — Fine-tune Doctor (v0.71.27)
Add `soup data doctor` and `soup data lint`, killing the top *silent*
fine-tune failures before a single training step: EOS-missing-from-labels
(the #1 "model never stops generating" bug), BOS duplication, no-system-role
templates, and preference-data length bias (the #1 silent DPO degradation) —
none of which any competitor (Unsloth/Axolotl/LlamaFactory) checks for.

- utils/data_doctor.py: 8-check chat-template compat report over a
  tokenizer + sampled rows, OK/MINOR/MAJOR taxonomy mirroring diagnose;
  --show-mask N renders per-token trained/masked colouring through the
  SAME masking dispatch (_build_row_labels) the report itself uses, so
  the two can never disagree about what's actually trained.
- utils/data_lint.py: preference-data linter (dpo/orpo/simpo/ipo/bco/kto)
  — length bias (Cohen's d), label imbalance, near-duplicates (MinHash),
  identical chosen==rejected pairs, prompt leakage.
- commands/data_doctor.py: Typer layer for both commands; strips C0
  control bytes before untrusted dataset content reaches the terminal.
- commands/diagnose.py: hardens the --evidence loader against a TOCTOU
  symlink swap (O_NOFOLLOW + fstat-on-open-fd), backporting the pattern
  soup ship shipped in v0.71.25 (closes v0.71.25 known-limitation (4)).

Live smoke against the real HuggingFaceTB/SmolLM2-135M-Instruct tokenizer
(Windows + RTX 3050) found and fixed two genuine bugs beyond the synthetic
fixtures: the EOS check needed to span-search the whole trained region
(not just the last token), and two apply_chat_template call sites needed
a broad except Exception for jinja2.exceptions.TemplateError.

+173 tests (14788 -> 15042). 5 sequential ECC reviews, every finding fixed.
2026-07-04 14:03:30 +05:00
..
README.md docs: index soup ship in docs/README + CONTRIBUTING utils list 2026-06-28 00:02:23 +05:00
adapters-and-governance.md feat(edit): GPT-2 Conv1D edits, covariance ROME, atomic governor, Mixtral LongLoRA (v0.71.16) 2026-06-07 16:11:56 +05:00
backends-and-ops.md feat(precision,rollout): live fp8/nvfp4 + vLLM sleep + openenv rollout + apple-adapter + delinearize-llama4 (v0.71.21) 2026-06-10 16:29:52 +05:00
commands.md feat(data): soup data doctor + soup data lint — Fine-tune Doctor (v0.71.27) 2026-07-04 14:03:30 +05:00
data.md feat(data): soup data doctor + soup data lint — Fine-tune Doctor (v0.71.27) 2026-07-04 14:03:30 +05:00
evaluation.md feat(eval): soup ship — SHIP / DON'T-SHIP verdict (v0.71.25) 2026-06-27 23:37:35 +05:00
models.md feat(recipes): 2026 model-family expansion — 17 SFT recipes, catalog 116→133 (v0.71.24) 2026-06-21 13:00:47 +05:00
peft-and-efficiency.md docs: refresh quant-menu modality + multipack sharding notes (v0.71.19) 2026-06-09 13:02:30 +05:00
performance-and-quantization.md feat(precision,rollout): live fp8/nvfp4 + vLLM sleep + openenv rollout + apple-adapter + delinearize-llama4 (v0.71.21) 2026-06-10 16:29:52 +05:00
serving-and-export.md feat(recipes): add ready-made SFT recipe for Qwen2.5-Coder-7B-Instruct (#285) 2026-06-28 19:22:29 +05:00
training.md docs(train): v0.71.26 release — closed-loop reward-hacking mitigation 2026-07-01 16:55:59 +05:00

README.md

Soup Documentation

← Back to the main README

The main README is the 5-minute front door. This directory holds the full feature reference — every soup capability, grouped by area.

Guide Covers
Training tasks & methods SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, tool-calling, PRM, pre-training, distillation, classification, vision/audio/TTS, unlearning, RAFT/RA-DIT, loop-hardening detectors
PEFT, long context & efficiency DoRA, LoRA+, rsLoRA, VeRA, OLoRA, NEFTune, PiSSA, ReLoRA, optimizer & PEFT zoo, LLaMA Pro, GaLore, YaRN/LongLoRA, packing, curriculum, auto-tuning
Performance & quantization QAT, FP8, Quant Menu (I + II), KV-cache, NVFP4, save formats, Cut Cross-Entropy, gradient checkpointing, kernels, activation offloading, multi-GPU / DeepSpeed / FSDP
Data engineering Formats, the Axolotl/LF-parity pipeline, data tools, synthetic generation & forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs
Evaluation & probes Eval design/gate, eval-gated training, benchmarks, NLG metrics, calibration, Elo arena, diagnose, soup ship verdict, post-train X-ray probes, A/B, drift, tunability, soup advise
Serving & export OpenAI-compatible server, batch inference, benchmarking, merge/export, Anthropic Messages endpoint, speculative decoding, deploy autopilot, Web UI, Agent Forge
Adapters, registry & governance Adapter lifecycle/management, model registry, Soup Cans, the data flywheel (soup loop), knowledge editing, steering, supply-chain controls
Backends, platform & ops MLX/Unsloth backends, Modal cloud GPU training, alternative hubs, HF Hub integration, autopilot, experiment tracking, plan/apply, env lockfiles, hardware-fit, completions, plugins, utility commands
Command reference The full soup command list
Supported models & extras Recommended model families, the VRAM size guide, the pip extras matrix

Per-release notes live on the GitHub Releases page; see also the repo-root CHANGELOG.md.