mirror of https://github.com/razor-ai/soup.git
Add `soup data doctor` and `soup data lint`, killing the top *silent* fine-tune failures before a single training step: EOS-missing-from-labels (the #1 "model never stops generating" bug), BOS duplication, no-system-role templates, and preference-data length bias (the #1 silent DPO degradation) — none of which any competitor (Unsloth/Axolotl/LlamaFactory) checks for. - utils/data_doctor.py: 8-check chat-template compat report over a tokenizer + sampled rows, OK/MINOR/MAJOR taxonomy mirroring diagnose; --show-mask N renders per-token trained/masked colouring through the SAME masking dispatch (_build_row_labels) the report itself uses, so the two can never disagree about what's actually trained. - utils/data_lint.py: preference-data linter (dpo/orpo/simpo/ipo/bco/kto) — length bias (Cohen's d), label imbalance, near-duplicates (MinHash), identical chosen==rejected pairs, prompt leakage. - commands/data_doctor.py: Typer layer for both commands; strips C0 control bytes before untrusted dataset content reaches the terminal. - commands/diagnose.py: hardens the --evidence loader against a TOCTOU symlink swap (O_NOFOLLOW + fstat-on-open-fd), backporting the pattern soup ship shipped in v0.71.25 (closes v0.71.25 known-limitation (4)). Live smoke against the real HuggingFaceTB/SmolLM2-135M-Instruct tokenizer (Windows + RTX 3050) found and fixed two genuine bugs beyond the synthetic fixtures: the EOS check needed to span-search the whole trained region (not just the last token), and two apply_chat_template call sites needed a broad except Exception for jinja2.exceptions.TemplateError. +173 tests (14788 -> 15042). 5 sequential ECC reviews, every finding fixed. |
||
|---|---|---|
| .. | ||
| README.md | ||
| adapters-and-governance.md | ||
| backends-and-ops.md | ||
| commands.md | ||
| data.md | ||
| evaluation.md | ||
| models.md | ||
| peft-and-efficiency.md | ||
| performance-and-quantization.md | ||
| serving-and-export.md | ||
| training.md | ||
README.md
Soup Documentation
The main README is the 5-minute front door. This directory holds the full
feature reference — every soup capability, grouped by area.
| Guide | Covers |
|---|---|
| Training tasks & methods | SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, tool-calling, PRM, pre-training, distillation, classification, vision/audio/TTS, unlearning, RAFT/RA-DIT, loop-hardening detectors |
| PEFT, long context & efficiency | DoRA, LoRA+, rsLoRA, VeRA, OLoRA, NEFTune, PiSSA, ReLoRA, optimizer & PEFT zoo, LLaMA Pro, GaLore, YaRN/LongLoRA, packing, curriculum, auto-tuning |
| Performance & quantization | QAT, FP8, Quant Menu (I + II), KV-cache, NVFP4, save formats, Cut Cross-Entropy, gradient checkpointing, kernels, activation offloading, multi-GPU / DeepSpeed / FSDP |
| Data engineering | Formats, the Axolotl/LF-parity pipeline, data tools, synthetic generation & forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs |
| Evaluation & probes | Eval design/gate, eval-gated training, benchmarks, NLG metrics, calibration, Elo arena, diagnose, soup ship verdict, post-train X-ray probes, A/B, drift, tunability, soup advise |
| Serving & export | OpenAI-compatible server, batch inference, benchmarking, merge/export, Anthropic Messages endpoint, speculative decoding, deploy autopilot, Web UI, Agent Forge |
| Adapters, registry & governance | Adapter lifecycle/management, model registry, Soup Cans, the data flywheel (soup loop), knowledge editing, steering, supply-chain controls |
| Backends, platform & ops | MLX/Unsloth backends, Modal cloud GPU training, alternative hubs, HF Hub integration, autopilot, experiment tracking, plan/apply, env lockfiles, hardware-fit, completions, plugins, utility commands |
| Command reference | The full soup command list |
| Supported models & extras | Recommended model families, the VRAM size guide, the pip extras matrix |
Per-release notes live on the GitHub Releases page; see also the repo-root CHANGELOG.md.