mirror of https://github.com/razor-ai/soup.git
Closes #141, #124, #125, #228, #97. - #141: apply_fp8_attention (torchao float8 on attention projections, Hopper gate) + apply_nvfp4 (NVFP4Config, Blackwell gate); partial-conversion honesty; wired into the v0.28 speed/memory pipeline with yellow-advisory degrade. - #124: vllm_sleep_mode live - create_vllm_engine(sleep_mode=True) + vllm_sleep_cycle ctx (wake in finally) + TRL GRPOConfig hook probe. - #125: openenv rollout fully live via training.rollout_func module:fn resolver; rows replace the prompt dataset; art/ruler/nemo_gym honest dep gates + _EXTERNAL_ROLLOUT_RUNNERS seam. Real GRPO train on SmolLM2-135M. - #228: convert_apple_adapter live - PEFT LoRA <-> mlx-lm (both matrices transpose, bf16 upcast, adapters.safetensors + num_layers, npz legacy read, np.ascontiguousarray fix for safetensors non-contiguous mangling); *-to-apple upstream-gated exit 3. - #97: delinearize-llama4 live - [E*din,dout] -> [E,din,dout] per shard, config.json expert-count probe + --num-experts, sidecar copy, atomic writes. Review waves: 3 HIGH + ~8 MEDIUM + ~12 LOW fixed. Tests: 13874 -> 14084 (+210 in tests/test_v07121.py). Full suite: 13967 passed, 117 skipped. ruff clean. |
||
|---|---|---|
| .. | ||
| README.md | ||
| adapters-and-governance.md | ||
| backends-and-ops.md | ||
| commands.md | ||
| data.md | ||
| evaluation.md | ||
| models.md | ||
| peft-and-efficiency.md | ||
| performance-and-quantization.md | ||
| serving-and-export.md | ||
| training.md | ||
README.md
Soup Documentation
The main README is the 5-minute front door. This directory holds the full
feature reference — every soup capability, grouped by area.
| Guide | Covers |
|---|---|
| Training tasks & methods | SFT, DPO/GRPO/PPO/KTO/ORPO/SimPO/IPO/BCO, tool-calling, PRM, pre-training, distillation, classification, vision/audio/TTS, unlearning, RAFT/RA-DIT, loop-hardening detectors |
| PEFT, long context & efficiency | DoRA, LoRA+, rsLoRA, VeRA, OLoRA, NEFTune, PiSSA, ReLoRA, optimizer & PEFT zoo, LLaMA Pro, GaLore, YaRN/LongLoRA, packing, curriculum, auto-tuning |
| Performance & quantization | QAT, FP8, Quant Menu (I + II), KV-cache, NVFP4, save formats, Cut Cross-Entropy, gradient checkpointing, kernels, activation offloading, multi-GPU / DeepSpeed / FSDP |
| Data engineering | Formats, the Axolotl/LF-parity pipeline, data tools, synthetic generation & forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs |
| Evaluation & probes | Eval design/gate, eval-gated training, benchmarks, NLG metrics, calibration, Elo arena, diagnose, post-train X-ray probes, A/B, drift, tunability, soup advise |
| Serving & export | OpenAI-compatible server, batch inference, benchmarking, merge/export, Anthropic Messages endpoint, speculative decoding, deploy autopilot, Web UI, Agent Forge |
| Adapters, registry & governance | Adapter lifecycle/management, model registry, Soup Cans, the data flywheel (soup loop), knowledge editing, steering, supply-chain controls |
| Backends, platform & ops | MLX/Unsloth backends, Modal cloud GPU training, alternative hubs, HF Hub integration, autopilot, experiment tracking, plan/apply, env lockfiles, hardware-fit, completions, plugins, utility commands |
| Command reference | The full soup command list |
| Supported models & extras | Recommended model families, the VRAM size guide, the pip extras matrix |
Per-release notes live on the GitHub Releases page; see also the repo-root CHANGELOG.md.