diff --git a/CHANGELOG.md b/CHANGELOG.md index 8eb1258..fbf6c29 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,10 +12,31 @@ reproducing 70+ versions of notes. ## [Unreleased] +## [0.71.24] - 2026-06-21 + ### Added +- **2026 model-family recipe expansion (catalog 116 → 133).** 17 new ready-made + SFT recipes for the open-weight models released Feb–Jun 2026, each with its + Hugging Face repo-ID verified to resolve: + - **Qwen 3.5 (Apache-2.0):** `qwen3.5-0.8b-sft`, `qwen3.5-2b-sft`, + `qwen3.5-4b-sft`, `qwen3.5-9b-sft`, `qwen3.5-27b-sft`, and the + `qwen3.5-35b-a3b-sft` / `qwen3.5-122b-a10b-sft` / `qwen3.5-397b-a17b-sft` + MoE sizes. + - **Qwen 3.6 (Apache-2.0):** `qwen3.6-27b-sft`, `qwen3.6-35b-a3b-sft`. + - **DeepSeek-V4 (MIT):** `deepseek-v4-flash-sft`, `deepseek-v4-pro-sft`. + - **GLM (MIT):** `glm-5.1-sft`. + - **Kimi (Modified MIT):** `kimi-k2.5-sft`, `kimi-k2.6-sft`. + - **MiniMax (MiniMax Community License — commercial use needs a separate + agreement):** `minimax-m3-sft`. + - **Mistral Large 3 (Apache-2.0, 675B/41B-active multimodal MoE):** + `mistral-large-3-sft`. - Unit-test coverage for the `warmup.py` auto-warmup-steps helper ([#274](https://github.com/MakazhanAlpamys/Soup/pull/274) by [@shatakshi-1404](https://github.com/shatakshi-1404)). +### Fixed +- **Stale recipe repo-ID:** `glm-5-sft` now points at `zai-org/GLM-5` (the org + migrated from `THUDM`). + ## [0.71.23] - 2026-06-12 ### Added diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index c2575a6..1bda23c 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -110,7 +110,7 @@ src/soup_cli/ experiment/ - SQLite experiment tracking eval/ - Eval platform (custom tasks, LLM judge, human eval, leaderboard) migrate/ - Config migration (LLaMA-Factory, Axolotl, Unsloth) - recipes/ - Ready-made configs for popular models (116 recipes) + recipes/ - Ready-made configs for popular models (133 recipes) autopilot/ - Zero-config decision engine (v0.25.0) registry/ - Model Registry (hashing, store, diff, attach) (v0.26.0 + v0.33.0) cans/ - Shareable .can artifact format + run/publish orchestrator (v0.26.0 + v0.33.0) @@ -120,7 +120,7 @@ src/soup_cli/ templates/ - 17 built-in soup.yaml templates (YAML + manifest.json) with load_template loader (v0.39.0, +bco v0.40.0) ui/ - Web UI (FastAPI + HTML/JS SPA) -tests/ - Test suite (294 files, 14278 tests) +tests/ - Test suite (296 files, 14514 tests) examples/ - Real-world config examples and datasets ``` @@ -150,7 +150,7 @@ pytest tests/test_data.py::test_detect_alpaca_format -v pytest tests/ --cov=soup_cli --cov-report=html ``` -### Test Files (293 files) +### Test Files (296 files) > A representative sample of the suite below. The full table lives in > [`.claude/CLAUDE.md`](.claude/CLAUDE.md); run `pytest tests/ -v` for the complete list. diff --git a/README.md b/README.md index 6b9aa4c..b47f277 100644 --- a/README.md +++ b/README.md @@ -49,19 +49,17 @@ infrastructure instead of improving models. Soup fixes that. ## What's New -**v0.71.23 — Native Spectrum targeted training.** Fine-tune only the layers that matter — no 8×H100 needed to find them: +**v0.71.24 — 2026 model-family recipes (catalog 116 → 133).** 17 new ready-made SFT recipes for the open-weight models that shipped Feb–Jun 2026: -- **`soup spectrum scan --model --top-percent 50`** — streams a model's weights one - tensor at a time (no model load, runs on a CPU box even for very large models) and computes a - singular-value signal-to-noise ratio per weight matrix (Marchenko-Pastur, arXiv:2406.06623). -- **Ready-to-paste patch** — prints a per-group SNR table and a `training.unfrozen_parameters` - YAML block; pipe it straight into your `soup.yaml`. Results cache under `~/.soup/spectrum/`. -- **Targeted full fine-tuning** — with `training.unfrozen_parameters` set, `soup train` freezes - every parameter and unfreezes only the matched high-SNR layers (full FT, LoRA off) — train - fewer parameters, keep more of the base model intact. -- **Honest guards** — the scan is pure-numpy and transpose-invariant; patterns are ReDoS-checked, - Hub downloads go through the SSRF-hardened loader, and the config gates `quantization: none` + - `task: sft` so a mis-set flag fails loudly, not silently. +- **Qwen 3.5 family (Apache-2.0)** — `qwen3.5-0.8b/2b/4b/9b/27b-sft` dense, plus the + `35b-a3b` / `122b-a10b` / `397b-a17b` MoE sizes (262K context, native vision). +- **Qwen 3.6 (Apache-2.0)** — `qwen3.6-27b-sft` + `qwen3.6-35b-a3b-sft`. +- **Frontier MoE** — `deepseek-v4-flash-sft` / `deepseek-v4-pro-sft` (MIT), + `glm-5.1-sft` (MIT), `kimi-k2.5-sft` / `kimi-k2.6-sft` (Modified MIT), + `minimax-m3-sft` (MiniMax Community License), `mistral-large-3-sft` (Apache-2.0). +- **Stale repo-ID fix** — `glm-5-sft` now points at `zai-org/GLM-5` (the org migrated from `THUDM`). +- Every base repo-ID was verified to resolve on Hugging Face. Grab one with + `soup recipes use ` or browse all 133 via `soup recipes list`. Full history: [CHANGELOG.md](CHANGELOG.md) · [GitHub Releases](https://github.com/MakazhanAlpamys/Soup/releases). diff --git a/docs/commands.md b/docs/commands.md index 0d327fc..a2401cb 100644 --- a/docs/commands.md +++ b/docs/commands.md @@ -129,7 +129,7 @@ soup migrate --from llamafactory config.yaml Import config from LLaMA-Factory soup migrate --from axolotl config.yml Import config from Axolotl soup migrate --from unsloth notebook.ipynb Import config from Unsloth notebook soup migrate --from llamafactory c.yaml --dry-run Preview without writing -soup recipes list List all 116 ready-made recipes +soup recipes list List all 133 ready-made recipes soup recipes show llama3.1-8b-sft Print recipe YAML soup recipes use llama3.1-8b-sft Copy recipe to soup.yaml soup recipes search "reasoning" Search by keyword/task/size diff --git a/docs/models.md b/docs/models.md index 803db59..38d25fa 100644 --- a/docs/models.md +++ b/docs/models.md @@ -16,11 +16,15 @@ Soup works with **any** of the **340,000+** text-generation models on [HuggingFa | **Llama 3.x** | Llama-3.1-8B-Instruct, Llama-3.3-70B-Instruct | 1B–70B | Chat, instruction following | | **Llama 3.2 Vision** | Llama-3.2-11B-Vision-Instruct, Llama-3.2-90B-Vision | 11B–90B | Image understanding | | **Gemma 3** | Gemma-3-4B-IT, Gemma-3-9B-IT, Gemma-3-27B-IT | 4B–27B | Efficient, multilingual | +| **Qwen 3.5 / 3.6** | Qwen3.5-0.8B…397B-A17B, Qwen3.6-27B, Qwen3.6-35B-A3B | 0.8B–397B | 262K context, native vision, MoE | | **Qwen 3** | Qwen3-8B, Qwen3-14B, Qwen3-32B, Qwen3-235B-A22B | 0.6B–235B | Reasoning, code, MoE | | **Qwen 2.5** | Qwen2.5-7B-Instruct, Qwen2.5-Coder-32B-Instruct | 0.5B–72B | Code, math | -| **DeepSeek** | DeepSeek-R1-Distill-Llama-8B, DeepSeek-V3-0324 | 1.5B–671B | Reasoning (GRPO), code | +| **DeepSeek** | DeepSeek-R1-Distill-Llama-8B, DeepSeek-V3-0324, DeepSeek-V4-Flash/Pro | 1.5B–1.6T | Reasoning (GRPO), code, MoE | +| **GLM** | GLM-5, GLM-5.1 | 9B–754B | Chinese + English, MoE | +| **Kimi** | Kimi-K2, Kimi-K2.5, Kimi-K2.6 | ~1T (MoE) | Long-context agentic, MoE | +| **MiniMax** | MiniMax-M2, MiniMax-M3 | 230B–428B | Agentic, MoE (community license) | | **Phi-4** | Phi-4-14B, Phi-4-mini-reasoning | 3.8B–14B | Compact reasoning | -| **Mistral** | Mistral-7B-Instruct-v0.3, Mistral-Small-24B-Instruct | 7B–24B | Fast, efficient | +| **Mistral** | Mistral-7B-Instruct-v0.3, Mistral-Small-24B, Mistral-Large-3 | 7B–675B | Fast, efficient, MoE | | **Mixtral** | Mixtral-8x7B-Instruct-v0.1, Mixtral-8x22B | 47B–141B | MoE architecture | | **CodeLlama** | CodeLlama-7b-Instruct-hf, CodeLlama-34b-Instruct | 7B–34B | Code generation | | **StarCoder 2** | StarCoder2-15B, StarCoder2-7B | 3B–15B | Code completion | diff --git a/docs/serving-and-export.md b/docs/serving-and-export.md index 5cd4773..d2cd887 100644 --- a/docs/serving-and-export.md +++ b/docs/serving-and-export.md @@ -418,7 +418,7 @@ soup ui **Pages:** - **Dashboard** — view all experiment runs, loss charts, system info, multi-run comparison -- **New Training** — create configs from templates or 116 ready-made recipes, validate, start training with live SSE log streaming and progress bar +- **New Training** — create configs from templates or 133 ready-made recipes, validate, start training with live SSE log streaming and progress bar - **Data Explorer** — browse and inspect datasets (JSONL, JSON, CSV, Parquet) - **Model Chat** — chat with streaming responses, configurable temperature/top_p/max_tokens, system prompt, adapter selection, markdown rendering, chat export @@ -427,7 +427,7 @@ soup ui - **Enhanced Metrics** — 2x2 chart grid (loss, LR, grad_norm, throughput) + GPU memory chart, eval results table - **Multi-Run Compare** — overlay loss curves from up to 5 runs side-by-side - **Chat Upgrade** — SSE streaming via proxy, typing indicator, cancel button, markdown renderer (bold, italic, code blocks), chat export as JSON -- **Config Builder** — recipe dropdown (116 recipes), config schema API for dynamic form generation +- **Config Builder** — recipe dropdown (133 recipes), config schema API for dynamic form generation **Security:** The Web UI generates a random auth token at startup (printed to console). All mutating endpoints (start/stop training, delete runs, inspect data, validate config) require `Authorization: Bearer ` header. CORS is restricted to the served origin. Data inspection is sandboxed to the working directory. diff --git a/pyproject.toml b/pyproject.toml index a67b4e0..85043d2 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "soup-cli" -version = "0.71.23" +version = "0.71.24" description = "Fine-tune and post-train LLMs in one command. No SSH, no config hell." readme = "README.md" license = "Apache-2.0" diff --git a/src/soup_cli/__init__.py b/src/soup_cli/__init__.py index 5a1909d..572a4a3 100644 --- a/src/soup_cli/__init__.py +++ b/src/soup_cli/__init__.py @@ -1,3 +1,3 @@ """Soup CLI — Fine-tune and post-train LLMs in one command.""" -__version__ = "0.71.23" +__version__ = "0.71.24" diff --git a/src/soup_cli/recipes/catalog.py b/src/soup_cli/recipes/catalog.py index c91de76..ba24f68 100644 --- a/src/soup_cli/recipes/catalog.py +++ b/src/soup_cli/recipes/catalog.py @@ -48,7 +48,7 @@ def search_recipes( # --------------------------------------------------------------------------- -# Recipe catalog (116 recipes) +# Recipe catalog (133 recipes) # --------------------------------------------------------------------------- RECIPES: Dict[str, RecipeMeta] = { @@ -2594,13 +2594,13 @@ output: ./output """, ), "glm-5-sft": RecipeMeta( - model="THUDM/glm-5", + model="zai-org/GLM-5", task="sft", size="9B", - tags=("glm", "thudm", "chat", "next-gen"), + tags=("glm", "zai-org", "chat", "next-gen"), description="GLM 5 SFT (next-gen GLM family)", yaml_str="""\ -base: THUDM/glm-5 +base: zai-org/GLM-5 task: sft data: @@ -3536,6 +3536,553 @@ training: alpha: 32 target_modules: auto +output: ./output +""", + ), + # ------------------------------------------------------------------ + # v0.71.24 — 2026 model-family expansion (catalog 116 -> 133) + # 17 SFT recipes for the open-weight models released Feb-Jun 2026. + # Every base repo-ID verified to resolve on Hugging Face. + # ------------------------------------------------------------------ + "qwen3.5-0.8b-sft": RecipeMeta( + model="Qwen/Qwen3.5-0.8B", + task="sft", + size="0.8B", + tags=("qwen", "qwen3.5", "sft", "tiny", "edge", "mobile"), + description="Qwen 3.5 0.8B SFT (Apache-2.0, tiny / mobile)", + yaml_str="""\ +base: Qwen/Qwen3.5-0.8B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 2048 + +training: + epochs: 3 + lr: 3e-4 + batch_size: auto + lora: + r: 8 + alpha: 16 + target_modules: auto + quantization: 4bit + +output: ./output +""", + ), + "qwen3.5-2b-sft": RecipeMeta( + model="Qwen/Qwen3.5-2B", + task="sft", + size="2B", + tags=("qwen", "qwen3.5", "sft", "small", "edge"), + description="Qwen 3.5 2B SFT (Apache-2.0, small / edge)", + yaml_str="""\ +base: Qwen/Qwen3.5-2B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 2048 + +training: + epochs: 3 + lr: 2e-4 + batch_size: auto + lora: + r: 8 + alpha: 16 + target_modules: auto + quantization: 4bit + +output: ./output +""", + ), + "qwen3.5-4b-sft": RecipeMeta( + model="Qwen/Qwen3.5-4B", + task="sft", + size="4B", + tags=("qwen", "qwen3.5", "sft", "small"), + description="Qwen 3.5 4B SFT (Apache-2.0, 262K context)", + yaml_str="""\ +base: Qwen/Qwen3.5-4B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 3 + lr: 2e-4 + batch_size: auto + lora: + r: 16 + alpha: 32 + target_modules: auto + quantization: 4bit + +output: ./output +""", + ), + "qwen3.5-9b-sft": RecipeMeta( + model="Qwen/Qwen3.5-9B", + task="sft", + size="9B", + tags=("qwen", "qwen3.5", "sft", "chat", "instruction"), + description="Qwen 3.5 9B SFT (Apache-2.0, 262K context)", + yaml_str="""\ +base: Qwen/Qwen3.5-9B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 3 + lr: 2e-4 + batch_size: auto + lora: + r: 16 + alpha: 32 + target_modules: auto + quantization: 4bit + +output: ./output +""", + ), + "qwen3.5-27b-sft": RecipeMeta( + model="Qwen/Qwen3.5-27B", + task="sft", + size="27B", + tags=("qwen", "qwen3.5", "sft", "large", "deepspeed"), + description="Qwen 3.5 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2", + yaml_str="""\ +base: Qwen/Qwen3.5-27B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 3 + lr: 1e-5 + batch_size: auto + gradient_accumulation_steps: 8 + lora: + r: 16 + alpha: 32 + target_modules: auto + quantization: 4bit + +output: ./output +""", + ), + "qwen3.5-35b-a3b-sft": RecipeMeta( + model="Qwen/Qwen3.5-35B-A3B", + task="sft", + size="35B", + tags=("qwen", "qwen3.5", "sft", "moe", "mixture-of-experts"), + description="Qwen 3.5 35B-A3B MoE SFT (Apache-2.0, 3B active)", + yaml_str="""\ +base: Qwen/Qwen3.5-35B-A3B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 3 + lr: 1e-4 + batch_size: auto + gradient_accumulation_steps: 8 + lora: + r: 16 + alpha: 32 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + +output: ./output +""", + ), + "qwen3.5-122b-a10b-sft": RecipeMeta( + model="Qwen/Qwen3.5-122B-A10B", + task="sft", + size="122B", + tags=("qwen", "qwen3.5", "sft", "moe", "large", "multi-gpu"), + description=( + "Qwen 3.5 122B-A10B MoE SFT (Apache-2.0, 10B active). " + "Multi-GPU recommended (8 x A100/H100 80GB)." + ), + yaml_str="""\ +base: Qwen/Qwen3.5-122B-A10B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 1 + lr: 1e-5 + batch_size: 1 + gradient_accumulation_steps: 16 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + +output: ./output +""", + ), + "qwen3.5-397b-a17b-sft": RecipeMeta( + model="Qwen/Qwen3.5-397B-A17B", + task="sft", + size="397B", + tags=("qwen", "qwen3.5", "sft", "moe", "large", "multi-gpu"), + description=( + "Qwen 3.5 397B-A17B flagship MoE SFT (Apache-2.0, 17B active). " + "Requires multi-node DeepSpeed / FSDP." + ), + yaml_str="""\ +base: Qwen/Qwen3.5-397B-A17B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 1 + lr: 5e-6 + batch_size: 1 + gradient_accumulation_steps: 32 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + +output: ./output +""", + ), + "qwen3.6-27b-sft": RecipeMeta( + model="Qwen/Qwen3.6-27B", + task="sft", + size="27B", + tags=("qwen", "qwen3.6", "sft", "large", "deepspeed"), + description="Qwen 3.6 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2", + yaml_str="""\ +base: Qwen/Qwen3.6-27B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 3 + lr: 1e-5 + batch_size: auto + gradient_accumulation_steps: 8 + lora: + r: 16 + alpha: 32 + target_modules: auto + quantization: 4bit + +output: ./output +""", + ), + "qwen3.6-35b-a3b-sft": RecipeMeta( + model="Qwen/Qwen3.6-35B-A3B", + task="sft", + size="35B", + tags=("qwen", "qwen3.6", "sft", "moe", "mixture-of-experts"), + description="Qwen 3.6 35B-A3B MoE SFT (Apache-2.0, 3B active)", + yaml_str="""\ +base: Qwen/Qwen3.6-35B-A3B +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 3 + lr: 1e-4 + batch_size: auto + gradient_accumulation_steps: 8 + lora: + r: 16 + alpha: 32 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + +output: ./output +""", + ), + "deepseek-v4-flash-sft": RecipeMeta( + model="deepseek-ai/DeepSeek-V4-Flash", + task="sft", + size="N/A", + tags=("deepseek", "deepseek-v4", "sft", "moe", "mixture-of-experts"), + description="DeepSeek V4 Flash MoE SFT (MIT, efficiency-tier)", + yaml_str="""\ +base: deepseek-ai/DeepSeek-V4-Flash +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 3 + lr: 1e-4 + batch_size: auto + gradient_accumulation_steps: 8 + lora: + r: 16 + alpha: 32 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + +output: ./output +""", + ), + "deepseek-v4-pro-sft": RecipeMeta( + model="deepseek-ai/DeepSeek-V4-Pro", + task="sft", + size="N/A", + tags=("deepseek", "deepseek-v4", "sft", "moe", "large", "multi-gpu"), + description=( + "DeepSeek V4 Pro flagship MoE SFT (MIT, 1.6T-class). " + "Requires multi-node DeepSpeed." + ), + yaml_str="""\ +base: deepseek-ai/DeepSeek-V4-Pro +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 1 + lr: 5e-6 + batch_size: 1 + gradient_accumulation_steps: 32 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + +output: ./output +""", + ), + "glm-5.1-sft": RecipeMeta( + model="zai-org/GLM-5.1", + task="sft", + size="754B", + tags=("glm", "zai-org", "sft", "moe", "large", "multi-gpu"), + description=( + "GLM 5.1 MoE SFT (MIT, 754B). Multi-GPU / multi-node recommended." + ), + yaml_str="""\ +base: zai-org/GLM-5.1 +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 8192 + +training: + epochs: 1 + lr: 1e-5 + batch_size: 1 + gradient_accumulation_steps: 16 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + +output: ./output +""", + ), + "kimi-k2.5-sft": RecipeMeta( + model="moonshotai/Kimi-K2.5", + task="sft", + size="1T", + tags=("kimi", "moonshot", "sft", "moe", "large", "multi-gpu"), + description=( + "Kimi K2.5 MoE SFT (Modified MIT, ~1T / 32B active). " + "Requires multi-node DeepSpeed." + ), + yaml_str="""\ +base: moonshotai/Kimi-K2.5 +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 8192 + +training: + epochs: 1 + lr: 1e-5 + batch_size: 1 + gradient_accumulation_steps: 16 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + +output: ./output +""", + ), + "kimi-k2.6-sft": RecipeMeta( + model="moonshotai/Kimi-K2.6", + task="sft", + size="1T", + tags=("kimi", "moonshot", "sft", "moe", "large", "multi-gpu"), + description=( + "Kimi K2.6 MoE SFT (Modified MIT, ~1T / 32B active). " + "Requires multi-node DeepSpeed." + ), + yaml_str="""\ +base: moonshotai/Kimi-K2.6 +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 8192 + +training: + epochs: 1 + lr: 5e-6 + batch_size: 1 + gradient_accumulation_steps: 32 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + +output: ./output +""", + ), + "minimax-m3-sft": RecipeMeta( + model="MiniMaxAI/MiniMax-M3", + task="sft", + size="428B", + tags=("minimax", "sft", "moe", "large", "multi-gpu"), + description=( + "MiniMax M3 MoE SFT (428B / 23B active). MiniMax Community License " + "- commercial use requires a separate agreement. Multi-GPU recommended." + ), + yaml_str="""\ +base: MiniMaxAI/MiniMax-M3 +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 1 + lr: 1e-5 + batch_size: 1 + gradient_accumulation_steps: 16 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + +output: ./output +""", + ), + "mistral-large-3-sft": RecipeMeta( + model="mistralai/Mistral-Large-3-675B-Instruct-2512", + task="sft", + size="675B", + tags=("mistral", "mistral-large", "sft", "moe", "large", "multi-gpu"), + description=( + "Mistral Large 3 MoE SFT (Apache-2.0, 675B / 41B active, multimodal). " + "Requires multi-node DeepSpeed." + ), + yaml_str="""\ +base: mistralai/Mistral-Large-3-675B-Instruct-2512 +task: sft + +data: + train: ./data/train.jsonl + format: auto + max_length: 4096 + +training: + epochs: 1 + lr: 5e-6 + batch_size: 1 + gradient_accumulation_steps: 32 + lora: + r: 32 + alpha: 64 + target_modules: auto + quantization: 4bit + moe_lora: true + moe_aux_loss_coeff: 0.01 + gradient_checkpointing: true + output: ./output """, ), diff --git a/tests/test_recipes.py b/tests/test_recipes.py index 0aecb48..d082663 100644 --- a/tests/test_recipes.py +++ b/tests/test_recipes.py @@ -260,7 +260,7 @@ class TestV025NewRecipes: assert cfg.base == recipe.model assert cfg.task == recipe.task - def test_catalog_size_is_116(self): + def test_catalog_size_is_133(self): """Total catalog size — grew with each release. v0.25.0 shipped 43 recipes (29 + 9 Part A + 2 Part B tools + 3 Part E MLX). @@ -270,10 +270,11 @@ class TestV025NewRecipes: v0.52.0 added 6 (5 TTS + Falcon-E BitNet) -> 112. v0.53.5 added 1 (deepseek-v3-reasoning) -> 113. v0.62.0 added 3 (raft-llama3-8b, ra-dit-retriever, ra-dit-llama3-8b) -> 116. + v0.71.24 added 17 (2026 model-family expansion) -> 133. """ from soup_cli.recipes.catalog import RECIPES - assert len(RECIPES) == 116 + assert len(RECIPES) == 133 def test_new_recipes_searchable(self): """Search returns the new recipes via keyword/task filter.""" diff --git a/tests/test_v07124.py b/tests/test_v07124.py new file mode 100644 index 0000000..5e00d46 --- /dev/null +++ b/tests/test_v07124.py @@ -0,0 +1,291 @@ +"""v0.71.24 — 2026 model-family recipe expansion (catalog 116 -> 133). + +Pure config/data release: 17 new ready-made SFT recipes for the open-weight +models released Feb-Jun 2026 (Qwen3.5 / Qwen3.6 / DeepSeek-V4 / GLM-5.1 / +Kimi-K2.5/K2.6 / MiniMax-M3 / Mistral-Large-3), plus a stale-repo-ID fix for +``glm-5-sft`` (``THUDM/glm-5`` -> ``zai-org/GLM-5``). CI validates each recipe by +``load_config_from_string`` parse only (no network). +""" +from __future__ import annotations + +import pytest +import yaml + +from soup_cli.config.loader import load_config_from_string +from soup_cli.config.schema import SoupConfig +from soup_cli.recipes.catalog import RECIPES, get_recipe, list_recipes, search_recipes + +# --------------------------------------------------------------------------- +# Step A — 17 new SFT recipes (every base verified to resolve on Hugging Face) +# --------------------------------------------------------------------------- + +V07124_RECIPE_NAMES = [ + # Qwen3.5 family (Apache-2.0) + "qwen3.5-0.8b-sft", + "qwen3.5-2b-sft", + "qwen3.5-4b-sft", + "qwen3.5-9b-sft", + "qwen3.5-27b-sft", + "qwen3.5-35b-a3b-sft", + "qwen3.5-122b-a10b-sft", + "qwen3.5-397b-a17b-sft", + # Qwen3.6 (Apache-2.0) + "qwen3.6-27b-sft", + "qwen3.6-35b-a3b-sft", + # DeepSeek-V4 (MIT) + "deepseek-v4-flash-sft", + "deepseek-v4-pro-sft", + # GLM (MIT, zai-org) + "glm-5.1-sft", + # Kimi (Modified MIT) + "kimi-k2.5-sft", + "kimi-k2.6-sft", + # MiniMax (MiniMax Community License) + "minimax-m3-sft", + # Mistral Large 3 (Apache-2.0) + "mistral-large-3-sft", +] + +# name -> expected base repo-ID (the verified Hugging Face repo) +V07124_RECIPE_BASES = { + "qwen3.5-0.8b-sft": "Qwen/Qwen3.5-0.8B", + "qwen3.5-2b-sft": "Qwen/Qwen3.5-2B", + "qwen3.5-4b-sft": "Qwen/Qwen3.5-4B", + "qwen3.5-9b-sft": "Qwen/Qwen3.5-9B", + "qwen3.5-27b-sft": "Qwen/Qwen3.5-27B", + "qwen3.5-35b-a3b-sft": "Qwen/Qwen3.5-35B-A3B", + "qwen3.5-122b-a10b-sft": "Qwen/Qwen3.5-122B-A10B", + "qwen3.5-397b-a17b-sft": "Qwen/Qwen3.5-397B-A17B", + "qwen3.6-27b-sft": "Qwen/Qwen3.6-27B", + "qwen3.6-35b-a3b-sft": "Qwen/Qwen3.6-35B-A3B", + "deepseek-v4-flash-sft": "deepseek-ai/DeepSeek-V4-Flash", + "deepseek-v4-pro-sft": "deepseek-ai/DeepSeek-V4-Pro", + "glm-5.1-sft": "zai-org/GLM-5.1", + "kimi-k2.5-sft": "moonshotai/Kimi-K2.5", + "kimi-k2.6-sft": "moonshotai/Kimi-K2.6", + "minimax-m3-sft": "MiniMaxAI/MiniMax-M3", + "mistral-large-3-sft": "mistralai/Mistral-Large-3-675B-Instruct-2512", +} + +# The MoE bases added in v0.71.24 — each recipe must enable moe_lora + aux loss. +# (Scope is intentionally limited to recipes added in THIS release, not every MoE +# recipe in the catalog.) +V07124_MOE_RECIPES = [ + "qwen3.5-35b-a3b-sft", + "qwen3.5-122b-a10b-sft", + "qwen3.5-397b-a17b-sft", + "qwen3.6-35b-a3b-sft", + "deepseek-v4-flash-sft", + "deepseek-v4-pro-sft", + "glm-5.1-sft", + "kimi-k2.5-sft", + "kimi-k2.6-sft", + "minimax-m3-sft", + "mistral-large-3-sft", +] + + +class TestV07124Recipes: + def test_recipe_count_target(self) -> None: + assert len(V07124_RECIPE_NAMES) == 17 + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_registered(self, name: str) -> None: + assert name in RECIPES, f"Recipe {name!r} missing from catalog" + assert get_recipe(name) is not None + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_base_matches_verified_repo(self, name: str) -> None: + assert RECIPES[name].model == V07124_RECIPE_BASES[name] + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_is_sft(self, name: str) -> None: + assert RECIPES[name].task == "sft" + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_metadata_well_formed(self, name: str) -> None: + meta = RECIPES[name] + assert meta.model, f"{name}: model is empty" + assert meta.size, f"{name}: size is empty" + assert meta.tags, f"{name}: tags is empty" + assert meta.description, f"{name}: description is empty" + assert meta.yaml_str, f"{name}: yaml_str is empty" + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_yaml_parses_as_soup_config(self, name: str) -> None: + cfg = load_config_from_string(RECIPES[name].yaml_str) + assert isinstance(cfg, SoupConfig) + assert cfg.base == RECIPES[name].model + assert cfg.task == "sft" + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_yaml_is_safe_loadable(self, name: str) -> None: + parsed = yaml.safe_load(RECIPES[name].yaml_str) + assert isinstance(parsed, dict) + assert "base" in parsed and "task" in parsed + # The inline ``base:`` / ``task:`` in the YAML must match the RecipeMeta. + assert parsed["base"] == RECIPES[name].model + assert parsed["task"] == RECIPES[name].task + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_model_id_no_null_or_whitespace(self, name: str) -> None: + meta = RECIPES[name] + assert "\x00" not in meta.model + assert " " not in meta.model + # No path-traversal segments or shell metacharacters in a repo id. + assert ".." not in meta.model, f"path-traversal segment in {meta.model!r}" + assert not any(c in meta.model for c in ';|&$`<>(){}*?!\\\n\t'), ( + f"shell metacharacter in {meta.model!r}" + ) + parts = meta.model.split("/") + assert len(parts) == 2 + assert all(p for p in parts), f"empty component in {meta.model!r}" + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_max_length_within_bounds(self, name: str) -> None: + cfg = load_config_from_string(RECIPES[name].yaml_str) + assert 64 <= cfg.data.max_length <= 1_048_576 + + @pytest.mark.parametrize("name", V07124_MOE_RECIPES) + def test_moe_recipes_enable_moe_lora(self, name: str) -> None: + cfg = load_config_from_string(RECIPES[name].yaml_str) + assert cfg.training.moe_lora is True, f"{name} should set moe_lora: true" + + @pytest.mark.parametrize("name", V07124_MOE_RECIPES) + def test_moe_recipes_set_aux_loss_coeff(self, name: str) -> None: + # Every new MoE recipe sets the aux-loss coefficient explicitly (0.01) so + # the whole batch is uniformly self-documenting (review HIGH-1). + cfg = load_config_from_string(RECIPES[name].yaml_str) + assert abs(cfg.training.moe_aux_loss_coeff - 0.01) < 1e-9, ( + f"{name} should set moe_aux_loss_coeff: 0.01, got " + f"{cfg.training.moe_aux_loss_coeff}" + ) + + @pytest.mark.parametrize("name", sorted(set(V07124_RECIPE_NAMES) - set(V07124_MOE_RECIPES))) + def test_dense_recipes_do_not_set_moe_lora(self, name: str) -> None: + # The complement of the MoE list: a dense recipe must NOT enable moe_lora + # (guards against pasting a MoE block into a dense recipe). + cfg = load_config_from_string(RECIPES[name].yaml_str) + assert cfg.training.moe_lora is not True, ( + f"{name} is a dense recipe but has moe_lora: true" + ) + + @pytest.mark.parametrize("name", V07124_MOE_RECIPES) + def test_moe_recipes_tagged_moe(self, name: str) -> None: + assert "moe" in RECIPES[name].tags, ( + f"{name} is a MoE recipe but lacks the 'moe' tag" + ) + + @pytest.mark.parametrize("name", [n for n in V07124_RECIPE_NAMES if n.startswith("qwen")]) + def test_qwen_recipes_tagged_qwen(self, name: str) -> None: + assert "qwen" in RECIPES[name].tags, f"{name} lacks the 'qwen' tag" + + def test_search_by_task_sft_includes_all_new_recipes(self) -> None: + sft_models = {r.model for r in search_recipes(task="sft")} + for name in V07124_RECIPE_NAMES: + assert RECIPES[name].model in sft_models, ( + f"{name} not returned by search_recipes(task='sft')" + ) + + def test_kimi_k2_5_size_known(self) -> None: + # K2.5 is a ~1T/32B MoE (verified on the model card) — not "N/A". + assert RECIPES["kimi-k2.5-sft"].size == "1T" + + def test_kimi_k2_6_size_known(self) -> None: + # K2.6 is a ~1T/32B MoE (verified on the model card). + assert RECIPES["kimi-k2.6-sft"].size == "1T" + + def test_minimax_recipe_notes_community_license(self) -> None: + desc = RECIPES["minimax-m3-sft"].description.lower() + assert "minimax" in desc + assert "license" in desc + # commercial-use caveat must be surfaced + assert "commercial" in desc + + def test_mistral_large_3_notes_apache_license(self) -> None: + # Plan guessed MRL; the published card is Apache-2.0 — describe accurately. + desc = RECIPES["mistral-large-3-sft"].description.lower() + assert "apache" in desc + + @pytest.mark.parametrize("name", ["kimi-k2.5-sft", "kimi-k2.6-sft"]) + def test_kimi_recipes_note_modified_mit(self, name: str) -> None: + assert "mit" in RECIPES[name].description.lower() + + def test_deepseek_v4_sizes_intentionally_na(self) -> None: + # DeepSeek has not publicly confirmed the V4 parameter counts; "N/A" is the + # honest value. This documents the deliberate choice (vs Kimi K2.5/K2.6, + # whose ~1T sizes ARE on the card). + assert RECIPES["deepseek-v4-flash-sft"].size == "N/A" + assert RECIPES["deepseek-v4-pro-sft"].size == "N/A" + + @pytest.mark.parametrize("name", V07124_RECIPE_NAMES) + def test_recipe_searchable(self, name: str) -> None: + # The recipe key is part of the searchable text, so a query for the recipe's + # own name must return that exact recipe (deterministic — no token heuristic). + results = search_recipes(query=name) + assert any(r.model == RECIPES[name].model for r in results), ( + f"{name} not returned by search_recipes(query={name!r})" + ) + + +# --------------------------------------------------------------------------- +# Step B — stale repo-ID fix: glm-5-sft (THUDM/glm-5 -> zai-org/GLM-5) +# --------------------------------------------------------------------------- + + +class TestGlm5RepoIdFix: + def test_glm5_uses_zai_org(self) -> None: + meta = RECIPES["glm-5-sft"] + assert meta.model == "zai-org/GLM-5" + assert "THUDM" not in meta.model + + def test_glm5_yaml_base_updated(self) -> None: + meta = RECIPES["glm-5-sft"] + parsed = yaml.safe_load(meta.yaml_str) + assert parsed["base"] == "zai-org/GLM-5" + assert "THUDM/glm-5" not in meta.yaml_str + + def test_glm5_old_key_still_exists(self) -> None: + # The fix RENAMES the base field; it must NOT delete the old recipe key. + assert "glm-5-sft" in RECIPES, "glm-5-sft was accidentally removed" + assert "glm-5.1-sft" in RECIPES, "glm-5.1-sft is the new v0.71.24 recipe" + + def test_glm5_and_glm51_are_distinct_entries(self) -> None: + assert RECIPES["glm-5-sft"].model != RECIPES["glm-5.1-sft"].model + + def test_no_recipe_references_thudm_glm5(self) -> None: + for name, meta in RECIPES.items(): + assert "THUDM/glm-5" not in meta.model, ( + f"{name}: model still references THUDM/glm-5" + ) + assert "THUDM/glm-5" not in meta.yaml_str, ( + f"{name}: yaml_str still references THUDM/glm-5" + ) + + +# --------------------------------------------------------------------------- +# Catalog count — 116 -> 133 +# --------------------------------------------------------------------------- + + +class TestCatalogCount: + def test_total_recipe_count_is_133(self) -> None: + assert len(RECIPES) == 133 + + def test_list_recipes_matches_dict(self) -> None: + assert len(list_recipes()) == len(RECIPES) + + def test_all_recipes_parse(self) -> None: + # Defence-in-depth — the whole catalog must still parse after the additions. + for name, meta in RECIPES.items(): + cfg = load_config_from_string(meta.yaml_str) + assert cfg.base, name + + def test_all_recipes_yaml_base_matches_model(self) -> None: + # Catalog-wide integrity invariant — the inline ``base:`` must equal + # RecipeMeta.model for EVERY recipe (catches a half-applied repo-ID fix). + for name, meta in RECIPES.items(): + parsed = yaml.safe_load(meta.yaml_str) + assert parsed.get("base") == meta.model, ( + f"{name}: yaml base {parsed.get('base')!r} != model {meta.model!r}" + )