feat(recipes): 2026 model-family expansion — 17 SFT recipes, catalog 116→133 (v0.71.24)

Add ready-made SFT recipes for the open-weight models released Feb–Jun 2026,
each base repo-ID verified to resolve on Hugging Face:
- Qwen3.5 0.8B/2B/4B/9B/27B + MoE 35B-A3B/122B-A10B/397B-A17B (Apache-2.0)
- Qwen3.6 27B + 35B-A3B (Apache-2.0)
- DeepSeek-V4 Flash/Pro (MIT), GLM-5.1 (MIT)
- Kimi-K2.5/K2.6 (Modified MIT), MiniMax-M3 (MiniMax Community License)
- Mistral-Large-3 (Apache-2.0, 675B/41B-active multimodal MoE)

Fix stale glm-5-sft repo-ID THUDM/glm-5 -> zai-org/GLM-5 (org migration).
+220 tests (tests/test_v07124.py). Catalog count 116 -> 133.
This commit is contained in:
Alpamys 2026-06-21 13:00:47 +05:00
parent da81b61798
commit 418f86390a
11 changed files with 890 additions and 28 deletions

View File

@ -12,10 +12,31 @@ reproducing 70+ versions of notes.
## [Unreleased]
## [0.71.24] - 2026-06-21
### Added
- **2026 model-family recipe expansion (catalog 116 → 133).** 17 new ready-made
SFT recipes for the open-weight models released Feb–Jun 2026, each with its
Hugging Face repo-ID verified to resolve:
- **Qwen 3.5 (Apache-2.0):** `qwen3.5-0.8b-sft`, `qwen3.5-2b-sft`,
`qwen3.5-4b-sft`, `qwen3.5-9b-sft`, `qwen3.5-27b-sft`, and the
`qwen3.5-35b-a3b-sft` / `qwen3.5-122b-a10b-sft` / `qwen3.5-397b-a17b-sft`
MoE sizes.
- **Qwen 3.6 (Apache-2.0):** `qwen3.6-27b-sft`, `qwen3.6-35b-a3b-sft`.
- **DeepSeek-V4 (MIT):** `deepseek-v4-flash-sft`, `deepseek-v4-pro-sft`.
- **GLM (MIT):** `glm-5.1-sft`.
- **Kimi (Modified MIT):** `kimi-k2.5-sft`, `kimi-k2.6-sft`.
- **MiniMax (MiniMax Community License — commercial use needs a separate
agreement):** `minimax-m3-sft`.
- **Mistral Large 3 (Apache-2.0, 675B/41B-active multimodal MoE):**
`mistral-large-3-sft`.
- Unit-test coverage for the `warmup.py` auto-warmup-steps helper
([#274](https://github.com/MakazhanAlpamys/Soup/pull/274) by [@shatakshi-1404](https://github.com/shatakshi-1404)).
### Fixed
- **Stale recipe repo-ID:** `glm-5-sft` now points at `zai-org/GLM-5` (the org
migrated from `THUDM`).
## [0.71.23] - 2026-06-12
### Added

View File

@ -110,7 +110,7 @@ src/soup_cli/
experiment/ - SQLite experiment tracking
eval/ - Eval platform (custom tasks, LLM judge, human eval, leaderboard)
migrate/ - Config migration (LLaMA-Factory, Axolotl, Unsloth)
recipes/ - Ready-made configs for popular models (116 recipes)
recipes/ - Ready-made configs for popular models (133 recipes)
autopilot/ - Zero-config decision engine (v0.25.0)
registry/ - Model Registry (hashing, store, diff, attach) (v0.26.0 + v0.33.0)
cans/ - Shareable .can artifact format + run/publish orchestrator (v0.26.0 + v0.33.0)
@ -120,7 +120,7 @@ src/soup_cli/
templates/ - 17 built-in soup.yaml templates (YAML + manifest.json) with load_template loader (v0.39.0, +bco v0.40.0)
ui/ - Web UI (FastAPI + HTML/JS SPA)
tests/ - Test suite (294 files, 14278 tests)
tests/ - Test suite (296 files, 14514 tests)
examples/ - Real-world config examples and datasets
```
@ -150,7 +150,7 @@ pytest tests/test_data.py::test_detect_alpaca_format -v
pytest tests/ --cov=soup_cli --cov-report=html
```
### Test Files (293 files)
### Test Files (296 files)
> A representative sample of the suite below. The full table lives in
> [`.claude/CLAUDE.md`](.claude/CLAUDE.md); run `pytest tests/ -v` for the complete list.

View File

@ -49,19 +49,17 @@ infrastructure instead of improving models. Soup fixes that.
## What's New
**v0.71.23 — Native Spectrum targeted training.** Fine-tune only the layers that matter — no 8×H100 needed to find them:
**v0.71.24 — 2026 model-family recipes (catalog 116 → 133).** 17 new ready-made SFT recipes for the open-weight models that shipped Feb–Jun 2026:
- **`soup spectrum scan --model <id|path> --top-percent 50`** — streams a model's weights one
tensor at a time (no model load, runs on a CPU box even for very large models) and computes a
singular-value signal-to-noise ratio per weight matrix (Marchenko-Pastur, arXiv:2406.06623).
- **Ready-to-paste patch** — prints a per-group SNR table and a `training.unfrozen_parameters`
YAML block; pipe it straight into your `soup.yaml`. Results cache under `~/.soup/spectrum/`.
- **Targeted full fine-tuning** — with `training.unfrozen_parameters` set, `soup train` freezes
every parameter and unfreezes only the matched high-SNR layers (full FT, LoRA off) — train
fewer parameters, keep more of the base model intact.
- **Honest guards** — the scan is pure-numpy and transpose-invariant; patterns are ReDoS-checked,
Hub downloads go through the SSRF-hardened loader, and the config gates `quantization: none` +
`task: sft` so a mis-set flag fails loudly, not silently.
- **Qwen 3.5 family (Apache-2.0)** — `qwen3.5-0.8b/2b/4b/9b/27b-sft` dense, plus the
`35b-a3b` / `122b-a10b` / `397b-a17b` MoE sizes (262K context, native vision).
- **Qwen 3.6 (Apache-2.0)** — `qwen3.6-27b-sft` + `qwen3.6-35b-a3b-sft`.
- **Frontier MoE** — `deepseek-v4-flash-sft` / `deepseek-v4-pro-sft` (MIT),
`glm-5.1-sft` (MIT), `kimi-k2.5-sft` / `kimi-k2.6-sft` (Modified MIT),
`minimax-m3-sft` (MiniMax Community License), `mistral-large-3-sft` (Apache-2.0).
- **Stale repo-ID fix** — `glm-5-sft` now points at `zai-org/GLM-5` (the org migrated from `THUDM`).
- Every base repo-ID was verified to resolve on Hugging Face. Grab one with
`soup recipes use <name>` or browse all 133 via `soup recipes list`.
Full history: [CHANGELOG.md](CHANGELOG.md) &middot; [GitHub Releases](https://github.com/MakazhanAlpamys/Soup/releases).

View File

@ -129,7 +129,7 @@ soup migrate --from llamafactory config.yaml Import config from LLaMA-Factory
soup migrate --from axolotl config.yml Import config from Axolotl
soup migrate --from unsloth notebook.ipynb Import config from Unsloth notebook
soup migrate --from llamafactory c.yaml --dry-run Preview without writing
soup recipes list List all 116 ready-made recipes
soup recipes list List all 133 ready-made recipes
soup recipes show llama3.1-8b-sft Print recipe YAML
soup recipes use llama3.1-8b-sft Copy recipe to soup.yaml
soup recipes search "reasoning" Search by keyword/task/size

View File

@ -16,11 +16,15 @@ Soup works with **any** of the **340,000+** text-generation models on [HuggingFa
| **Llama 3.x** | Llama-3.1-8B-Instruct, Llama-3.3-70B-Instruct | 1B–70B | Chat, instruction following |
| **Llama 3.2 Vision** | Llama-3.2-11B-Vision-Instruct, Llama-3.2-90B-Vision | 11B–90B | Image understanding |
| **Gemma 3** | Gemma-3-4B-IT, Gemma-3-9B-IT, Gemma-3-27B-IT | 4B–27B | Efficient, multilingual |
| **Qwen 3.5 / 3.6** | Qwen3.5-0.8B…397B-A17B, Qwen3.6-27B, Qwen3.6-35B-A3B | 0.8B–397B | 262K context, native vision, MoE |
| **Qwen 3** | Qwen3-8B, Qwen3-14B, Qwen3-32B, Qwen3-235B-A22B | 0.6B–235B | Reasoning, code, MoE |
| **Qwen 2.5** | Qwen2.5-7B-Instruct, Qwen2.5-Coder-32B-Instruct | 0.5B–72B | Code, math |
| **DeepSeek** | DeepSeek-R1-Distill-Llama-8B, DeepSeek-V3-0324 | 1.5B–671B | Reasoning (GRPO), code |
| **DeepSeek** | DeepSeek-R1-Distill-Llama-8B, DeepSeek-V3-0324, DeepSeek-V4-Flash/Pro | 1.5B–1.6T | Reasoning (GRPO), code, MoE |
| **GLM** | GLM-5, GLM-5.1 | 9B–754B | Chinese + English, MoE |
| **Kimi** | Kimi-K2, Kimi-K2.5, Kimi-K2.6 | ~1T (MoE) | Long-context agentic, MoE |
| **MiniMax** | MiniMax-M2, MiniMax-M3 | 230B–428B | Agentic, MoE (community license) |
| **Phi-4** | Phi-4-14B, Phi-4-mini-reasoning | 3.8B–14B | Compact reasoning |
| **Mistral** | Mistral-7B-Instruct-v0.3, Mistral-Small-24B-Instruct | 7B–24B | Fast, efficient |
| **Mistral** | Mistral-7B-Instruct-v0.3, Mistral-Small-24B, Mistral-Large-3 | 7B–675B | Fast, efficient, MoE |
| **Mixtral** | Mixtral-8x7B-Instruct-v0.1, Mixtral-8x22B | 47B–141B | MoE architecture |
| **CodeLlama** | CodeLlama-7b-Instruct-hf, CodeLlama-34b-Instruct | 7B–34B | Code generation |
| **StarCoder 2** | StarCoder2-15B, StarCoder2-7B | 3B–15B | Code completion |

View File

@ -418,7 +418,7 @@ soup ui
**Pages:**
- **Dashboard** — view all experiment runs, loss charts, system info, multi-run comparison
- **New Training** — create configs from templates or 116 ready-made recipes, validate, start training with live SSE log streaming and progress bar
- **New Training** — create configs from templates or 133 ready-made recipes, validate, start training with live SSE log streaming and progress bar
- **Data Explorer** — browse and inspect datasets (JSONL, JSON, CSV, Parquet)
- **Model Chat** — chat with streaming responses, configurable temperature/top_p/max_tokens, system prompt, adapter selection, markdown rendering, chat export
@ -427,7 +427,7 @@ soup ui
- **Enhanced Metrics** — 2x2 chart grid (loss, LR, grad_norm, throughput) + GPU memory chart, eval results table
- **Multi-Run Compare** — overlay loss curves from up to 5 runs side-by-side
- **Chat Upgrade** — SSE streaming via proxy, typing indicator, cancel button, markdown renderer (bold, italic, code blocks), chat export as JSON
- **Config Builder** — recipe dropdown (116 recipes), config schema API for dynamic form generation
- **Config Builder** — recipe dropdown (133 recipes), config schema API for dynamic form generation
**Security:** The Web UI generates a random auth token at startup (printed to console). All mutating endpoints (start/stop training, delete runs, inspect data, validate config) require `Authorization: Bearer <token>` header. CORS is restricted to the served origin. Data inspection is sandboxed to the working directory.

View File

@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "soup-cli"
version = "0.71.23"
version = "0.71.24"
description = "Fine-tune and post-train LLMs in one command. No SSH, no config hell."
readme = "README.md"
license = "Apache-2.0"

View File

@ -1,3 +1,3 @@
"""Soup CLI — Fine-tune and post-train LLMs in one command."""
__version__ = "0.71.23"
__version__ = "0.71.24"

View File

@ -48,7 +48,7 @@ def search_recipes(
# ---------------------------------------------------------------------------
# Recipe catalog (116 recipes)
# Recipe catalog (133 recipes)
# ---------------------------------------------------------------------------
RECIPES: Dict[str, RecipeMeta] = {
@ -2594,13 +2594,13 @@ output: ./output
""",
),
"glm-5-sft": RecipeMeta(
model="THUDM/glm-5",
model="zai-org/GLM-5",
task="sft",
size="9B",
tags=("glm", "thudm", "chat", "next-gen"),
tags=("glm", "zai-org", "chat", "next-gen"),
description="GLM 5 SFT (next-gen GLM family)",
yaml_str="""\
base: THUDM/glm-5
base: zai-org/GLM-5
task: sft
data:
@ -3536,6 +3536,553 @@ training:
alpha: 32
target_modules: auto
output: ./output
""",
),
# ------------------------------------------------------------------
# v0.71.24 — 2026 model-family expansion (catalog 116 -> 133)
# 17 SFT recipes for the open-weight models released Feb-Jun 2026.
# Every base repo-ID verified to resolve on Hugging Face.
# ------------------------------------------------------------------
"qwen3.5-0.8b-sft": RecipeMeta(
model="Qwen/Qwen3.5-0.8B",
task="sft",
size="0.8B",
tags=("qwen", "qwen3.5", "sft", "tiny", "edge", "mobile"),
description="Qwen 3.5 0.8B SFT (Apache-2.0, tiny / mobile)",
yaml_str="""\
base: Qwen/Qwen3.5-0.8B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 2048
training:
epochs: 3
lr: 3e-4
batch_size: auto
lora:
r: 8
alpha: 16
target_modules: auto
quantization: 4bit
output: ./output
""",
),
"qwen3.5-2b-sft": RecipeMeta(
model="Qwen/Qwen3.5-2B",
task="sft",
size="2B",
tags=("qwen", "qwen3.5", "sft", "small", "edge"),
description="Qwen 3.5 2B SFT (Apache-2.0, small / edge)",
yaml_str="""\
base: Qwen/Qwen3.5-2B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 2048
training:
epochs: 3
lr: 2e-4
batch_size: auto
lora:
r: 8
alpha: 16
target_modules: auto
quantization: 4bit
output: ./output
""",
),
"qwen3.5-4b-sft": RecipeMeta(
model="Qwen/Qwen3.5-4B",
task="sft",
size="4B",
tags=("qwen", "qwen3.5", "sft", "small"),
description="Qwen 3.5 4B SFT (Apache-2.0, 262K context)",
yaml_str="""\
base: Qwen/Qwen3.5-4B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 3
lr: 2e-4
batch_size: auto
lora:
r: 16
alpha: 32
target_modules: auto
quantization: 4bit
output: ./output
""",
),
"qwen3.5-9b-sft": RecipeMeta(
model="Qwen/Qwen3.5-9B",
task="sft",
size="9B",
tags=("qwen", "qwen3.5", "sft", "chat", "instruction"),
description="Qwen 3.5 9B SFT (Apache-2.0, 262K context)",
yaml_str="""\
base: Qwen/Qwen3.5-9B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 3
lr: 2e-4
batch_size: auto
lora:
r: 16
alpha: 32
target_modules: auto
quantization: 4bit
output: ./output
""",
),
"qwen3.5-27b-sft": RecipeMeta(
model="Qwen/Qwen3.5-27B",
task="sft",
size="27B",
tags=("qwen", "qwen3.5", "sft", "large", "deepspeed"),
description="Qwen 3.5 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2",
yaml_str="""\
base: Qwen/Qwen3.5-27B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 3
lr: 1e-5
batch_size: auto
gradient_accumulation_steps: 8
lora:
r: 16
alpha: 32
target_modules: auto
quantization: 4bit
output: ./output
""",
),
"qwen3.5-35b-a3b-sft": RecipeMeta(
model="Qwen/Qwen3.5-35B-A3B",
task="sft",
size="35B",
tags=("qwen", "qwen3.5", "sft", "moe", "mixture-of-experts"),
description="Qwen 3.5 35B-A3B MoE SFT (Apache-2.0, 3B active)",
yaml_str="""\
base: Qwen/Qwen3.5-35B-A3B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 3
lr: 1e-4
batch_size: auto
gradient_accumulation_steps: 8
lora:
r: 16
alpha: 32
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
output: ./output
""",
),
"qwen3.5-122b-a10b-sft": RecipeMeta(
model="Qwen/Qwen3.5-122B-A10B",
task="sft",
size="122B",
tags=("qwen", "qwen3.5", "sft", "moe", "large", "multi-gpu"),
description=(
"Qwen 3.5 122B-A10B MoE SFT (Apache-2.0, 10B active). "
"Multi-GPU recommended (8 x A100/H100 80GB)."
),
yaml_str="""\
base: Qwen/Qwen3.5-122B-A10B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 1
lr: 1e-5
batch_size: 1
gradient_accumulation_steps: 16
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),
"qwen3.5-397b-a17b-sft": RecipeMeta(
model="Qwen/Qwen3.5-397B-A17B",
task="sft",
size="397B",
tags=("qwen", "qwen3.5", "sft", "moe", "large", "multi-gpu"),
description=(
"Qwen 3.5 397B-A17B flagship MoE SFT (Apache-2.0, 17B active). "
"Requires multi-node DeepSpeed / FSDP."
),
yaml_str="""\
base: Qwen/Qwen3.5-397B-A17B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 1
lr: 5e-6
batch_size: 1
gradient_accumulation_steps: 32
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),
"qwen3.6-27b-sft": RecipeMeta(
model="Qwen/Qwen3.6-27B",
task="sft",
size="27B",
tags=("qwen", "qwen3.6", "sft", "large", "deepspeed"),
description="Qwen 3.6 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2",
yaml_str="""\
base: Qwen/Qwen3.6-27B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 3
lr: 1e-5
batch_size: auto
gradient_accumulation_steps: 8
lora:
r: 16
alpha: 32
target_modules: auto
quantization: 4bit
output: ./output
""",
),
"qwen3.6-35b-a3b-sft": RecipeMeta(
model="Qwen/Qwen3.6-35B-A3B",
task="sft",
size="35B",
tags=("qwen", "qwen3.6", "sft", "moe", "mixture-of-experts"),
description="Qwen 3.6 35B-A3B MoE SFT (Apache-2.0, 3B active)",
yaml_str="""\
base: Qwen/Qwen3.6-35B-A3B
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 3
lr: 1e-4
batch_size: auto
gradient_accumulation_steps: 8
lora:
r: 16
alpha: 32
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
output: ./output
""",
),
"deepseek-v4-flash-sft": RecipeMeta(
model="deepseek-ai/DeepSeek-V4-Flash",
task="sft",
size="N/A",
tags=("deepseek", "deepseek-v4", "sft", "moe", "mixture-of-experts"),
description="DeepSeek V4 Flash MoE SFT (MIT, efficiency-tier)",
yaml_str="""\
base: deepseek-ai/DeepSeek-V4-Flash
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 3
lr: 1e-4
batch_size: auto
gradient_accumulation_steps: 8
lora:
r: 16
alpha: 32
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
output: ./output
""",
),
"deepseek-v4-pro-sft": RecipeMeta(
model="deepseek-ai/DeepSeek-V4-Pro",
task="sft",
size="N/A",
tags=("deepseek", "deepseek-v4", "sft", "moe", "large", "multi-gpu"),
description=(
"DeepSeek V4 Pro flagship MoE SFT (MIT, 1.6T-class). "
"Requires multi-node DeepSpeed."
),
yaml_str="""\
base: deepseek-ai/DeepSeek-V4-Pro
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 1
lr: 5e-6
batch_size: 1
gradient_accumulation_steps: 32
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),
"glm-5.1-sft": RecipeMeta(
model="zai-org/GLM-5.1",
task="sft",
size="754B",
tags=("glm", "zai-org", "sft", "moe", "large", "multi-gpu"),
description=(
"GLM 5.1 MoE SFT (MIT, 754B). Multi-GPU / multi-node recommended."
),
yaml_str="""\
base: zai-org/GLM-5.1
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 8192
training:
epochs: 1
lr: 1e-5
batch_size: 1
gradient_accumulation_steps: 16
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),
"kimi-k2.5-sft": RecipeMeta(
model="moonshotai/Kimi-K2.5",
task="sft",
size="1T",
tags=("kimi", "moonshot", "sft", "moe", "large", "multi-gpu"),
description=(
"Kimi K2.5 MoE SFT (Modified MIT, ~1T / 32B active). "
"Requires multi-node DeepSpeed."
),
yaml_str="""\
base: moonshotai/Kimi-K2.5
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 8192
training:
epochs: 1
lr: 1e-5
batch_size: 1
gradient_accumulation_steps: 16
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),
"kimi-k2.6-sft": RecipeMeta(
model="moonshotai/Kimi-K2.6",
task="sft",
size="1T",
tags=("kimi", "moonshot", "sft", "moe", "large", "multi-gpu"),
description=(
"Kimi K2.6 MoE SFT (Modified MIT, ~1T / 32B active). "
"Requires multi-node DeepSpeed."
),
yaml_str="""\
base: moonshotai/Kimi-K2.6
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 8192
training:
epochs: 1
lr: 5e-6
batch_size: 1
gradient_accumulation_steps: 32
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),
"minimax-m3-sft": RecipeMeta(
model="MiniMaxAI/MiniMax-M3",
task="sft",
size="428B",
tags=("minimax", "sft", "moe", "large", "multi-gpu"),
description=(
"MiniMax M3 MoE SFT (428B / 23B active). MiniMax Community License "
"- commercial use requires a separate agreement. Multi-GPU recommended."
),
yaml_str="""\
base: MiniMaxAI/MiniMax-M3
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 1
lr: 1e-5
batch_size: 1
gradient_accumulation_steps: 16
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),
"mistral-large-3-sft": RecipeMeta(
model="mistralai/Mistral-Large-3-675B-Instruct-2512",
task="sft",
size="675B",
tags=("mistral", "mistral-large", "sft", "moe", "large", "multi-gpu"),
description=(
"Mistral Large 3 MoE SFT (Apache-2.0, 675B / 41B active, multimodal). "
"Requires multi-node DeepSpeed."
),
yaml_str="""\
base: mistralai/Mistral-Large-3-675B-Instruct-2512
task: sft
data:
train: ./data/train.jsonl
format: auto
max_length: 4096
training:
epochs: 1
lr: 5e-6
batch_size: 1
gradient_accumulation_steps: 32
lora:
r: 32
alpha: 64
target_modules: auto
quantization: 4bit
moe_lora: true
moe_aux_loss_coeff: 0.01
gradient_checkpointing: true
output: ./output
""",
),

View File

@ -260,7 +260,7 @@ class TestV025NewRecipes:
assert cfg.base == recipe.model
assert cfg.task == recipe.task
def test_catalog_size_is_116(self):
def test_catalog_size_is_133(self):
"""Total catalog size — grew with each release.
v0.25.0 shipped 43 recipes (29 + 9 Part A + 2 Part B tools + 3 Part E MLX).
@ -270,10 +270,11 @@ class TestV025NewRecipes:
v0.52.0 added 6 (5 TTS + Falcon-E BitNet) -> 112.
v0.53.5 added 1 (deepseek-v3-reasoning) -> 113.
v0.62.0 added 3 (raft-llama3-8b, ra-dit-retriever, ra-dit-llama3-8b) -> 116.
v0.71.24 added 17 (2026 model-family expansion) -> 133.
"""
from soup_cli.recipes.catalog import RECIPES
assert len(RECIPES) == 116
assert len(RECIPES) == 133
def test_new_recipes_searchable(self):
"""Search returns the new recipes via keyword/task filter."""

291
tests/test_v07124.py Normal file
View File

@ -0,0 +1,291 @@
"""v0.71.24 — 2026 model-family recipe expansion (catalog 116 -> 133).
Pure config/data release: 17 new ready-made SFT recipes for the open-weight
models released Feb-Jun 2026 (Qwen3.5 / Qwen3.6 / DeepSeek-V4 / GLM-5.1 /
Kimi-K2.5/K2.6 / MiniMax-M3 / Mistral-Large-3), plus a stale-repo-ID fix for
``glm-5-sft`` (``THUDM/glm-5`` -> ``zai-org/GLM-5``). CI validates each recipe by
``load_config_from_string`` parse only (no network).
"""
from __future__ import annotations
import pytest
import yaml
from soup_cli.config.loader import load_config_from_string
from soup_cli.config.schema import SoupConfig
from soup_cli.recipes.catalog import RECIPES, get_recipe, list_recipes, search_recipes
# ---------------------------------------------------------------------------
# Step A — 17 new SFT recipes (every base verified to resolve on Hugging Face)
# ---------------------------------------------------------------------------
V07124_RECIPE_NAMES = [
# Qwen3.5 family (Apache-2.0)
"qwen3.5-0.8b-sft",
"qwen3.5-2b-sft",
"qwen3.5-4b-sft",
"qwen3.5-9b-sft",
"qwen3.5-27b-sft",
"qwen3.5-35b-a3b-sft",
"qwen3.5-122b-a10b-sft",
"qwen3.5-397b-a17b-sft",
# Qwen3.6 (Apache-2.0)
"qwen3.6-27b-sft",
"qwen3.6-35b-a3b-sft",
# DeepSeek-V4 (MIT)
"deepseek-v4-flash-sft",
"deepseek-v4-pro-sft",
# GLM (MIT, zai-org)
"glm-5.1-sft",
# Kimi (Modified MIT)
"kimi-k2.5-sft",
"kimi-k2.6-sft",
# MiniMax (MiniMax Community License)
"minimax-m3-sft",
# Mistral Large 3 (Apache-2.0)
"mistral-large-3-sft",
]
# name -> expected base repo-ID (the verified Hugging Face repo)
V07124_RECIPE_BASES = {
"qwen3.5-0.8b-sft": "Qwen/Qwen3.5-0.8B",
"qwen3.5-2b-sft": "Qwen/Qwen3.5-2B",
"qwen3.5-4b-sft": "Qwen/Qwen3.5-4B",
"qwen3.5-9b-sft": "Qwen/Qwen3.5-9B",
"qwen3.5-27b-sft": "Qwen/Qwen3.5-27B",
"qwen3.5-35b-a3b-sft": "Qwen/Qwen3.5-35B-A3B",
"qwen3.5-122b-a10b-sft": "Qwen/Qwen3.5-122B-A10B",
"qwen3.5-397b-a17b-sft": "Qwen/Qwen3.5-397B-A17B",
"qwen3.6-27b-sft": "Qwen/Qwen3.6-27B",
"qwen3.6-35b-a3b-sft": "Qwen/Qwen3.6-35B-A3B",
"deepseek-v4-flash-sft": "deepseek-ai/DeepSeek-V4-Flash",
"deepseek-v4-pro-sft": "deepseek-ai/DeepSeek-V4-Pro",
"glm-5.1-sft": "zai-org/GLM-5.1",
"kimi-k2.5-sft": "moonshotai/Kimi-K2.5",
"kimi-k2.6-sft": "moonshotai/Kimi-K2.6",
"minimax-m3-sft": "MiniMaxAI/MiniMax-M3",
"mistral-large-3-sft": "mistralai/Mistral-Large-3-675B-Instruct-2512",
}
# The MoE bases added in v0.71.24 — each recipe must enable moe_lora + aux loss.
# (Scope is intentionally limited to recipes added in THIS release, not every MoE
# recipe in the catalog.)
V07124_MOE_RECIPES = [
"qwen3.5-35b-a3b-sft",
"qwen3.5-122b-a10b-sft",
"qwen3.5-397b-a17b-sft",
"qwen3.6-35b-a3b-sft",
"deepseek-v4-flash-sft",
"deepseek-v4-pro-sft",
"glm-5.1-sft",
"kimi-k2.5-sft",
"kimi-k2.6-sft",
"minimax-m3-sft",
"mistral-large-3-sft",
]
class TestV07124Recipes:
def test_recipe_count_target(self) -> None:
assert len(V07124_RECIPE_NAMES) == 17
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_registered(self, name: str) -> None:
assert name in RECIPES, f"Recipe {name!r} missing from catalog"
assert get_recipe(name) is not None
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_base_matches_verified_repo(self, name: str) -> None:
assert RECIPES[name].model == V07124_RECIPE_BASES[name]
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_is_sft(self, name: str) -> None:
assert RECIPES[name].task == "sft"
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_metadata_well_formed(self, name: str) -> None:
meta = RECIPES[name]
assert meta.model, f"{name}: model is empty"
assert meta.size, f"{name}: size is empty"
assert meta.tags, f"{name}: tags is empty"
assert meta.description, f"{name}: description is empty"
assert meta.yaml_str, f"{name}: yaml_str is empty"
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_yaml_parses_as_soup_config(self, name: str) -> None:
cfg = load_config_from_string(RECIPES[name].yaml_str)
assert isinstance(cfg, SoupConfig)
assert cfg.base == RECIPES[name].model
assert cfg.task == "sft"
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_yaml_is_safe_loadable(self, name: str) -> None:
parsed = yaml.safe_load(RECIPES[name].yaml_str)
assert isinstance(parsed, dict)
assert "base" in parsed and "task" in parsed
# The inline ``base:`` / ``task:`` in the YAML must match the RecipeMeta.
assert parsed["base"] == RECIPES[name].model
assert parsed["task"] == RECIPES[name].task
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_model_id_no_null_or_whitespace(self, name: str) -> None:
meta = RECIPES[name]
assert "\x00" not in meta.model
assert " " not in meta.model
# No path-traversal segments or shell metacharacters in a repo id.
assert ".." not in meta.model, f"path-traversal segment in {meta.model!r}"
assert not any(c in meta.model for c in ';|&$`<>(){}*?!\\\n\t'), (
f"shell metacharacter in {meta.model!r}"
)
parts = meta.model.split("/")
assert len(parts) == 2
assert all(p for p in parts), f"empty component in {meta.model!r}"
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_max_length_within_bounds(self, name: str) -> None:
cfg = load_config_from_string(RECIPES[name].yaml_str)
assert 64 <= cfg.data.max_length <= 1_048_576
@pytest.mark.parametrize("name", V07124_MOE_RECIPES)
def test_moe_recipes_enable_moe_lora(self, name: str) -> None:
cfg = load_config_from_string(RECIPES[name].yaml_str)
assert cfg.training.moe_lora is True, f"{name} should set moe_lora: true"
@pytest.mark.parametrize("name", V07124_MOE_RECIPES)
def test_moe_recipes_set_aux_loss_coeff(self, name: str) -> None:
# Every new MoE recipe sets the aux-loss coefficient explicitly (0.01) so
# the whole batch is uniformly self-documenting (review HIGH-1).
cfg = load_config_from_string(RECIPES[name].yaml_str)
assert abs(cfg.training.moe_aux_loss_coeff - 0.01) < 1e-9, (
f"{name} should set moe_aux_loss_coeff: 0.01, got "
f"{cfg.training.moe_aux_loss_coeff}"
)
@pytest.mark.parametrize("name", sorted(set(V07124_RECIPE_NAMES) - set(V07124_MOE_RECIPES)))
def test_dense_recipes_do_not_set_moe_lora(self, name: str) -> None:
# The complement of the MoE list: a dense recipe must NOT enable moe_lora
# (guards against pasting a MoE block into a dense recipe).
cfg = load_config_from_string(RECIPES[name].yaml_str)
assert cfg.training.moe_lora is not True, (
f"{name} is a dense recipe but has moe_lora: true"
)
@pytest.mark.parametrize("name", V07124_MOE_RECIPES)
def test_moe_recipes_tagged_moe(self, name: str) -> None:
assert "moe" in RECIPES[name].tags, (
f"{name} is a MoE recipe but lacks the 'moe' tag"
)
@pytest.mark.parametrize("name", [n for n in V07124_RECIPE_NAMES if n.startswith("qwen")])
def test_qwen_recipes_tagged_qwen(self, name: str) -> None:
assert "qwen" in RECIPES[name].tags, f"{name} lacks the 'qwen' tag"
def test_search_by_task_sft_includes_all_new_recipes(self) -> None:
sft_models = {r.model for r in search_recipes(task="sft")}
for name in V07124_RECIPE_NAMES:
assert RECIPES[name].model in sft_models, (
f"{name} not returned by search_recipes(task='sft')"
)
def test_kimi_k2_5_size_known(self) -> None:
# K2.5 is a ~1T/32B MoE (verified on the model card) — not "N/A".
assert RECIPES["kimi-k2.5-sft"].size == "1T"
def test_kimi_k2_6_size_known(self) -> None:
# K2.6 is a ~1T/32B MoE (verified on the model card).
assert RECIPES["kimi-k2.6-sft"].size == "1T"
def test_minimax_recipe_notes_community_license(self) -> None:
desc = RECIPES["minimax-m3-sft"].description.lower()
assert "minimax" in desc
assert "license" in desc
# commercial-use caveat must be surfaced
assert "commercial" in desc
def test_mistral_large_3_notes_apache_license(self) -> None:
# Plan guessed MRL; the published card is Apache-2.0 — describe accurately.
desc = RECIPES["mistral-large-3-sft"].description.lower()
assert "apache" in desc
@pytest.mark.parametrize("name", ["kimi-k2.5-sft", "kimi-k2.6-sft"])
def test_kimi_recipes_note_modified_mit(self, name: str) -> None:
assert "mit" in RECIPES[name].description.lower()
def test_deepseek_v4_sizes_intentionally_na(self) -> None:
# DeepSeek has not publicly confirmed the V4 parameter counts; "N/A" is the
# honest value. This documents the deliberate choice (vs Kimi K2.5/K2.6,
# whose ~1T sizes ARE on the card).
assert RECIPES["deepseek-v4-flash-sft"].size == "N/A"
assert RECIPES["deepseek-v4-pro-sft"].size == "N/A"
@pytest.mark.parametrize("name", V07124_RECIPE_NAMES)
def test_recipe_searchable(self, name: str) -> None:
# The recipe key is part of the searchable text, so a query for the recipe's
# own name must return that exact recipe (deterministic — no token heuristic).
results = search_recipes(query=name)
assert any(r.model == RECIPES[name].model for r in results), (
f"{name} not returned by search_recipes(query={name!r})"
)
# ---------------------------------------------------------------------------
# Step B — stale repo-ID fix: glm-5-sft (THUDM/glm-5 -> zai-org/GLM-5)
# ---------------------------------------------------------------------------
class TestGlm5RepoIdFix:
def test_glm5_uses_zai_org(self) -> None:
meta = RECIPES["glm-5-sft"]
assert meta.model == "zai-org/GLM-5"
assert "THUDM" not in meta.model
def test_glm5_yaml_base_updated(self) -> None:
meta = RECIPES["glm-5-sft"]
parsed = yaml.safe_load(meta.yaml_str)
assert parsed["base"] == "zai-org/GLM-5"
assert "THUDM/glm-5" not in meta.yaml_str
def test_glm5_old_key_still_exists(self) -> None:
# The fix RENAMES the base field; it must NOT delete the old recipe key.
assert "glm-5-sft" in RECIPES, "glm-5-sft was accidentally removed"
assert "glm-5.1-sft" in RECIPES, "glm-5.1-sft is the new v0.71.24 recipe"
def test_glm5_and_glm51_are_distinct_entries(self) -> None:
assert RECIPES["glm-5-sft"].model != RECIPES["glm-5.1-sft"].model
def test_no_recipe_references_thudm_glm5(self) -> None:
for name, meta in RECIPES.items():
assert "THUDM/glm-5" not in meta.model, (
f"{name}: model still references THUDM/glm-5"
)
assert "THUDM/glm-5" not in meta.yaml_str, (
f"{name}: yaml_str still references THUDM/glm-5"
)
# ---------------------------------------------------------------------------
# Catalog count — 116 -> 133
# ---------------------------------------------------------------------------
class TestCatalogCount:
def test_total_recipe_count_is_133(self) -> None:
assert len(RECIPES) == 133
def test_list_recipes_matches_dict(self) -> None:
assert len(list_recipes()) == len(RECIPES)
def test_all_recipes_parse(self) -> None:
# Defence-in-depth — the whole catalog must still parse after the additions.
for name, meta in RECIPES.items():
cfg = load_config_from_string(meta.yaml_str)
assert cfg.base, name
def test_all_recipes_yaml_base_matches_model(self) -> None:
# Catalog-wide integrity invariant — the inline ``base:`` must equal
# RecipeMeta.model for EVERY recipe (catches a half-applied repo-ID fix).
for name, meta in RECIPES.items():
parsed = yaml.safe_load(meta.yaml_str)
assert parsed.get("base") == meta.model, (
f"{name}: yaml base {parsed.get('base')!r} != model {meta.model!r}"
)