docs(v0.71.31): judge-in-the-loop suite — version bump, CHANGELOG, docs, recipe, counts

Bump 0.71.30 -> 0.71.31; +1 recipe online-dpo-smollm2-135m (137->138); CHANGELOG
[0.71.31]; README What's New; docs (training/data/evaluation/commands/serving);
recipe-count test asserts 137->138.
This commit is contained in:
Alpamys 2026-07-06 13:47:14 +05:00
parent c0a9ba5435
commit 8f59b9b9d7
15 changed files with 146 additions and 33 deletions

View File

@ -12,6 +12,35 @@ reproducing 70+ versions of notes.
## [Unreleased]
## [0.71.31] - 2026-07-06
### Added
- **Judge-in-the-loop suite** — put an LLM judge in the loop across the workflow:
- **`task='online_dpo'`** — Online DPO training (wraps TRL `OnlineDPOTrainer`):
the model generates two completions per prompt on-policy each step and a
*judge* (a pairwise LLM judge over the existing ollama/openai-compatible
backend) OR a `reward_model` picks the winner. Config:
`training.online_dpo_judge: "ollama://model"` (or set `reward_model`
exactly one), `online_dpo_loss_type: sigmoid|ipo`, `online_dpo_max_new_tokens`;
`beta` reuses `dpo_beta`. Transformers + text only. Recipe:
`online-dpo-smollm2-135m`.
- **`soup data best-of-n`** — Best-of-N rejection sampling (BOND-lite): sample
N completions from `--base` locally, a `--judge` scores each pointwise, and
the winner is written as an SFT chat row (with provenance). `--emit-pairs`
also writes winner-vs-loser DPO pairs.
- **`soup data evolve`** — Evol-Instruct instruction evolution (WizardLM depth
/ breadth) over an ollama/vllm provider, completing the synthetic-data suite
(Magpie / Forge / Persona / evolve).
- **`soup ship --task-mode pairwise`** — a true pairwise judge win-rate as the
ship leg-1 task-win (base = 0.5 coin-flip, tuned = its win-rate; swap-debiased),
fusing with the catastrophic-forgetting guard into one SHIP / DON'T-SHIP verdict.
### Security
- `soup data best-of-n` / `evolve` write outputs via atomic `mkstemp` + `os.replace`
(re-validated cwd containment), closing the TOCTOU symlink-swap window between the
containment check and the write. All judge/provider URLs are SSRF-validated; model
loads probe `trust_remote_code`.
## [0.71.30] - 2026-07-05
### Added

View File

@ -110,17 +110,17 @@ src/soup_cli/
experiment/ - SQLite experiment tracking
eval/ - Eval platform (custom tasks, LLM judge, human eval, leaderboard)
migrate/ - Config migration (LLaMA-Factory, Axolotl, Unsloth)
recipes/ - Ready-made configs for popular models (137 recipes)
recipes/ - Ready-made configs for popular models (138 recipes)
autopilot/ - Zero-config decision engine (v0.25.0)
registry/ - Model Registry (hashing, store, diff, attach) (v0.26.0 + v0.33.0)
cans/ - Shareable .can artifact format + run/publish orchestrator (v0.26.0 + v0.33.0)
data/traces/ - Trace-to-Preference harvester (v0.26.0)
data/collators.py - CrossDocCollator for sample packing (v0.33.0)
utils/ - GPU, errors, MoE, GaLore, QAT, Unsloth, vLLM, SGLang, Liger, FlashAttn, FSDP, Ring Attention, long-context, quality, curriculum, freeze, dataset-registry, mlx, peft_builder, paths, topology, launcher, mii, pipeline, cut_ce, fp8, gradient_ckpt, kernel_picker, cross_doc_attn, activation_offload, hf, spec_pairing, structured_output, metrics, tracing, auto_quant, lr_finder, grad_accum, mixed_precision, warmup, spike_recovery, convergence, v028_features, multipack_sampler, multipack, neat_packing, jinja_analyzer, quant_menu, relora, peft_patches, peft_wiring, dpo_variants, optimizer_zoo, lr_groups, loftq_init, block_expansion, tts, classifier, distill, bitnet, ebft_gdpo, moe_quant, reasoning_effort, gguf_quant, kv_cache, advanced_precision, save_formats, deploy_measure, advise, advise_history, adapter_diff, adapter_merge, blame, adapter_branch, unlearning, unlearning_eval, knowledge_edit, edit_governor, edit_diff, ra_dit, steering, citation_faithful, grace_codebook, ingest_sources, prune_prompt, active_sampler, ab_test, drift_alarm, tunability, terraform_plan, env_lock, hardware_fit, completions, license_advisor, behavior_battery, capability_suite, checklist_dsl, irt, sae_diff, sleeper_probe, interference, probe_pack, cmaes_merge, vector_bank, mole_routing, adapter_pr, soup_lock, adapter_bisect, prompt_compile, prompt_distill, compile_tools, apple_adapter, local_rl, build_dag, expectations, magpie, persona_hub, brain_rot, reward_hacking, uld, minillm, rl_checkpoint, iterative_dpo, echo_trap, mod, local_rl_scheduler, spectrum_scan, ship_verdict, reward_hack_control, data_doctor, data_lint, shrink
utils/ - GPU, errors, MoE, GaLore, QAT, Unsloth, vLLM, SGLang, Liger, FlashAttn, FSDP, Ring Attention, long-context, quality, curriculum, freeze, dataset-registry, mlx, peft_builder, paths, topology, launcher, mii, pipeline, cut_ce, fp8, gradient_ckpt, kernel_picker, cross_doc_attn, activation_offload, hf, spec_pairing, structured_output, metrics, tracing, auto_quant, lr_finder, grad_accum, mixed_precision, warmup, spike_recovery, convergence, v028_features, multipack_sampler, multipack, neat_packing, jinja_analyzer, quant_menu, relora, peft_patches, peft_wiring, dpo_variants, optimizer_zoo, lr_groups, loftq_init, block_expansion, tts, classifier, distill, bitnet, ebft_gdpo, moe_quant, reasoning_effort, gguf_quant, kv_cache, advanced_precision, save_formats, deploy_measure, advise, advise_history, adapter_diff, adapter_merge, blame, adapter_branch, unlearning, unlearning_eval, knowledge_edit, edit_governor, edit_diff, ra_dit, steering, citation_faithful, grace_codebook, ingest_sources, prune_prompt, active_sampler, ab_test, drift_alarm, tunability, terraform_plan, env_lock, hardware_fit, completions, license_advisor, behavior_battery, capability_suite, checklist_dsl, irt, sae_diff, sleeper_probe, interference, probe_pack, cmaes_merge, vector_bank, mole_routing, adapter_pr, soup_lock, adapter_bisect, prompt_compile, prompt_distill, compile_tools, apple_adapter, local_rl, build_dag, expectations, magpie, persona_hub, brain_rot, reward_hacking, uld, minillm, rl_checkpoint, iterative_dpo, echo_trap, mod, local_rl_scheduler, spectrum_scan, ship_verdict, reward_hack_control, data_doctor, data_lint, shrink, best_of_n, evolve
templates/ - 17 built-in soup.yaml templates (YAML + manifest.json) with load_template loader (v0.39.0, +bco v0.40.0)
ui/ - Web UI (FastAPI + HTML/JS SPA)
tests/ - Test suite (307 files, 15334 tests)
tests/ - Test suite (308 files, 15421 tests)
examples/ - Real-world config examples and datasets
```

View File

@ -49,27 +49,29 @@ infrastructure instead of improving models. Soup fixes that.
## What's New
**v0.71.30 — Process-supervised RL: PRM-guided GRPO.** Use a trained Process Reward Model to score each *reasoning step* of a GRPO completion — the o1-era training signal — plus bundled toy environments so the openenv rollout path runs out-of-the-box.
**v0.71.31 — Judge-in-the-loop suite.** Put an LLM judge in the loop across the whole workflow: train against a judge, mine winners from a base model, grow instruction diversity, and decide SHIP with a true pairwise judge win-rate.
- **PRM as the GRPO reward.** Point `training.prm_reward` at a PRM you trained with
`soup train task=prm`; it splits each completion into steps, scores every step with the PRM's
reward head, and folds them (`min` / `prod` / `last`) into one reward that GRPO optimises —
replacing `reward_fn`. It rides the reward-shaping + reward-hack-mitigation seam, so the
v0.71.26 controller still watches it.
- **Bundled rollout environments.** `soup_cli.envs.calculator` / `retrieval_qa` / `guess_number`
each expose a `rollout(prompts)` entry point; wire one with `rollout_backend=openenv` +
`rollout_func=soup_cli.envs.calculator:rollout` (three ready-made `grpo-env-*` recipes ship it).
- **Producer fixes too.** `soup train task=prm` now saves its tokenizer and no longer crashes on
its own summary — a PRM checkpoint is loadable standalone.
- **Proof-of-mechanism, honestly.** Live-validated on SmolLM2-135M (CPU): the PRM reward scores
good completions above bad and drives GRPO's advantages (`rewards/prm_reward` logged). Tiny
model + synthetic PRM — not a production reward-model claim (scale help wanted: #286).
- **`task='online_dpo'`.** Online DPO (wraps TRL `OnlineDPOTrainer`): the model generates two
completions per prompt on-policy each step and a *judge* (a pairwise LLM judge over your local
ollama / OpenAI-compatible endpoint) — or a `reward_model` — picks the winner. Set
`training.online_dpo_judge: "ollama://llama3.1"` (or `reward_model`, exactly one).
- **`soup data best-of-n`.** Best-of-N rejection sampling: sample N completions from `--base`
locally, a `--judge` scores each, and the winner becomes an SFT row (`--emit-pairs` also writes
winner-vs-loser DPO pairs).
- **`soup data evolve`.** WizardLM Evol-Instruct (depth / breadth) over an ollama / vllm provider —
completing the synthetic-data suite (Magpie / Forge / Persona / evolve).
- **`soup ship --task-mode pairwise` (#284).** A true swap-debiased judge win-rate as the ship
leg-1 task-win, fused with the catastrophic-forgetting guard into one SHIP / DON'T-SHIP verdict.
- **Proof-of-mechanism, honestly.** Live-validated on SmolLM2-135M: online-DPO trains with a
synthetic judge (`rewards/*` + `objective/kl` logged), best-of-N samples + picks, evolve loops.
Not a production RLHF claim (scale help wanted: #286).
```yaml
task: grpo
task: online_dpo
training:
prm_reward: ./my-prm # a `soup train task=prm` checkpoint dir
prm_aggregate: min # weakest-link (default) | prod | last
online_dpo_judge: "ollama://llama3.1" # a pairwise judge (or set reward_model)
online_dpo_loss_type: sigmoid # sigmoid | ipo
online_dpo_max_new_tokens: 64
```
Full history: [CHANGELOG.md](CHANGELOG.md) · [GitHub Releases](https://github.com/MakazhanAlpamys/Soup/releases).

View File

@ -133,7 +133,7 @@ soup migrate --from llamafactory config.yaml Import config from LLaMA-Factory
soup migrate --from axolotl config.yml Import config from Axolotl
soup migrate --from unsloth notebook.ipynb Import config from Unsloth notebook
soup migrate --from llamafactory c.yaml --dry-run Preview without writing
soup recipes list List all 137 ready-made recipes
soup recipes list List all 138 ready-made recipes
soup recipes show llama3.1-8b-sft Print recipe YAML
soup recipes use llama3.1-8b-sft Copy recipe to soup.yaml
soup recipes search "reasoning" Search by keyword/task/size
@ -155,6 +155,7 @@ soup why [run_id] Explain training anomalies (heuris
soup ship --base <m> --adapter <lora> --task-eval t.jsonl SHIP / DON'T-SHIP verdict: task win AND no forgetting (exit 0=SHIP / 2=DON'T / 1=error) (v0.71.25)
soup ship --evidence ev.json [--output v.json] Decide offline from pre-computed scores (no model load)
soup ship ... --task-mode judge_score --judge-model ollama://llama3.1 Leg-1 via LLM-as-a-judge
soup ship ... --task-mode pairwise --judge-model ollama://llama3.1 Leg-1 via swap-debiased judge win-rate (base=0.5) (v0.71.31)
soup ship ... --general-suite mmlu,gsm8k --baseline base.json lm-eval leg-2 + recorded base scores
soup mcp serve MCP server over stdio (drive Soup from Claude Code / Cursor / Cline; requires [mcp] extra) (v0.71.28)
soup mcp serve --allow-mutating Also expose plan-only train_start / export tools (never execute) (v0.71.28)
@ -250,6 +251,8 @@ soup local-rl train --db <path> --model <id> [--scheduler-dir <dir>] [--hour H]
soup build <manifest.yaml> [--dry-run] [--output-dir <dir>] dbt-for-SFT DAG: validate + plan + live materialise (v0.69.0; live v0.71.6)
soup expect <data.jsonl> <suite.yaml> Expectations suite: PII / token-length / refusal / judge (v0.69.0)
soup data gen-magpie --base <m> --provider ollama|vllm --target N --output <jsonl> [--base-url <url>] [--quality-filter] Magpie synthetic generator — live (v0.69.0; live v0.71.6)
soup data best-of-n --base <m> --prompts <jsonl> --n 8 --judge <url> -o <sft.jsonl> [--emit-pairs <dpo.jsonl>] Best-of-N rejection sampling: sample N locally, judge picks winner -> SFT (+ DPO) rows (v0.71.31)
soup data evolve --input <seeds.jsonl> --provider ollama|vllm --model <m> --strategy depth|breadth --rounds N -o <jsonl> Evol-Instruct (WizardLM) instruction evolution (v0.71.31)
soup data persona-mix --prompts <jsonl> --n N --output <jsonl> Persona-Hub diversity sampler (v0.69.0)
soup data brain-rot <data.jsonl> [--strict] Brain-rot detector — arXiv 2510.13928 (v0.69.0)
soup iterative-dpo --base-model <m> --reward-model <rm> --prompts <p.jsonl> --output-dir <out> --rounds N --pairs-per-round N [--plan-only] Iterative DPO loop driver — LIVE sample→score→pair→train (v0.70.0; live v0.71.11)

View File

@ -62,6 +62,15 @@ soup data persona-mix --prompts prompts.jsonl --n 500 --output mixed.jsonl
# Brain-rot detector (arXiv 2510.13928) — refuses to train on excessive slop
soup data brain-rot data.jsonl --strict --max-major-fraction 0.10
# Best-of-N rejection sampling (v0.71.31) — sample N locally, a judge picks the winner
soup data best-of-n --base HuggingFaceTB/SmolLM2-135M-Instruct \
--prompts prompts.jsonl --n 8 --judge ollama://llama3.1 \
-o best_of_n.jsonl --emit-pairs pairs.jsonl
# Evol-Instruct (WizardLM depth/breadth, v0.71.31) — grow instruction diversity
soup data evolve --input seeds.jsonl --provider ollama --model llama3.1 \
--strategy depth --rounds 2 -o evolved.jsonl
```
Every command applies the project-wide TOCTOU policy (`os.lstat + S_ISLNK` symlink rejection before any open) and cwd containment via the shared `paths.enforce_under_cwd_and_no_symlink` helper. All five are LIVE: `soup build` materialises with five built-in transforms (`identity` / `drop_empty` / `lowercase` / `strip` / `dedup_exact`) and SQLite-tracked incremental re-transform (v0.71.6); `soup data gen-magpie` harvests via raw completion against `--provider ollama|vllm` (loopback-only; `anthropic` rejected — no raw-completion endpoint, v0.71.6).

View File

@ -330,6 +330,11 @@ soup ship --base HuggingFaceTB/SmolLM2-135M-Instruct --adapter ./out \
soup ship --base <m> --adapter ./out --task-eval tasks.jsonl \
--task-mode judge_score --judge-model ollama://llama3.1
# Leg-1 via a true pairwise judge win-rate (v0.71.31 #284) — the judge picks
# base vs tuned per prompt (swap-debiased); base = 0.5 coin-flip, won <=> winrate > 0.5
soup ship --base <m> --adapter ./out --task-eval tasks.jsonl \
--task-mode pairwise --judge-model ollama://llama3.1
# Leg-2 via lm-eval benchmarks, base scores supplied by --baseline
soup ship --base <m> --tuned ./out --task-eval tasks.jsonl \
--general-suite mmlu,hellaswag --baseline registry://abc123

View File

@ -418,7 +418,7 @@ soup ui
**Pages:**
- **Dashboard** — view all experiment runs, loss charts, system info, multi-run comparison
- **New Training** — create configs from templates or 137 ready-made recipes, validate, start training with live SSE log streaming and progress bar
- **New Training** — create configs from templates or 138 ready-made recipes, validate, start training with live SSE log streaming and progress bar
- **Data Explorer** — browse and inspect datasets (JSONL, JSON, CSV, Parquet)
- **Model Chat** — chat with streaming responses, configurable temperature/top_p/max_tokens, system prompt, adapter selection, markdown rendering, chat export
@ -427,7 +427,7 @@ soup ui
- **Enhanced Metrics** — 2x2 chart grid (loss, LR, grad_norm, throughput) + GPU memory chart, eval results table
- **Multi-Run Compare** — overlay loss curves from up to 5 runs side-by-side
- **Chat Upgrade** — SSE streaming via proxy, typing indicator, cancel button, markdown renderer (bold, italic, code blocks), chat export as JSON
- **Config Builder** — recipe dropdown (137 recipes), config schema API for dynamic form generation
- **Config Builder** — recipe dropdown (138 recipes), config schema API for dynamic form generation
**Security:** The Web UI generates a random auth token at startup (printed to console). All mutating endpoints (start/stop training, delete runs, inspect data, validate config) require `Authorization: Bearer <token>` header. CORS is restricted to the served origin. Data inspection is sandboxed to the working directory.

View File

@ -401,6 +401,34 @@ advantages). Not a production reward-model claim; scale validation is help-wante
(#286).
## Online DPO (`task='online_dpo'`) — judge in the loop (v0.71.31)
Unlike offline DPO (static `prompt/chosen/rejected` rows), **Online DPO** generates
two completions per prompt **on-policy** each step and asks a *judge* — or a
*reward model* — which is better; the winner becomes `chosen`, the loser
`rejected`. The judge closes the loop. Wraps TRL `OnlineDPOTrainer`; data is
prompt-only (like GRPO). Transformers + text only.
```yaml
base: HuggingFaceTB/SmolLM2-135M-Instruct
task: online_dpo
data:
train: ./data/prompts.jsonl # prompt-only (or any format — prompts are extracted)
training:
online_dpo_judge: "ollama://llama3.1" # a pairwise judge (ollama://|https://|http://localhost)
# OR: reward_model: ./my-reward-model # exactly one of judge / reward_model
online_dpo_loss_type: sigmoid # sigmoid | ipo
online_dpo_max_new_tokens: 64
dpo_beta: 0.1
lora: { r: 8, alpha: 16, target_modules: auto }
```
The judge is Soup's own OpenAI-compatible `JudgeEvaluator` adapted to TRL's
`BasePairwiseJudge` (swap-debiased: a winner is only recorded when both A,B and
B,A orders agree). Recipe: `online-dpo-smollm2-135m`. Proof-of-mechanism was
validated on SmolLM2-135M with a synthetic judge (not a production RLHF claim; #286).
## Weighted Multi-Objective Preference Loss
Mix DPO / SimPO / ORPO / IPO terms in one training run by setting

View File

@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "soup-cli"
version = "0.71.30"
version = "0.71.31"
description = "Fine-tune and post-train LLMs in one command. No SSH, no config hell."
readme = "README.md"
license = "Apache-2.0"

View File

@ -1,3 +1,3 @@
"""Soup CLI — Fine-tune and post-train LLMs in one command."""
__version__ = "0.71.30"
__version__ = "0.71.31"

View File

@ -2141,6 +2141,42 @@ training:
target_modules: auto
quantization: none
output: ./output
""",
),
"online-dpo-smollm2-135m": RecipeMeta(
model="HuggingFaceTB/SmolLM2-135M-Instruct",
task="online_dpo",
size="135M",
tags=("smollm", "smollm2", "online_dpo", "judge", "rlhf", "tiny", "edge"),
description="SmolLM2 135M Online DPO — on-policy generation judged by a "
"pairwise LLM judge (point --online-dpo-judge at a local ollama model)",
yaml_str="""\
base: HuggingFaceTB/SmolLM2-135M-Instruct
task: online_dpo
data:
train: ./data/prompts.jsonl
format: auto
max_length: 1024
training:
# On-policy: the model generates 2 completions per prompt each step and the
# judge picks chosen/rejected. Point online_dpo_judge at a local judge, OR set
# reward_model instead (exactly one of the two).
online_dpo_judge: "ollama://llama3.1"
online_dpo_loss_type: sigmoid
online_dpo_max_new_tokens: 64
dpo_beta: 0.1
epochs: 1
lr: 5e-5
batch_size: auto
lora:
r: 8
alpha: 16
target_modules: auto
quantization: none
output: ./output
""",
),

View File

@ -260,7 +260,7 @@ class TestV025NewRecipes:
assert cfg.base == recipe.model
assert cfg.task == recipe.task
def test_catalog_size_is_137(self):
def test_catalog_size_is_138(self):
"""Total catalog size — grew with each release.
v0.25.0 shipped 43 recipes (29 + 9 Part A + 2 Part B tools + 3 Part E MLX).
@ -273,10 +273,11 @@ class TestV025NewRecipes:
v0.71.24 added 17 (2026 model-family expansion) -> 133.
v0.71.25 added 1 (qwen2.5-coder-7b-sft) -> 134.
v0.71.30 added 3 (grpo-env-calculator/retrieval-qa/guess-number) -> 137.
v0.71.31 added 1 (online-dpo-smollm2-135m) -> 138.
"""
from soup_cli.recipes.catalog import RECIPES
assert len(RECIPES) == 137
assert len(RECIPES) == 138
def test_new_recipes_searchable(self):
"""Search returns the new recipes via keyword/task filter."""

View File

@ -269,8 +269,8 @@ class TestGlm5RepoIdFix:
class TestCatalogCount:
def test_total_recipe_count_is_137(self) -> None:
assert len(RECIPES) == 137
def test_total_recipe_count_is_138(self) -> None:
assert len(RECIPES) == 138
def test_list_recipes_matches_dict(self) -> None:
assert len(list_recipes()) == len(RECIPES)

View File

@ -721,10 +721,10 @@ class TestRecipes:
assert cfg.training.rollout_backend == "openenv"
assert cfg.training.rollout_func.startswith("soup_cli.envs.")
def test_catalog_size_is_137(self):
def test_catalog_size_is_138(self):
from soup_cli.recipes.catalog import RECIPES
assert len(RECIPES) == 137
assert len(RECIPES) == 138
# ---------------------------------------------------------------------------

View File

@ -8,7 +8,6 @@ pairwise`` (#284), ``task='online_dpo'`` (schema + trainer + routing),
from __future__ import annotations
# ---------------------------------------------------------------------------
# Shared test doubles
# ---------------------------------------------------------------------------
@ -566,9 +565,10 @@ def _online_dpo_wrapper():
class TestBuildJudgeOrReward:
def test_judge_url_branch(self):
import soup_cli.trainer.online_dpo as od
from trl import BasePairwiseJudge
import soup_cli.trainer.online_dpo as od
od._ONLINE_DPO_JUDGE_OVERRIDE = None
result = _online_dpo_wrapper()._build_judge_or_reward(_Tcfg(judge="ollama://m"))
assert isinstance(result["judge"], BasePairwiseJudge)