3 macOS CI failures from the v0.55.0 push — Rich wraps option names
with ANSI escapes when the terminal is narrow (macOS CI runners
default to a smaller width than Linux/Windows), so substring searches
like `"--goal" in result.output` fail because the actual output
contains `\x1b[1;36m-\x1b[0m\x1b[1;36m-goal\x1b[0m`.
Project precedent: v0.53.5 / v0.53.6 / v0.53.8 / v0.53.9 all hit the
same pattern; tests/test_auto_tuning.py and tests/test_eval_platform.py
already ship `_ANSI_RE` + `_strip_ansi` helpers.
Failures fixed:
tests/test_v0550.py::TestCLIPlumbing::test_eval_design_help
tests/test_v0550.py::TestEvalAgainst::test_against_help
tests/test_v0550_followups.py::TestEvalAgainst::test_against_cli_help_lists_flag
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two failures on ubuntu/macos/windows × py3.9/3.11/3.12 after v0.54.0
push:
1. test_env_null_byte_falls_back: monkeypatch.setenv can't set raw
NUL into the OS env layer (POSIX execve + Win32 SetEnv both
refuse). Switched to a temporary `advise_history.os.environ` swap
so the helper's defence-in-depth NUL guard is still exercised
without going through the C env layer.
2. test_default_missing_data: Click 8.0–8.1 returns rc=0 on
`no_args_is_help=True` invocations; Click 8.2+ returns rc=2 (the
"missing command" convention). CI runners had the newer Click;
dev box had the older. Accept both renderings.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
`soup advise <data.jsonl> --goal "..."` returns one of PROMPT_ENG /
RAG / SFT / DPO / GRPO with a confidence, reason, and reverse-when
criterion BEFORE the user spends 8 hours on a GPU. Layer above
autopilot — autopilot picks hyperparams AFTER the training decision;
advise picks the training decision itself.
Three Parts:
- Part A: Verdict engine — TASK_CATEGORIES + CHOICES allowlists,
frozen Verdict / DatasetProfile / ROIEstimate dataclasses, pure-
Python classify_task + compute_dataset_profile + build_verdict
rubric (DPO / GRPO floor 500 / PROMPT_ENG floor 50 / RAG / SFT).
- Part B: Probe runner — synth_probe_baselines + synth_probe_lora_delta
heuristic stubs with forward-compat model/device/lr/timeout_seconds
kwargs (v0.54.1 lifts to live model loading per stub-then-live
cadence used by v0.27.0 MII / v0.37.0 multipack / v0.50.0 GRPO Plus).
- Part C: Cross-project learning — ~/.soup/advise_history.jsonl with
cross-process file locking (fcntl on POSIX, sidecar <path>.lock +
msvcrt on Windows). `soup advise compare` reads history; env
override SOUP_ADVISE_HISTORY_PATH containment-checked to $HOME /
$CWD / tempdir (mirrors v0.36.0 SOUP_BATCH_CACHE_PATH policy).
CLI: Typer subcommand group `run` / `explain` / `compare` plus argv
preprocessor in cli.py that maps `soup advise data.jsonl` →
`soup advise run data.jsonl`. Scoped to argv[1] == "advise" only
(code-review HIGH fix — defends against rewrites when an unrelated
arg contains the literal string "advise").
Schema: AdviseConfig (goal / probe / record) field on SoupConfig
honors the plan's cross-cutting bullet.
Security: cwd-containment + os.lstat + S_ISLNK symlink reject on
every path input; atomic writes via tempfile.mkstemp + os.replace
on scratch + history; per-line 64 KB cap + 16 MiB file cap on
history reads; bool / finite / NUL / oversize guards on every public
input; Rich markup escape on user-controlled output.
Reviewed by python / code / security / tdd / architect agents — every
finding fixed before commit (0 CRITICAL + 5 HIGH + 7 MEDIUM + 4 LOW).
Test count: 8400 → 8571 (+136 in tests/test_v0540.py, +35 net
adjustments to v0.53.x version-pin assertions to forward-compat >=).
Note: Windows CRLF / LF warnings during stage are .gitattributes-
governed and benign. CI runs on ubuntu-latest / windows-latest /
macos-latest × Python 3.9 / 3.11 / 3.12.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes v0.50.1 (#123, #126, #127), v0.49.1 (#119), v0.40.1 (#68).
#123 — live math kernels for 6 GRPO variants (gspo/dapo/dr_grpo/bnpo/
two_sided/rft) + `_GRPOTrainerVariant` HF Trainer subclass via
`make_grpo_trainer_variant` factory. Variant compute_loss reads kernel
inputs FIRST (no double-forward); falls back to super() only on missing
attrs. Case-insensitive variant normalisation before lru_cache.
#126 — PRMTrainerWrapper + `_PRMTrainer` HF Trainer subclass with real
compute_loss (gather hidden states at step_positions -> reward_head ->
MSE via compute_prm_loss). Dataset wrapped in datasets.Dataset.from_list
for HF Trainer compatibility. Bool-before-isinstance guard on batch_size.
#127 — GRPOStabilityCallback inherits transformers.TrainerCallback
(lazy), live EMA ref-model update in on_step_end with strict=True +
fallback-to-strict=False-with-WARNING on key mismatch (silent corruption
defence). math.isfinite guard on alpha.
#119 — LongLoRA forward override via LongLoRAForwardOverride context
manager with idempotent install (_soup_longlora_patched marker prevents
re-entry double-wrap), 256-char class name cap on regex match, restore
on __exit__ AND on exception.
#68 — true per-batch weighted-sum preference combine reading policy/ref
logps from TRL inputs + each compute_*_term kernel + combine_losses.
Explicit None checks on trainer attrs (no `or` on possibly-tensor),
DEBUG log on per-term skip.
Review fixes from 4 agents (python/code/security/tdd): 10 HIGH + 8
MEDIUM + 7 LOW — see CLAUDE.md v0.53.11 entry for the full list.
Test count: 8330 -> 8400 (+75 in test_v05311.py: 54 initial + 21
review-fix coverage gaps).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
7 issues closed:
- #150 [mix] pyproject extra bundles scikit-optimize so `soup data mix
--optimize` runs the Bayesian loop instead of the v0.48.0 Dirichlet
fallback; new describe_default_optimizer() helper labels the active
backend without paying skopt's import cost.
- #113 [data-pro] extras (langdetect + presidio-analyzer) with lazy
fall-through helpers in utils/data_score (broader language coverage +
Presidio entity recognition on top of the v0.47.0 regex baseline).
Llama-Guard-3-1B documented as a manual recipe (license + size).
- #154 SOUP_POSTHOG_KEY / SOUP_POSTHOG_ENDPOINT env override via
sentinel-based explicit-vs-env precedence; HTTPS-only +
RFC1918/link-local rejection on the endpoint; null-byte / control-char
/ >256-char rejection on the key.
- #152 --hub flag plumbed on chat / serve / infer / merge / export /
push via shared utils/hubs.apply_hub_to_cli_model +
prefetch_model_from_hub helpers; push uses upload_repo (skips
HF-specific Collections + model-card auto-render on non-HF hubs).
- #153 `soup data download --hub modelscope|modelers` live SDK
(lifts the v0.53.8 advisory-only path); friendly ImportError
advisory when the SDK is missing.
- #155 Web UI Tool Outputs panel — `loadToolOutputs` polls
/api/tool-outputs every 3s; XSS-safe DOM-built table (textContent
per cell, no innerHTML for user-controlled fields); Bearer token
threaded via the v0.53.9 window._authToken bootstrap.
- #156 SoupTrainerCallback.on_step_end records tool_calls counts
from kwargs['inputs'] into the global tool buffer. Best-effort
(# noqa: BLE001 per project policy — training must never crash).
13 review-fixes applied (4 HIGH / 5 MEDIUM / 4 LOW):
- HIGH PostHog explicit-endpoint precedence sentinel
- HIGH absolute path leak in local_path advisory reduced to relpath
- HIGH Rich markup escape on base / local_path / cache_dir
- HIGH callback # noqa: BLE001 per project policy
- MED `import time` moved out of try block
- MED oversize key + explicit-empty key rejection tests
- MED source-grep regression guards (advisory-removal, helper imports
across 5 non-push commands)
- MED `prefetch_model_from_hub` outside-cwd cache_root rejection
- LOW empty-list + bool-True tool_calls no-op tests
- LOW push.py uses upload_repo + validate_hub_name regression guard
Test count: 8285 -> 8330 (+45 in tests/test_v05310.py).
Full suite green; ruff clean; on Win+Py3.10.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI's Rich pipeline emits styled output that splits `--vocab-size` into
multiple ANSI-bracketed spans (e.g. `\x1b[36m-\x1b[0m\x1b[36m-vocab\x1b[0m\x1b[36m-size\x1b[0m`),
breaking naive `"--vocab-size" in result.output` checks. Locally Rich
auto-detects non-TTY and skips the codes, so the regression only shows
on CI (ubuntu/macos/windows × 3.9/3.11/3.12).
Fix: small `_plain()` helper using `re.sub(r"\x1b\[[0-9;]*m", "", ...)`
applied to the 7 failing assertions. Same approach already used in
several other v0.5x test modules.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Eight features that close out the v0.44.x live-monitoring deferrals
plus a long tail of standalone CLI wins:
- #94 /api/train/stream async SSE with per-subscriber cursor + JS
EventSource consumer; SoupTrainerCallback pushes TrainEvent on
each on_log.
- #95 soup ui --public derives LAN IP via SOCK_DGRAM connect-trick,
prints scannable QR; --auth-token override; SPA bootstrap
hydrates window._authToken from ?token= + sessionStorage and
cleans the URL via history.replaceState; CORS regex auto-widens
to loopback + RFC1918 in public mode; set_auth_token rotation
race fixed via threading.Lock.
- #98 soup serve --reasoning-parser strips <think>...</think> (and
OpenThinker tags); pre-compiled regex with marker-token
fast-path + 1 MiB cap + leading-newline-only strip.
- #100 ToolOutputsBuffer global singleton + /api/tool-outputs JSON
endpoint; best-effort observation hook in callback.on_log.
- #15 soup tokenizer train: BPE training CLI with raw-path lstat
symlink rejection, 50 MiB total / 8 KiB per-line caps,
post-mkdir output-dir re-check, --special-token NUL/oversize
dedup, vocab bounds [256, 200000].
- #26 soup bench --p50 --p95 renders extra per-prompt tail-latency
Rich table; --prompts-file gains symlink rejection.
- #28 soup bench --backend auto: MLX weights.npz probe (per-entry
lstat) -> config.json model_type keyword -> transformers
fallback; SOUP_BENCH_BACKEND env hint.
- #12 examples/synthetic_workflow.{md,yaml} end-to-end walkthrough.
Review fixes: 0 CRITICAL + 11 HIGH + 14 MEDIUM + 9 LOW across the
python / code / security / tdd review agents. Notable HIGH:
- QR token now consumed by SPA (was unreachable previously).
- set_auth_token rotation lock-protected, 8-thread stress tested.
- Tokenizer input + output symlink TOCTOU defence on raw path.
- SSE generator switched to async (asyncio.sleep) for non-blocking
multi-subscriber operation.
- CORS regex for --public LAN mode (the old fixed allowlist of
http://0.0.0.0:port never matched a real Origin header).
- _has_mlx_weights per-entry lstat so a symlinked weights.npz can't
trigger MLX dispatch.
Test count: 8257 -> 8285 (+57 in tests/test_v0539.py, minus the
relaxed v0.53.8 version-pin asserts in tests/test_v0538.py).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
v0.53.8 PyPI publish failed with:
400 Invalid distribution file. ZIP archive not accepted:
Duplicate filename in local headers
Root cause: `[tool.hatch.build.targets.wheel.force-include]` shipped
`soup_cli/data/_fixtures/` AND `packages = ["soup_cli"]` recursed into
the same path, so both the wheel and sdist contained each JSONL twice.
Fix: switch from force-include to `artifacts = [...]` which adds
non-Python files to the existing package tree exactly once. Standard
hatchling pattern for shipping data files inside an already-packaged
directory.
Version bumped to v0.53.8.1 (patch) — same code surface, just a build
config fix. v0.53.8 GitHub release remains as the feature changelog;
PyPI ships under v0.53.8.1.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Five v0.53.8 CI failures (ubuntu/macos × py3.9/3.11/3.12):
1. test_help_lists_hub_flag — Typer's Rich-rendered help wraps long
option help across ANSI box-drawing lines; "--hub" appears as
"│ --\nhub" in the CI terminal renderer. Strip ANSI + collapse
whitespace before asserting.
2-5. test_pyproject_version / test_*_extra_present / test_force_include
— used `Path("pyproject.toml")` (relative to cwd). CI invokes pytest
from a different cwd than the repo root on at least one matrix
entry. Switched to a `_repo_root()` helper that derives from
`__file__` (matches v0.43.0 Part D demo_bundles approach).
Local re-run: 66/66 v0.53.8 tests pass after the fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- #85 fsspec live loaders — data/loader.py routes the v0.42.0 fsspec
scheme allowlist (s3:// / gs:// / gcs:// / az:// / abfs:// / abfss:// /
oci://) through fsspec.open with validate_remote_uri containment
BEFORE connection. Friendly Rich panel names the pip install
advisory when the backend SDK is missing. Threads data.streaming
+ data.buffer_size. Row count capped at 1M.
- #130 Hub dispatcher live — utils/hubs.download_repo() and
upload_repo() lazy-import per backend (huggingface_hub /
modelscope / openmind_hub). Shared _validate_repo_id_shape (bool /
null-byte / leading-slash / .. / control-char / oversize) + cwd
containment on local_dir / folder_path. commands/train.py pre-fetches
non-HF base into .soup_hub_cache/ (sanitised slug, idempotent on
resume, cfg.base updated via model_copy). soup data download --hub
flag plumbed. Multi-command rollout for chat / serve / infer / merge
/ export / push tracked for v0.53.9.
- #89 [trackers] pyproject extra bundles mlflow / swanlab / trackio;
tracker_missing_dep_message surfaces a friendly pip install advisory
via importlib.util.find_spec (non-executing probe).
- #90 utils/trackers.send_telemetry_payload — opt-IN via SOUP_TELEMETRY=1;
lazy httpx; 1s hard timeout; HTTPS-only with SSRF re-validation
(mirrors v0.51.0 hub endpoint policy); silent-fail on every exception.
- #93 Fixtures migrated to soup_cli/data/_fixtures/ — zipapp /
namespace-package safe via [tool.hatch.build.targets.wheel.force-include];
_bundle_source_path falls back to examples/data/ for editable installs.
- #69 utils/hf_space.detect_space_sdk(requirements_text) — picks
"streamlit" / "gradio" from the rendered requirements.txt; closes
the v0.40.2 known limitation that custom Spaces always defaulted to
gradio. Wired into commands/deploy.py.
Review pass: python-review + code-review + security-review ran in
parallel; 16 findings fixed (3 HIGH + 8 MEDIUM + 5 LOW). Highlights:
cwd-containment on local_dir/folder_path, Windows ..\ traversal
defence on .soup_hub_cache slug, Pydantic model_copy(update=...)
instead of attribute mutation, idempotent pre-fetch via cache probe,
1M-row cap on remote materialisation, SSRF re-validation on
telemetry endpoint override, 256 KB cap on detect_space_sdk input,
modelscope.push_model commit_message kwarg removed (would TypeError
at runtime), find_spec instead of __import__ to avoid swanlab
side-effects.
Test count: 8162 -> 8257 (+66 in tests/test_v0538.py + 29 net adjustments).
Lint clean. CPU smoke: version, --help, load_config_from_string with
hub: modelscope passes; mlx + non-HF rejected; data download --hub
modelscope advisory rendered; detect_space_sdk live on real
requirements.txt bodies; package-data fixtures resolve from
soup_cli/data/_fixtures/.
v0.53.7 known limitation #1 (bash 501 marker) bumped to v0.53.9.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
test_load_jsonl_rows_rejects_symlink: when the symlink target is outside
cwd, is_under_cwd's realpath resolution catches it before the lstat check
fires. Both rejections are valid security guards; broaden the regex to
match either error message ("symlink" or "under cwd").
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI run 25805598433 failed across all 9 OS×Python cells:
- TestToolEndpointsLive / TestAnthropicMessagesStreaming /
TestReviewFixesVllmAnthropicLive ModuleNotFoundError: fastapi (CI does
not install [serve] extra) → autouse fixture pytest.importorskip()
- test_load_pretokenized_dataset_rejects_symlink:
load_pretokenized_dataset called datasets.load_from_disk on the symlink
target before the lstat check ran → moved the lstat + S_ISLNK check
to the entry of the helper so symlinks reject before any load attempt
- test_redact_exc_message_handles_windows_paths: hardened _redact_exc_message
to strip both POSIX absolute paths and Windows-style paths regardless
of host platform
Local pytest tests/test_v0537.py: 112 passed, 7 skipped.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI failures on `tests/test_v0536.py`:
1. `ModuleNotFoundError: No module named 'fastapi'` on Ubuntu/macOS runners
— fastapi is in the [serve] extra, not [dev]. Added
`pytest.importorskip("fastapi")` to all 7 tests that use TestClient
(matches the pattern in `_build_app`).
2. `--execute` / `--output` / `v0.53.7` substring checks failed because
CI terminals render Typer/Rich help text with style spans (`-` and
`-execute` end up in separate `\x1b[...]m` runs). Added `_strip_ansi`
helper + wrapped 5 substring checks. Same fix pattern as the v0.53.5
`--live` CI fix on test_v0535.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Rich splits the `--live` token across ANSI colour codes on narrow CI
terminals (`-\x1b[0m\x1b[1;36m-live`), so the raw-output substring
assertion failed on macOS/Windows runners but passed locally on a wide
terminal. Strip ANSI escapes before the check — matches how earlier test
files (e.g. test_v0402_part_b) handle Rich-coloured `--help` output.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Six closes lifting the v0.49.0 LongLoRA hardening + v0.41.0 LLaMA Pro
deferred stubs, plus a UX upgrade to the CUDA-OOM friendly message:
- #11 utils/errors.py: OOM hint now names --batch-size / --grad-accum
- #122 flash_attn.is_flash_attn_v3_available() + LongLoRA+FA3 schema reject
- #120 LongLoRA arch allowlist expansion (Mistral / Qwen / Phi); Mixtral
intentionally excluded (regex matches the bare 'mistral' token only)
- #121 apply_long_context_config auto-detects 'llama3' when caller passes
rope_scaling_type=None and the model config carries a Llama 3.1
rope_scaling block
- #83 block_expansion.expand_model_blocks LIVE (deepcopy last-N blocks,
zero-init residual projections, append, bump num_hidden_layers) +
apply_llama_pro_freeze + shared apply_block_expansion_if_configured
helper wired into SFT + Pretrain (mirrors v0.40.6 peft_wiring
centralisation policy so SFT and Pretrain stay in lock-step)
- #74 HF push surface QA — test plan recorded in tests/qa/v053_qa.md;
live execution against a private HF repo deferred to a credentialed
contributor
Review pipeline (python / code / security / tdd agents) ran; every
CRITICAL -> LOW finding addressed:
- bool-first guards in _check_model_name (defends against int subclass)
- is_supported_longlora_arch defensive non-string surface (returns False,
never raises) matching v0.53.3 is_known_vlm_base policy
- _truncate_for_message(value, limit=64) bounds the base echo in
LongLoRA error messages (security MEDIUM, mirrors v0.34.0 crash.py)
- null-byte + non-string TypeError guards on validate_longlora_compat
task / backend params (matches v0.50.0 validate_long_context_grpo_compat)
- _get_layers_module uses explicit `is None` not falsy shortcut (defends
against nn.Module.__bool__ overrides on subclasses)
- _zero_init_block_residual returns bool + warnings.warn when neither
standard projection matches the cloned block (non-Llama-shaped arches
still train but lose the LLaMA Pro identity-init guarantee)
Test count: 7879 -> 7935 (+56 net; +49 in new tests/test_v0534.py).
Lint clean. CPU smoke verified on a real transformers.LlamaForCausalLM:
4 -> 6 layers, down_proj + o_proj actually zeroed on PyTorch tensors,
old blocks frozen + new blocks trainable, forward pass finite.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two surgical fixes from the v0.50.0 GRPO Plus deferred-stub family land:
- #128 grpo_fp16 routing: GRPOTrainerWrapper._build_precision_kwargs
returns {fp16, bf16} per (device, grpo_fp16) matrix (CPU/MPS/XPU →
both False, CUDA + grpo_fp16=True → fp16/!bf16, default CUDA →
legacy bf16). SoupConfig._validate_grpo_fp16_amp_exclusive rejects
the silent-mutex combo with auto_mixed_precision=True; short-circuits
when task != 'grpo' so the v0.50.0 task-gate diagnosis fires first.
- #129 vision-GRPO base probe: KNOWN_VLM_REGEX covers 10 VLM families
(Qwen2-VL/Qwen2.5-VL/QVQ/Pixtral/InternVL/Llama-3.2-Vision/LLaVA/
MiniCPM-V/Idefics/ShareGPT4V/Fuyu) with word-boundary anchors;
is_known_vlm_base returns False (never raises) on bad input;
validate_vision_grpo_compat now accepts optional base kwarg with
64-char error-message truncation. YAML pairing vision_grpo: true
with a non-VLM base is rejected at schema load with a friendly
families listing instead of a cryptic runtime AttributeError.
Scope: 4 larger v0.53.3 items (#127 stability callback, #123 GRPO
variant losses, #126 PRMTrainerWrapper, #68 multi-objective preference
live combine) are scope-deferred to v0.53.4 — each warrants its own
focused release per the v0.40.x stub-then-live cadence.
Tests: 7842 -> 7879 (+37 in tests/test_v0533.py). Four review agents
(python/code/security/tdd) ran; every HIGH/MEDIUM/LOW finding fixed
(task-gate priority short-circuit, MPS branch documented, 64-char
error truncation, QVQ regex coverage, 512-byte boundary test).
Two pre-existing v0.50.0 Part E tests migrated `base: test-llama` ->
`base: Qwen/Qwen2-VL-7B-Instruct` to clear the new probe.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The three CLI tests added in the previous commit (test_help_lists_measure_flag,
test_save_format_help_lists_flag, test_torchao_help_lists_quant_config)
assumed the literal option name (e.g. `--measure`, `--save-format`,
`--quant-config`) would appear contiguously in CliRunner-captured
output. On CI runners the terminal defaults to 80-col and Rich wraps
long option names across lines, splitting the literal string.
Fix: walk `typer.main.get_command(app).params` and collect every `opt`
+ `secondary_opt` into a set, then assert membership. The test now
verifies what we actually care about (the option is registered) without
depending on Rich's wrapping behaviour.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Lift six v0.53.0 deferred stubs from NotImplementedError to live wiring:
- #82 autopilot pre-quantized base detection
utils name regex over gptq/awq/aqlm/eetq/fp8/mxfp4 with word-boundary
anchoring + HQQ Nbit extraction + config.json quantization_config probe
(cwd-contained + symlink-rejected). decide_quantization() short-circuits
the VRAM heuristic when prequantized is set. autopilot pipeline auto-
applies so TheBloke/Llama-2-7B-Chat-GPTQ is recommended gptq instead of
4bit-on-top-of-quantized.
- #142 merge_4bit + export_torchao live writers
soup merge --save-format {fp16|4bit|4bit_forced}: single BNB-4bit
merged checkpoint without the dequant->merge->requant cycle (fixes
wrong-name llm_int8_skip_modules to bnb_4bit_skip_modules per code-
review). soup export --format torchao --quant-config <yaml>: torchao
.quantize_ + save_pretrained with per-scheme closed kwarg allowlist
(Int4WeightOnly accepts {group_size, inner_k_tiles}, NVFP4 accepts
nothing extra; dunder + unknown keys rejected per security-review H1).
load_quant_config enforces yaml.safe_load + 256 KB cap + extension
allowlist + cwd containment + S_ISLNK rejection.
- #139 export_advanced_gguf via llama.cpp imatrix
3-stage pipeline: convert_hf_to_gguf.py -> optional imatrix ->
quantize. argv-list subprocess (no shell), 30-min timeout, realpath-
verified convert script stays inside llama_cpp_dir (security-review
M5). _prepare_calibration_text accepts JSONL with text/prompt/content
field aliases + raw text fallback; strips null bytes, collapses
newlines, 8 KB per-line + 50 MB total cap (security-review M1); POSIX
O_NOFOLLOW closes the TOCTOU window between dispatch-time check and
open() (security-review M3). UD- prefix stripped before passing to
llama-quantize. _safe_stderr Rich-escapes subprocess stderr before
embedding in RuntimeError (security-review L4).
- #109 soup deploy autopilot --measure
Live Quant-Lobotomy scorecard: classifies each candidate quant OK /
MINOR / MAJOR (thresholds 2% / 5% mirror v0.26.0 Part D). Results
cached at ~/.soup/deploy_autopilot_cache.json (atomic write, 0o600
perms on POSIX, S_ISLNK rejection on BOTH load and save). pick_best
soft-fallback now picks max-by-delta (was max-by-after) matching the
v0.33.0 #54 design intent. _DEPLOY_MEASURE_BEFORE_GEN / _AFTER_FACTORY
module-level hooks act as the stop-gap escape hatch until v0.46.1
ships first-party transformers / vLLM generator factories.
- #70/#72 manual QA log scripted at tests/qa/v053_qa.md with exact
reproduction recipes + acceptance criteria for the CUDA + llama.cpp
smokes that can't run on the CI runners.
Shared cleanup:
- soup_cli/utils/paths.enforce_under_cwd_and_no_symlink consolidates the
v0.33.0 #22 TOCTOU pattern previously copy-pasted in save_formats.py
and gguf_quant.py (code-review HIGH fix).
Reviews ran: python / code / security / tdd. Every CRITICAL / HIGH /
MEDIUM / LOW finding fixed or documented.
Test count: 7610 -> 7722 (+112 across 4 new files).
Known limitations: live GPU + bitsandbytes / torchao smokes for the
new merge / export paths remain pending (recipes in QA log); injected-
generator escape hatch is non-public until v0.46.1; cache key truncates
base_sha to 16 hex (1-in-2^32 collision floor).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
22 features across 5 internal Parts shipped as schema-only — closed
allowlists, Pydantic validators, NotImplementedError stubs for live
wiring deferred to v0.50.1 (mirrors v0.27.0 MII / v0.37.0 multipack /
v0.41.0 LLaMA Pro / v0.45.0 plugins / v0.48.0 curriculum / v0.49.0
LongLoRA stub-then-live pattern).
Part A — 7 GRPO objective variants (gspo / dapo / dr_grpo / bnpo /
two_sided / rft / standard) with `validate_grpo_variant` + frozen
`GRPOVariantSpec` metadata + `MappingProxyType`-wrapped registry.
`validate_grpo_delta` is bool-first / math.isfinite / (0, 1] bounded.
`apply_variant_loss` raises NotImplementedError with v0.50.1 marker
for the 6 new variants and is a no-op for standard.
Part B — long_context_grpo + vllm_sleep_mode schema gates with
compat validators (null-byte rejection on task + backend, bool guard
on use_ring_attention). vllm_sleep_mode requires task='grpo' AND a
transformers/unsloth backend (code-review HIGH fix — sleep is a
between-rollouts feature).
Part C — 4 multi-turn rollout backends (art / ruler / nemo_gym /
openenv) with closed allowlist + per-entry required_package mapping.
Part D — 7 stability/efficiency knobs (ref_model_ema_alpha,
replay_buffer_size, async_grpo_prefetch, tis_threshold,
mask_truncated_completions, defer_rerolling, skip_zero_advantage,
off_policy_mask_threshold). Every numeric field rejects bool via a
shared `_reject_bool_on_grpo_numerics` field_validator (tdd-guide
HIGH fix — Pydantic v2 coerces True→1 by default). The
`mask_truncated_completions` + `tis_threshold` pairing is enforced
by a cross-validator (matches v0.32.0 spike-recovery+watchdog
policy).
Part E — top-level task='prm' Literal addition (Process Reward
Model / stepwise-supervised, paired with data.format='prm' from
v0.42.0) + `vision_grpo: bool` flag for VLM-RL on Qwen2-VL /
Pixtral / InternVL. Compat helpers gate task / modality / backend.
Review-round fixes applied (5 sequential reviews per CLAUDE.md):
- python-review: list_variants annotation, frozenset[str] params,
Optional[str] → str | None, module-level math import, D401
imperative docstrings, dropped *args/**kwargs on stubs.
- code-review: grpo_fp16 added to GRPO-only task-gate;
vllm_sleep_mode requires task='grpo'.
- security-review: explicit field_validator for grpo_delta NaN/Inf
rejection (Pydantic le=1.0 incidentally rejects NaN, made
explicit); null-byte rejection on backend/task in grpo_long_context
helpers; use_ring_attention bool guard.
- tdd-guide: bool-rejecting field_validator on all Part D numeric
fields + grpo_delta; missing bool-rejection tests added on
validate_grpo_variant / validate_rollout_backend; null-byte test
on validate_vllm_sleep_mode_compat; required_rollout_package
rejection path; RolloutBackendSpec.live_wired; PPO+vision_grpo
round-trip; _DEFERRED_LIVE invariant.
Test count: 6490 → 6729 (+239 across 5 new test files).
Notes for future maintainers:
- v0.50.0 has zero new CLI commands and zero new trainer wirings;
every step 6d/6e is intentionally n/a. Step 6 smoke runs schema
happy + every documented cross-validator rejection.
- All `task='grpo'` gates use `if self.task != 'grpo'` literal
comparisons; do NOT switch to a set membership check until Part D
knobs are wired into PPO/preference trainers in v0.50.x.
- Multi-modal Vision RL does not yet verify the base model is
actually a VLM — upstream trainer surfaces the error loudly when
it fails to load the vision tower.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
BETA. Adds `training.curriculum_dynamic: true` schema flag with online
uncertainty estimation: every N steps, aggregate per-sample loss + grad-norm
into per-bucket softmax weights, water-filled to enforce a minimum
`curriculum_dynamic_floor`. DDP/grad-accum safety via
`validate_distributed_curriculum` cross-validator that rejects un-coordinated
multi-rank runs upfront — the well-known footgun where divergent per-rank
stats silently desynchronise the sampler.
New `soup runs curriculum-curve <run_id>` visualiser with TOCTOU
(`os.lstat + S_ISLNK`) + 50 MB file-size cap + 100k-line streaming cap on
the history file. Schema gated to sft/pretrain on transformers backend;
mlx + non-SFT rejected with distinct messages.
Live HF Trainer callback wiring deferred to v0.48.1 (stub-then-live
pattern; mirrors v0.27.0 MII / v0.37.0 multipack / v0.41.0 LLaMA Pro).
Review fixes:
- water-fill design fix (code-review HIGH): removed trailing renorm that
could push elements below `floor` when accumulated float error left
sum slightly > 1.0. Softmax already sums to 1.0, so water-fill output
also sums to 1.0 (drift bounded by nb*eps).
- DoS caps on `render_curve` + `parse_history_jsonl`
(`_MAX_HISTORY_ROWS=100_000`) — without these an attacker-controlled
JSONL with 10M rows would OOM the process.
- `curriculum-curve` CLI: symlink rejection + 50 MB + 100k-line caps,
null-byte rejection on tracker-supplied `output_dir`.
+74 tests.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds a public plugin/hook system plus the schema scaffolding for 20+ ecosystem
integrations. Live trainer-callbacks, Anthropic /v1/messages route, server-tool
HTTP endpoints, and the recipe runner ship in v0.45.1 (matches v0.27.0 MII /
v0.37.0 multipack / v0.41.0 LLaMA Pro stub-then-live pattern).
Part A — Plugin / hook system
* New soup_cli/plugins/ package: BasePlugin Protocol, PluginSpec frozen
dataclass, register_plugin / discover_hooks / enable_plugin /
disable_plugin / load_plugins. Kebab-case name regex, semver-ish version,
null-byte rejection on every string. Idempotency check covers
(version, plugin object, templates, model_groups, description) — review-fix
added description after first-cut omitted it. Per-list caps on templates
and model_groups (32 entries, 128-char per name).
* New soup plugins list/install/enable/disable Typer CLI; all user-controlled
output passes through rich.markup.escape.
Part B — API extensions (schema-only)
* utils/anthropic_messages.py: to_anthropic / from_anthropic /
validate_anthropic_payload converters. Multiple system messages join with
\n\n; tool role with structured (list) content concatenated into single
tool_result text block (review-fix MEDIUM — first-cut silently dropped).
max_tokens cap 16384, temperature [0.0, 2.0], bool rejection on numerics.
* utils/server_tools.py: closed {python, bash, web_search} allowlist,
WebSearchConfig with domain allowlist + leading-dot subdomain pattern,
rate_limit [1, 600]. is_domain_allowed strips :port suffix and rejects
IPv6 literals (review-fix MEDIUM).
* utils/ngram_spec.py: NgramSpecConfig validators with bounded n / draft
tokens / prompt-lookup-max; bool rejection on every numeric field.
Part C — External integrations catalog
* utils/integrations.py: 15-entry MappingProxyType catalog of ecosystem
targets (lm-studio, comfyui, ollama, claude-code, cursor, continue, ...).
Part D — Advanced trainer-plugin allowlist
* utils/trainer_plugins.py: 6-entry allowlist (grokfast, spectrum,
llmcompressor, sonicmoe, cce_plugin, math_verify) + validate_trainer_
plugin_list (Sequence[str], dedup, _MAX_PLUGINS_PER_RUN=8).
Part E — Data Recipe DAG
* utils/recipe_dag.py: closed NODE_KINDS frozenset, Kahn's topological
sort via collections.deque (review-fix HIGH — first-cut had O(N^2 log N)
queue.sort() inside the BFS body), cycle / self-loop / dangling-edge
rejection, _MAX_NODES=256 / _MAX_EDGES=1024 / _MAX_FILE_BYTES=1MiB.
load_recipe_yaml enforces is_under_cwd containment AND os.lstat + S_ISLNK
symlink rejection (review-fix MEDIUM — TOCTOU defence; mirrors v0.33.0 #22
/ v0.43.0 Part C / v0.44.0 Part B policy).
* New soup data recipe <path> CLI validates topology and prints planned
topo order; live runner deferred to v0.45.1.
Reviews: python-review, security-review, code-review, tdd-guide all run;
verification-loop replaced by manual smoke (CLI happy + failure paths
exercised on real fixtures).
Test count: 5820 -> 5989 (+169). Test files: 164 -> 165. Ruff clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Schema-first surface for the data pipeline gap with Axolotl + LlamaFactory.
Ships in one release: 5 new formats (prm, pre_tokenized, input_output,
video, multimodal), remote URI allowlist (s3/gs/gcs/az/abfs/abfss/oci) +
streaming + sharding, AOT preprocess cache + `soup data preprocess` CLI,
multi-dataset interleave (concat/under/over/probs) + 8 advanced masking
fields, vocab expansion (add_new_tokens / new_special_tokens / resize_vocab)
+ custom prompt_strategy, and document ingestion (`soup data ingest` for
PDF/DOCX/MD/TXT).
Live wiring for fsspec backends, AOT tokenize loop, custom prompt-strategy
runtime, and PRM trainer integration is deferred to v0.42.1+ (stub-then-live
pattern from v0.27.0 / v0.37.0 / v0.41.0). Schema gates fire at config
load so misconfiguration fails fast.
Security: full v0.42.0 hardening matrix — `_REMOTE_SCHEMES` MappingProxyType
allowlist; bucket regex 1-63 chars per S3/GCS spec; userinfo / fragment /
query-string rejection on remote URIs (query-string forwarded to fsspec is
SSRF-adjacent); null-byte + length caps on every string-shaped input;
bool-rejected-before-int on every numeric input; frozen InterleaveSpec
dataclass; 10k caps on add_new_tokens; `is_under_cwd` containment on
video_dir + tokenized_path schema fields and on preprocess --config / both
ingest paths; `os.lstat + S_ISLNK` symlink rejection on ingest input;
PRM converter type-checks completions (str) + labels (bool, not int);
video field null-byte + 2KB cap; field-name threading on image-pixels
validator so error messages name the actual field.
5242 tests → 5389 (+147 net). 11 review findings addressed across
python-review / code-review / security-review / tdd-guide (CRITICAL→LOW).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Extends the v0.38.0 train-time quantization menu (gptq / awq / hqq:Nbit /
aqlm / eetq / mxfp4 / fp8) from SFT-only to all 11 transformer-backend
trainers (DPO / GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel /
Pretrain / Embedding / BCO). Closes the v0.38.0 known gap.
Schema gate: SoupConfig._validate_quant_menu_supported_tasks removes the
`task != "sft"` rejection branch. MLX backend rejection retained with
distinct message; modality=text gate retained (vision/audio Quant Menu
deferred — modality-specific kwargs not yet threaded through the unified
loader).
Trainer wiring (11 sites): each non-SFT _setup_transformers replaces the
inline BitsAndBytesConfig(load_in_4bit=True, ...) block with a call to
build_quantization_config_for_loader(tcfg=tcfg, base=cfg.base, console=console)
— mirrors sft.py:420-440 exactly. kbit-prep tuple widened from
("4bit", "8bit") to ("4bit", "8bit", "mxfp4"). BitsAndBytesConfig import
removed from each non-SFT trainer.
Review fix — PPO reward model: _load_reward_model gains optional tcfg
kwarg; when supplied, the reward checkpoint loads with the same Quant
Menu config as the policy. Both PPO call sites forward tcfg=tcfg.
Defends against silent fp16 OOM on a GPTQ/AWQ/HQQ policy run.
Review fix — defence-in-depth: new TrainingConfig.reward_model field
validator rejects null bytes and caps length at 512 chars (matches
cfg.base policy). The Quant Menu loader's per-call null-byte check
in _check_local_marker remains as the runtime backstop.
Tests: +131 net new (4930 -> 5061). New tests/test_v0405_part_a.py
parametrizes 11 tasks x 7 quant formats; covers MLX rejection per task,
quantization_aware x Quant Menu cross-validator regression for non-SFT,
source-level invariants (no inline BNB literal, kbit tuple regex),
and a live mock-based dispatch test for _load_reward_model proving the
Quant Menu path is reachable when tcfg is supplied and skipped when
tcfg=None.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI failure root cause: Rich line-wraps `--trust-remote-code` with ANSI
colour escapes between `-`, `-trust`, `-remote-code` on narrow CI
terminals (mirrors v0.40.3 ANSI fix). Substring assertion missed
because the ANSI escapes were embedded mid-flag.
Also: original test used `[cmd, "--help"]` then fell back to
`["data", cmd, "--help"]`. For diff/export/merge/infer the first
invocation worked (top-level commands) but the substring miss
triggered the fallback into `data` subcommand, which then errored
"No such command 'X'" — masking the real ANSI issue. Switched to
explicit per-command argv lists.
Adds the `_strip_ansi` helper from tests/test_trust_remote_code.py
and routes `data generate` directly via `["data", "generate", "--help"]`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#65 (deferred from v0.40.3).
make_multipack_trainer_class adds a get_train_dataloader override
that builds a MultipackBatchSampler(real_batches=False) — yields a
flat list[int] per packed sequence, which is the contract HF
DataLoader.batch_sampler expects — and installs it via
DataLoader(..., batch_sampler=sampler, collate_fn=self.data_collator,
num_workers=args.dataloader_num_workers,
pin_memory=args.dataloader_pin_memory). drop_last is forwarded from
TrainingArguments.dataloader_drop_last.
Falls back to super().get_train_dataloader() when state was never
attached OR when train_dataset is unset — defence-in-depth so the
subclass remains safe to instantiate even when multipack is later
disabled.
The state-presence guard switched from falsy (`not max_seq`) to
explicit `is None` (plus `not lengths` for empty-list defence) —
attach_multipack_state already rejects non-positive ints, so the
falsy guard would only mask configurator bugs.
_get_train_sampler override stays as a defensive no-op fallback that
ALWAYS delegates to super (review-fix from v0.40.4 code-review:
returning a multipack list[list[int]] from this method would cause a
shape mismatch if any HF eval / prediction loop bypasses
get_train_dataloader and calls _get_train_sampler directly).
SFT and Pretrain trainer wrappers now invoke
make_multipack_trainer_class(SFTTrainer) and attach_multipack_state(...)
when multipack: true. The v0.40.3 yellow advisory + standard-sampler
fallback is gone. Architecture allowlist
(validate_multipack_architecture) still gates at build time.
tests/test_v0403_part_b.py: TestSftAndPretrainWiringDeferred renamed
to TestSftAndPretrainWiringLive; the deferred-state test
(_get_train_sampler returns MultipackBatchSampler when state is set)
is replaced by the live-state test (_get_train_sampler always
delegates to super even with state attached).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes the v0.36.0 #63 known gap. Every non-SFT trainer wrapper (DPO /
GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel / Pretrain /
Embedding / BCO + the unified Preference dispatcher) now accepts
trust_remote_code: bool = False on __init__, resolves once via the
v0.36.0 helper (model_requires_trust_remote_code +
resolve_trust_remote_code), and stores the resolved value on
self._trust_remote_code. Every from_pretrained call site reads from
the resolved attribute — no remaining trust_remote_code=True literal
in any trainer file (asserted by tests/test_v0404_part_a.py).
Five standalone commands gain a --trust-remote-code Typer flag with
the same default-deny + KNOWN_SAFE_PREFIXES allowlist behaviour as
soup train: soup diff, soup export, soup merge, soup infer,
soup data generate.
commands/train.py removes the v0.36.0 sft_kwargs split — every trainer
receives trust_remote_code from the same trainer_kwargs dict.
PreferenceTrainerWrapper forwards the raw bool to the inner DPO /
SimPO / ORPO / IPO / BCO wrapper kwargs at both _build_inner and
_build_multi_objective sites; the resolver fires inside the inner
wrapper at construction time.
_load_reward_model (module-level helper in ppo.py) accepts a
trust_remote_code: bool parameter and resolves internally — design
intent is that the helper is independently safe to call outside
PPOTrainerWrapper.
_export_onnx / _export_tensorrt / _export_awq / _export_gptq and
_merge_adapter helpers all gain a trust_remote_code: bool = False
parameter threaded from the Typer flag.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI failure on previous v0.40.3 hotfix: Typer renders Rich-styled help
with ANSI escape codes BETWEEN flag fragments (`--trace\x1b[0m\x1b[1;36m-log`),
so a whitespace-only strip still failed to find `--trace-log` in the
flattened output. Strip both ANSI codes (`\x1b\[[0-9;]*m`) and whitespace
in one pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI failures on this commit:
- macOS Typer help text wraps `--judge` / `--trace-log` to two lines on
narrow CI terminals; tests asserted the raw string. Strip whitespace
before match (mirrors v0.40.2 width-independent fix).
- `fastapi` is not in the base CI deps (only `[serve]` extra); two
`_create_app` tests ImportError-ed. Skip those tests when fastapi is
unavailable.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three v0.X.0 deferred-stub features become live runtime — closes#33
(harvester judge filter + serve trace log) and #64 (live CUDA OOM probe).
#65 (multipack live wiring in HF Trainer) remains deferred to v0.40.4
after the adversarial 5th-review pass surfaced a Sampler[int] vs
list[list[int]] shape mismatch with HF Trainer's DataLoader; helpers
(`make_multipack_trainer_class`, `attach_multipack_state`,
`lengths_from_dataset`, `detect_arch_name`) ship as a stub used by
unit tests, but the SFT/Pretrain wrappers print a yellow advisory and
fall back to the standard sampler when `multipack: true`.
Live CUDA OOM probe (#64): `make_cuda_probe_fn` builds a closure that
runs ONE forward+backward+step on a synthetic batch per candidate.
`model.zero_grad(set_to_none=True)` runs BEFORE forward; intermediate
ids/attn/labels/outputs are del-ed before `loss.backward()` so peak
VRAM reflects a realistic training step (matches v0.35.0 #45 policy).
`pad_id` is bounded by `len(tokenizer)` (not `vocab_size`) so extended
vocabs (Llama-3 + `<|pad|>`) don't fold pad to a random byte token.
SFT-only this release.
Trace-to-Preference judge filter (#33 (a)): `judge_filter_pairs` reuses
v0.19.0 JudgeEvaluator backends (openai/server/ollama). Threshold
rejects bool/NaN/out-of-[0,1]; `_MAX_BATCH=100_000` cap applied via
lazy `itertools.islice`; per-pair backend exceptions caught and DEBUG-
logged (matches v0.33.0 #47 policy); `judge_provider` validated against
the allowlist at the CLI boundary BEFORE constructor with a Rich-escape
error message; yellow projected-call-count warning before the loop
(2× per pair).
Inference Server trace log (#33 (b)): `TraceLogWriter` is thread-safe
(single-process lock — multi-worker documented as known limitation);
path containment via shared `is_under_cwd`; null-byte/empty/non-string
path rejected; cap_mb bounds [1, 10000] with explicit bool rejection.
Rotation (one backup retained) refuses symlink at the backup path via
`os.lstat + stat.S_ISLNK` (matches v0.33.0 #22 TOCTOU policy). Secret
redaction (`hf_*` ≥8, `sk-*` ≥16, `Bearer …` ≥8 with `.` excluded so
end-of-sentence period survives) applied to prompt + response and
recursively to caller-supplied `extra` dict values. Streaming SSE path
also records (was a coverage gap caught in adversarial review).
Behaviour change: v0.40.2 users with `auto_batch_size_strategy: probe`
were silently getting the static fallback. v0.40.3 actually runs a
CUDA probe on first run (~5–30s, cached per (model, max_length, quant,
lora_r, gpu) tuple).
Reviews: 5 agents (python, code, security, tdd, verification-loop).
Verification-loop run twice — once shallow smoke (PASS), once
adversarial bug-hunt which found C1/C2 (multipack live wiring crash —
demoted to v0.40.4), H1 (streaming SSE missing trace log — fixed),
H4 (vocab_size vs len(tokenizer) on extended vocabs — fixed), H3
(Bearer regex consumed trailing period — fixed), H2 (judge cost
shock — warning added), M2 (empty lengths accepted — rejected), L1
(extra dict bypassed redaction — recursive walk added).
Tests: 4756 → 4855 (+99 net new) across test_v0403_part_a/b/c.py.
Lint clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI on Linux/macOS runners has narrower terminals than the local Windows
shell. Rich wraps long option names like ``--template-dir`` across two
lines (``-\n-template\x1b...-dir``) which makes a substring check on the
raw output string fail.
Updated `_plain` helper in both v0.40.2 test files to strip whitespace
in addition to ANSI escapes — matches the option name regardless of
terminal width.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes 3 originally-scheduled GitHub issues plus 7 v0.40.1 long-tail UX
papercuts. No new schema fields, no new trainers — pure polish.
Originally scheduled:
- #36 format_gate_row helper for the eval-gate dashboard row (pure formatter
in soup_cli/monitoring/display.py; passed=is True so missing field renders
neutral; supports stop/warn action suffixes; multi-task " | " join).
- #50 prepare_hf_resume now skips snapshot_download when local checkpoint-N
is greater-or-equal to the remote highest-N. New _find_highest_local_checkpoint
helper handles missing dirs / OSError / non-directories cleanly.
- #51 soup deploy hf-space --template-dir <path> via new
soup_cli/utils/hf_space.py:render_custom_template_dir. Containment via
is_under_cwd; validate_repo_id BEFORE substitution; per-file 256 KB cap;
symlinks + non-regular files rejected (TOCTOU defence per v0.33.0 #22).
v0.40.1 carry-overs:
- H2: data filter --min-coherence alias; data split --train no-op; data
register/unregister positional <name> <path> + Optional --name/--path
with conflict detection.
- H3: soup quickstart --output DIR (containment-checked) routes data,
config, run dir under the chosen directory.
- N1/G2: apply_logging_level pushes parsed --log-level tier into the root
logger so transformers / peft / trl actually respect QUIET / DEBUG.
- N7: shared _resolve_model_source in commands/infer.py (used by bench.py
too) — path-like-but-missing raises FileNotFoundError; non-path-like
values fall through to HF download via from_pretrained.
- G13: verified ONNX/AWQ/GPTQ/TensorRT install hints already correct.
- M4: verified data dedup --threshold already exposed.
- M5: soup runs --cwd-only + _filter_runs_by_cwd helper using
os.path.realpath + commonpath (Windows 8.3 + cross-drive safe).
Review-fix follow-ups landed in the same release:
- soup_cli/commands/infer.py: from __future__ import annotations (Py3.9
PEP 604 fix); --output containment via is_under_cwd, late-evaluated to
preserve pre-existing test contracts.
- soup_cli/commands/data.py register_data + soup_cli/commands/bench.py
prompts file: Path.resolve()+relative_to() → is_under_cwd (project rule
for Windows 8.3 short-name safety).
- soup_cli/commands/runs.py: typed _filter_runs_by_cwd, removed redundant
inner import os.
- soup_cli/commands/deploy.py: confirmation panel now shows --template-dir
path when set, not the unused --template default.
Tests: 4720 → 4756 (+36) across two new files (test_v0402_part_a.py,
test_v0402_part_b.py). 5 review agents (python / code / security / tdd /
verification) all clean after fixes.
Known limitations:
- Custom HF Space templates always create the Space with sdk=gradio
regardless of the supplied app.py. Use --template streamlit-chat with
the inline registry for Streamlit. Tracked for v0.40.3+.
- _resolve_model_source returns ("hf", repo_id) without validate_repo_id;
transformers.from_pretrained will raise loudly on malformed ids.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Part A — BCO Trainer (Binary Classifier Optimization): new task='bco',
training.bco_beta, bco.yaml template, train+sweep routing. Internal
_split_dpo_rows_to_bco adapts paired DPO input to TRL's BCO unpaired
schema; skipped rows logged at DEBUG (mirrors v0.33.0 #47 policy).
Part B — Unified preference dispatcher: additive task='preference' +
training.preference_loss Literal {dpo,simpo,orpo,ipo,bco}. Legacy
task='dpo' / 'simpo' / 'orpo' / 'ipo' / 'bco' remain first-class —
the new surface is purely additive, not a breaking collapse.
_make_inner_cfg uses model_copy so re-validation never sees an
intermediate inconsistent state and the caller's cfg is never mutated.
Part C — KL-controlled DPO variants: dpo_beta_schedule (linear /
cosine / exponential) + dpo_beta_end + dpo_ref_regen_epochs [1, 1000].
BetaScheduleCallback resolves total_steps lazily in on_train_begin
(closes a first-cut bug where total_steps=0 silently emitted beta_end
for every step). RefModelRegenCallback uses load_state_dict(strict=True)
with WARNING-on-mismatch (closes a first-cut silent partial-copy
hazard). Gated to DPO-family tasks only; rejected on mlx backend with
distinct error message.
Part D — Multi-objective preference_loss_weights (2-5 entries, key
allowlist + null-byte rejection, sum-to-1 ±1e-6). Schema-level surface
only; live runtime weighted-loss combination deferred to v0.40.1 with
NotImplementedError stub-then-live (mirrors v0.27.0 MII / v0.37.0
multipack / v0.38.0 quant menu / v0.39.0 ReLoRA pattern).
Net +118 tests (4538 → 4656). All four review-agent waves clean
(Python / Code / Security / TDD).
Known limitation: BCOTrainerWrapper still hardcodes
trust_remote_code=True (carry-over of the v0.36.0 #63 family across
non-SFT trainers).
Also: add docs/ to .gitignore (internal-only docs going forward;
existing docs/QUANTIZATION.md from v0.38.0 stays tracked).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Train-time support for 7 new quantization formats — close the width gap
with LlamaFactory. Wired into SFT trainer + transformers backend + text
modality; multi-trainer/modality expansion deferred to v0.38.1 (mirrors
v0.27.0 MII / v0.37.0 multipack stub-then-live pattern).
- Part A — GPTQ: quantization='gptq' + gptq_disable_exllama (PEFT triton)
- Part B — AWQ: quantization='awq' + GEMM/GEMV builder
- Part C — HQQ: hqq:1bit..hqq:8bit (no 7bit; not supported upstream)
- Part D — AQLM: locked-2-bit
- Part E — EETQ: locked-8-bit
- Part F — MXFP4 + FP8 dequantize-on-load
- Part G — bnb_4bit_quant_storage for FSDP+QLoRA
("crucial for fsdp+qlora" — LlamaFactory quantization.py:178)
- Part H — check_quant_distributed_compat matrix + docs/QUANTIZATION.md.
HQQ/EETQ/AQLM x {FSDP, ZeRO-3} hard-fail; BNB-4bit + FSDP without
quant_storage warns. Wired into commands/train.py startup.
Three new schema validators:
- _validate_prequantized_no_qat — pre-quantized + QAT incompatible
- _validate_bnb_quant_storage_only_with_4bit — silent no-op guard
- _validate_quant_menu_supported_tasks — sft + transformers + text gate
Net: +61 tests (4374 -> 4435). Four review-agent waves clean before tag
(python-review / code-review / security-review / tdd-guide);
verification-loop performed as manual equivalent per CLAUDE.md allowance.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI macOS runners render Rich panel help at a narrower terminal width
than local Windows. The flag --trust-remote-code is split with ANSI
colour escapes between segments, so the literal substring match in
the three CLI plumbing tests failed even though the flag was correct
in --help output. Mirrors the existing _strip_ansi helper in
tests/test_log_level.py (v0.34.0 fix for the same class of issue).
Tests-only follow-up; no soup_cli/ changes; no version bump needed
per release checklist policy on tests-only commits.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Four silent-failure modes Soup had → loud failures, plus a
security default-deny.
- Part A: assistant-only loss masking (default true). Replaces TRL's
multi-turn heuristic with explicit IGNORE_INDEX masking. New
data.train_on_responses_only / train_on_messages_with_train_field
+ per-message train: bool field. Preferred path uses
return_assistant_tokens_mask; fallback uses incremental tokenize
delta with add_special_tokens=False to avoid double-BOS drift.
- Part B: --trust-remote-code opt-in default-deny on soup train /
chat / serve / data download / eval auto. KNOWN_SAFE_PREFIXES
allowlist (15 first-party orgs) suppresses warning panel.
Replaces 9 unconditional trust_remote_code=True call sites in
the SFT path. Non-SFT trainers + diff/export/merge/infer/generate
still hardcode trust_remote_code=True — documented v0.36.x patch.
- Part C: chat-template hardening. Tokenizers without chat_template
raise loudly instead of silent f"{role}: {content}" fallback.
New data.chat_template (registered name or raw Jinja). Filesystem
-touching Jinja directives (include/import/from/macro/extends)
blocked at config-load. Override application warns that soup push
will persist the new Jinja into tokenizer_config.json.
- Part D: OOM-probe auto batch-size. New
training.auto_batch_size_strategy: auto|static|probe. Try-halve
-then-double-to-ceiling loop, max 8 doublings, ceiling = static
× 4. ~/.soup/batch_cache.json (0600 perms, env-override
containment-checked against ~/cwd/tempdir). make_cache_key
rejects bool inputs.
Net +134 tests (4115 → 4249). All 5 review-agent waves clean
before commit; 5 HIGH / 10 MEDIUM / 5 LOW findings fixed in one
review-fix wave.
Smoke: python -m soup_cli.cli version → soup v0.36.0; all 5 new
--trust-remote-code flags surface in --help; ruff clean; pytest
4249 passed / 3 skipped / 0 failed in 2m41s on Windows py3.10.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PR #62 added fp8_recipe support but only wired sft.py directly.
The other 10 trainers (dpo, pretrain, grpo, kto, orpo, simpo, ipo,
ppo, reward_model, embedding) all route through apply_v028_speed_memory,
which was calling apply_fp8_training(model) without recipe -- meaning
user-set fp8_recipe='rowwise' was a silent no-op on every non-SFT task.
- v028_features.apply_v028_speed_memory: read tcfg.fp8_recipe and pass
through to apply_fp8_training; surface the picked recipe in the
green status line so the run record reflects the actual dispatch
- sft.py: drop the defensive getattr (fp8_recipe is a Pydantic field
with a default, not optional) -- use tcfg.fp8_recipe directly
- tests: add TestFP8RecipeViaV028Features (4 tests) verifying the
recipe propagates through apply_v028_speed_memory for tensorwise /
rowwise / rowwise_with_gw_hp, plus the int8-QAT path is unaffected
Add fp8_recipe config field to TrainingConfig with three torchao-backed
scaling recipes: tensorwise (default, v0.28.0 behavior), rowwise (more
accurate via CUTLASS), and rowwise_with_gw_hp (most accurate, grad_weight
in high precision). Dispatches via Float8LinearConfig.from_recipe_name().
- schema.py: add fp8_recipe Literal field with validator requiring
quantization_aware='fp8' for non-default recipes
- fp8.py: update apply_fp8_training() to accept recipe parameter
- sft.py: pass tcfg.fp8_recipe to apply_fp8_training()
- README.md: document recipe options with comparison table
- tests: 24 tests covering schema, dispatch, validation, backward compat
Wires v0.28.0 speed/memory features into every transformer-backend
trainer (grpo / kto / orpo / simpo / ipo / ppo / reward_model /
embedding) plus closes the v0.33.0 #43 oversight where dpo / pretrain
accepted activation_offloading without installing offload hooks.
Auto-quant --auto-quant now forwards the picked candidate's
quantization to vLLM via an explicit named parameter (kwarg-splat
hazard removed). Kernel auto-compose runs a forward-only benchmark
loop on the trainer's actual model under torch.no_grad() so live
training gradients aren't polluted (this was a critical-class bug
caught by code-review pre-tag and fixed before merge).
Schema gate lifted with distinct MLX-backend vs unknown-task error
messages so users get the right fix. fp8 / int8 QAT guard fixed in
6 trainers (the legacy unguarded `if tcfg.quantization_aware:` would
have crashed the int8 path with the string "fp8").
Net +187 tests (3928 -> 4115). New file
tests/test_trainer_coverage_v035.py provides a parametrised matrix
proof that every trainer x every feature is exercised on every CI
matrix job. All four review-agent waves (python / code / security /
tdd) clean with every CRITICAL -> LOW finding fixed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI's narrower terminal width forced Rich to split the flag literal across
ANSI colour escapes (`\x1b[1;36m-\x1b[0m\x1b[1;36m-log\x1b[0m\x1b[1;36m-level\x1b[0m`),
so the contiguous substring `--log-level` was not present in result.output
even though the flag is registered correctly. Strip ANSI codes before the
substring check — same pattern applied to similar Typer/Rich help-text
tests in other Python projects. No code change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds soup why, soup tui, soup runs replay, soup train --profile, .crash
bundle on training exception, per-run cost in SQLite, --log-level global
flag. Net +110 tests (3818 → 3928); all five review-agent waves
(python / code / security / tdd / smoke) clean.
- Part A: --log-level quiet|normal|verbose|debug → Rich-formatted logger
on the "soup" namespace; idempotent + tier-change replaces handler.
- Part B: SQLite gains cost_usd / cost_gpu_label via lazy ALTER TABLE
(race-tolerant against duplicate-column on concurrent first-boot);
rendered in soup runs show / replay / TUI; bool num_gpus rejected;
LIKE wildcards escaped in tracker.get_run prefix match.
- Part C: soup why — heuristic NaN / plateau / divergence / grad-norm /
LR bounds; severity-ordered findings.
- Part D: .crash bundle generator with recursive hf_*/sk-*/Bearer
redaction, output_dir basename-only, os.path.realpath containment,
secrets.token_hex filename, ValueError (not PermissionError) on
outside-cwd; train.py except-handler writes the bundle without
masking the original exception.
- Part E: soup runs replay <id> — summary panel + downsampled loss
curve (≤2000 points) from SQLite history.
- Part F: soup train --profile — torch.profiler Chrome trace to
<output>/profiles/<run_id>.trace.json; run_id rejects '.', '..',
'/', '\\', null bytes; profiles dir created only on torch import.
- Part G: soup tui — Textual dashboard with lazy ExperimentTracker
import; markup_escape on every DB-sourced string; new [tui] extra.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The test patched sys.platform="linux" but called the cached
_get_isolation_strategy() — on macOS / Windows CI the cache had
already been populated with the host's strategy ("sandbox-exec" on
darwin, "best-effort" on win32) before the test ran, so the platform
patch was a no-op.
Use _compute_isolation_strategy() (uncached) like every other test in
the class. Also inject a fake os.unshare via monkeypatch so the
"namespaces" branch is reachable regardless of host kernel.
Tightened the assertion from `in {"namespaces", "best-effort"}` to
`== "namespaces"` since the unshare-unavailable case is covered by
the dedicated test_isolation_strategy_linux_unshare_unavailable test.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses findings from 5-agent review wave (python-reviewer,
code-reviewer, security-reviewer, tdd-guide, smoke-verification).
CRITICAL:
- cans/run.py _deploy_target ollama path: rglob *.gguf result is now
realpath+commonpath checked against extract_dir before forwarding to
`soup deploy ollama --gguf`. Prevents a crafted symlink in the can
from making rglob point at an arbitrary on-disk path.
HIGH:
- cans/publish.py: removed dead update_repo_settings + bare-except tag
block (was a no-op network round-trip). Tag attachment via README
front-matter is documented as a v0.33.x docs follow-up.
- registry/attach.py lookup_entry_by_output_dir: emits ResourceWarning
when the 1000-row scan limit is hit (was a silent miss).
- data/collators.py CrossDocCollator: stops mutating input dicts via
pop() — uses get + dict comprehension. HF Dataset rows are cached and
reused; mutation broke subsequent batches silently. Bare-except now
logs at DEBUG level so production degradation is inspectable.
- monitoring/callback.py _write_spike_recovery_hint: added is_under_cwd
guard. args.output_dir came from raw HF TrainingArguments without
separate path-containment check.
- trainer/rewards.py MACOS_SANDBOX_PROFILE: narrowed (allow mach-lookup)
to a 3-name allowlist (SecurityServer, notification_center,
opendirectoryd.libinfo). Broad mach-lookup permitted DNS / NSURLSession
via launchd, defeating (deny network*).
- cans/run.py: PermissionError → ValueError so a caller wrapping in
`except OSError` cannot silently swallow the consent gate.
PermissionError is an OSError subclass.
- commands/can.py run_cmd: assigns result=None up front + explicit None
guard so a future _fail bypass cannot trigger NameError on result.
- utils/v028_features.py: added type annotations on apply_v028_speed_memory
(model: Any, tcfg: TrainingConfig via TYPE_CHECKING, console: Console)
and warn_unsupported_features.
- cans/run.py: confirm_callback now annotated
Callable[[Manifest], bool] for IDE introspection.
- tests/test_part_b.py reexec test: drops env-var contamination
(RANK/WORLD_SIZE/LOCAL_RANK/ACCELERATE_*) before run, patches
imported names on train module, and forces assertion that
os.execvp was called — no more silent skip-on-bypass.
- tests/test_part_d.py: added TestGenerateResponseSignature
source-level guard that catches the lenient logits_processor mock
silently passing.
MEDIUM:
- cans/run.py _run_subprocess: catches subprocess.TimeoutExpired and
returns rc=124 (coreutils convention) so callers see a clean
CanRunResult instead of an unhandled traceback after the 24h cap.
- cans/run.py: temp dir created via mkdtemp is now cleaned up on
extract_can failure (try/except + cleanup_extract_dir).
- cans/run.py cleanup_extract_dir: switched startswith path check to
os.path.commonpath (project-standard idiom; Windows-safe).
- cans/schema.py DeployTarget._safe_relpath: normalises mixed
separators before splitting on '/' so foo/..\bar can no longer
bypass the .. check.
- utils/lr_finder.py run_lr_sweep: removed redundant local
`import math as _math` (math already at module level).
LOW:
- eval/gate.py _parse_judge_url: removed bare http:// catchall after
scheme allowlist. Defence-in-depth for callers that bypass the
Pydantic GateTask validator.
- utils/auto_quant.py evaluate_candidate: latency mean now divides by
*completed* prompts (excludes crashed). Crashed candidate no longer
appears artificially fast.
- utils/auto_quant.py Candidate.__post_init__: explicitly rejects bool
in score / latency_ms (bool is a subclass of int, was sneaking past).
- utils/mii.py: removed `noqa: F401` on Optional import (now actually
used in type annotation since we restored it).
Tests added (+7, total 3811→3818):
- test_part_a_wave1: attach_artifact outside-cwd rejection.
- test_part_a_wave2: PermissionError→ValueError migration in 2 tests.
- test_part_c: CrossDocCollator mismatched doc_lengths fallback,
does-not-mutate-input-dict regression guard.
- test_part_d: source-level _generate_response signature guard.
- test_part_e: should_recover at max_attempts, outside-cwd skip.
Lint: clean. Full suite: 3818 passed in 156s.
Findings deliberately not actioned (with rationale):
- code-review M1 (mii Pydantic at import-time): forward-ref resolution
requires module-level definitions for FastAPI; documented in mii.py.
- code-review M4 (supports_v028_features vs validator divergence):
the v0.33.0 schema validator was renamed to
_validate_v028_speed_memory_supported_tasks and now imports
supports_v028_features — they cannot drift.
- python-review LOW (_deploy_target vllm silent no-op): documented in
the docstring as advisory; logging requires a console arg the
helper does not currently take.
- security-review LOW 8/9 (TOCTOU window, CLONE_NEWPID): theoretical;
documented in CLAUDE.md security section in the next commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#37, #38. Final Part of v0.33.0 implementation phase.
#37 Auto-reexec under accelerate launch when --gpus N>1:
- soup_cli/commands/train.py gains --no-reexec opt-out flag (default
behaviour: auto-reexec).
- When --gpus N>1 and not already in a distributed env (RANK/WORLD_SIZE
+ ACCELERATE_* markers absent), train() reconstructs argv via
utils.launcher.build_accelerate_argv and calls os.execvp("accelerate",
argv). os.execvp replaces the current process — no leftover PID tree,
stdio passes through unchanged.
- Critical flags (--fsdp, --deepspeed, --resume, --wandb, --tensorboard,
--yes) are forwarded to the reexec'd run so users see the same
behaviour they'd get from running accelerate launch by hand.
- OSError from execvp falls back to the v0.27.0 advisory (printed
command) so misconfigured PATH doesn't dead-lock the user.
- --no-reexec preserves the v0.27.0 print-and-exit behaviour for users
who want to control env vars / stdio explicitly.
#38 DeepSpeed-MII live serve:
- soup_cli/utils/mii.py gains build_mii_app(pipeline, model_name) which
returns a FastAPI app with /v1/chat/completions + /v1/models matching
the v0.30.0 transformers backend's contract.
- Pipeline is held by closure (single MII instance, thread-safe across
concurrent generations). Loopback-only CORS mirrors v0.30.0
transformers backend policy.
- max_tokens bounds [1, 16384], stream=True rejected (MII v0.x lacks
stable streaming), pipeline crashes return 500 with generic message
(no stack-trace leak). Empty response → 500.
- soup_cli/commands/serve.py replaces the v0.27.0 stub-warning + Exit(1)
with create_mii_pipeline → build_mii_app → uvicorn.run.
Tests: +9 in tests/test_part_b.py covering /v1/models endpoint, chat
happy-path with mocked pipeline returning .generated_text, streaming
rejection, max_tokens bounds (low + high), pipeline failure → 500,
empty pipeline response → 500, --no-reexec parameter exists,
--no-reexec advisory fallback, --gpus 2 reexec calls os.execvp with
accelerate argv (via monkeypatched os.execvp).
Known limitations:
- MII server has no streaming, no LoRA hot-swap, no /metrics dashboard,
no OpenTelemetry — those are v0.30.0 transformers-backend features
not yet ported. Documented in the build_mii_app docstring.
- Auto-reexec assumes accelerate is on PATH; OSError path prints the
command instead, matching the v0.27.0 baseline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#43, #44, #47.
#43 Multi-trainer wiring (sft/dpo/pretrain):
- New utils/v028_features.apply_v028_speed_memory(model, tcfg, base_model,
console) — single shared helper for use_cut_ce, quantization_aware="fp8",
kernel_auto_compose. Each feature degrades silently to a yellow advisory
if the underlying lib is missing; never crashes training kick-off.
- Helpers supports_v028_features(task) and warn_unsupported_features(tcfg, task)
drive both the schema validator and runtime advisories.
- soup_cli/trainer/dpo.py and trainer/pretrain.py now call the helper after
model load (post-LoRA, post-QAT) — same hook point as SFT.
- soup_cli/config/schema.py validator
_validate_v028_speed_memory_sft_only renamed
_validate_v028_speed_memory_supported_tasks; allowlist now {sft, dpo,
pretrain}. GRPO/KTO/ORPO/SimPO/IPO/PPO/RewardModel/Embedding still error
out at config-load with a precise multi-trainer message.
#44 Selective gradient-checkpoint hooks:
- New utils/gradient_ckpt.install_selective_hooks(model, granularity)
iterates ``model.named_modules()`` looking for transformer-block-shaped
names (numeric suffix on layer path), wraps each module's ``forward``
with torch.utils.checkpoint.checkpoint based on tier:
- selective: only attention sub-modules
- medium: every second transformer block
- full: every transformer block
- Returns hook count so callers can fall back to HF native checkpointing
when zero blocks were found.
#47 CrossDocCollator:
- New soup_cli/data/collators.CrossDocCollator wraps any base data
collator and injects a block-diagonal causal ``cross_doc_attn_mask``
built from per-example ``doc_lengths``. Preferred over TRL's
``packing_strategy="attention_free"`` flag (best-effort across TRL
versions). Degrades gracefully when doc_lengths is missing or shapes
don't match — base attention_mask preserved, no crash.
Tests: +16 in tests/test_part_c.py covering apply_v028_speed_memory
(no-features, cut_ce graceful failure), supports/warn helpers extension,
schema gate (dpo + pretrain accept, kto still rejects), selective hook
installation across full/medium/selective with fake transformer-shaped
models, CrossDocCollator passthrough + strip + injection. One existing
test in test_training_speed.py updated: dpo+use_cut_ce now accepted.
Known limitations:
- 7 trainers (GRPO/KTO/ORPO/SimPO/IPO/PPO/RewardModel/Embedding) still
reject v0.28.0 flags at config-load. Each is a 5-line addition once
schema validation is satisfied; tracked as a v0.33.x follow-up.
- install_selective_hooks doesn't undo earlier hooks — caller must be
re-init aware. Not an issue for the typical "construct wrapper, train,
exit" flow but worth noting.
- CrossDocCollator emits ``cross_doc_attn_mask`` (not ``attention_mask``)
to avoid clobbering the base collator's contract; downstream consumers
must read the new key explicitly. The plan calls for "preferred over
TRL's packing_strategy" which we satisfy via opt-in collation, not
silent override.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#49, #53, #54.
#53 Wire --structured-output into transformers generation loop:
- New utils/structured_output.build_logits_processors(constraint, tok)
returns a HF LogitsProcessor list. Tries outlines first (broader
coverage), falls back to lm-format-enforcer, returns [] if neither
installed or factory crashes — server degrades to free-form rather
than 500 on a missing dep.
- _generate_response gains logits_processor kwarg, forwarded to
model.generate(...). Chat-completions handler builds the processor
list per request (cheap; per-request build keeps the descriptor
mutable for future /v1/output_constraint endpoints) and passes it
down. Empty list path is unchanged from v0.30.0 free-form behaviour.
#54 --auto-quant live eval loop:
- New utils/auto_quant.evaluate_candidate(name, eval_fn, prompts):
times mean per-prompt latency, scores correctness, marks ok=False
when any prompt crashes or score < min_correct_fraction.
- New utils/auto_quant.run_auto_quant_picker(candidate_specs, prompts,
min_score): evaluates each candidate, calls pick_best, soft-falls-
back to highest-scored ok candidate if no candidate clears the
threshold so the server still binds.
- serve.py replaces the v0.30.0 deferral warning with a real picker
run over a fixed 3-prompt set across default_candidate_order().
Logs the picked (name, score, latency) on stdout.
#49 End-to-end --push-as integration test (mocked HF):
- New tests/test_part_d.py::TestPushAsResumeIntegration uses a fake
huggingface_hub module via patch.dict to verify HFPushCallback
constructs cleanly with a token, exposes the _repo_failed sticky
flag (v0.29.0 review fix), and that prepare_hf_resume rejects
output_dir outside cwd. The full HF Hub network roundtrip needs a
paid sandbox repo — keeping it mocked-only is a deliberate trade
(prevents flaky CI on rate limits / token rotation).
Tests: +15 in tests/test_part_d.py covering build_logits_processors
graceful-degrade paths (None / off / unknown / no-libs / factory
crash), generate_response logits_processor plumbing, evaluate_candidate
(empty / all-correct / crash / below-threshold), run_auto_quant_picker
(threshold pass + soft fallback), HF push smoke. One existing test in
test_inference_advanced.py updated: TestAutoQuantCLIWarning no longer
expects the v0.30.1 deferral message — it now expects --auto-quant to
actually run.
Known limitations:
- #49: full HF Hub roundtrip is mocked-only; live integration test
requires a paid sandbox repo and rotating token, deferred to a
separate end-to-end CI job.
- #54: live re-loading of the model at the picked quant is NOT done
in this commit — the picker logs the choice but the already-loaded
model is served. Live re-load needs an additional bnb / awq round-
trip per candidate, which is heavy for a startup-time decision;
follow-up tracked.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#56, #57, #58, #59.
#56 Live --find-lr in-process LR-sweep:
- New utils/lr_finder.run_lr_sweep(model, dataloader, schedule,
optimizer_factory, device): per-step LR mutation + forward + backward,
records loss until exhaustion or NaN/Inf divergence.
- commands/train.py wires it via _live_lr_sweep_from_config (loads model
+ tokenizer + first N rows of cfg.data.train), with synthetic-curve
fallback so users without GPU/torch still get a parseable report.
#57 Loss-spike recovery hint:
- SoupTrainerCallback gains spike_recovery / spike_recovery_max_attempts
/ spike_recovery_lr_decay; on watchdog fire writes
output_dir/spike_recovery.json with previous_lr, recommended_lr (per
SpikeRecoveryStrategy.compute_new_lr), should_recover, attempts. A
wrapper / re-launch can resume with the decayed LR. Live optimizer
rewind is intentionally NOT done — HF Trainer has no safe public API
for mid-loop optimizer-state mutation; the JSON hint is the contract.
#58 auto_mixed_precision push to TrainingArguments:
- New SFTTrainerWrapper._resolve_mixed_precision: when
tcfg.auto_mixed_precision is True, queries torch.cuda compute
capability and calls pick_mixed_precision(base, cc) to set
bf16=/fp16= flags. CPU short-circuits to (False, False). When the
flag is False, legacy default preserved (bf16=cuda).
#59 Grad-accum advisory (Phase 1):
- SoupTrainerCallback gains grad_accum_auto_tune /
grad_accum_pressure_threshold / grad_accum_total_vram_gb /
grad_accum_current_steps / grad_accum_current_batch.
- on_log samples torch.cuda.max_memory_allocated each step; if
GradAccumMonitor.should_adjust crosses the threshold once,
prints (batch, accum) -> (new_batch, new_accum) advisory and
short-circuits (one-shot). Phase 2 (live DataLoader rebuild)
needs a small TRL upstream PR — tracked as a known limitation.
Wiring:
- soup_cli/trainer/sft.py: _resolve_mixed_precision helper, batch_size
preserved on self, SoupTrainerCallback constructor passes through new
spike + grad-accum knobs.
- soup_cli/monitoring/callback.py: rich Console import added (was
previously module-relative); spike + grad-accum state fields and
one-shot helpers.
Tests: +15 in tests/test_part_e.py covering the LR-sweep loop with
mocked model + optimizer (records, divergence break), mixed-precision
resolver across cpu/cuda + auto-flag combinations + qwen2 fp16 quirk on
Ampere, spike recovery hint write + attempts increment + disabled
no-op, grad-accum advisory one-shot semantics + threshold + cuda-absent
+ disabled.
Known limitations (release notes):
- #57 spike recovery is a JSON hint, not in-process optimizer rewind
- #59 Phase 2 (live DataLoader rebuild on advisory) deferred
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#34. Completes Part A — all of #32, #34, #35 now shipped.
Schema bump: CAN_FORMAT_VERSION 1 -> 2 (additive).
- SUPPORTED_CAN_FORMAT_VERSIONS = (1, 2): old cans still inspect/extract.
- New DeployTarget Pydantic model with kind in {ollama, gguf, vllm}, name
validation (no null bytes / newlines), path validation (relative-only,
rejects '..' and absolute paths).
- Manifest gains optional deploy_targets: list[DeployTarget] field.
cans/run.py — orchestrator:
- run_can(can_path, yes, deploy, extract_dir, capture_env_to, ...)
validates path containment, requires --yes or explicit confirm_callback
(security: auto-downloads data + auto-trains), extracts, optionally
captures env, then invokes `soup train --config ... --yes` via
subprocess so the trainer dispatch stays a single source of truth.
- capture_env: best-effort pip freeze + python version + GPU detection,
never blocks training on env-capture failure.
- _deploy_target dispatches per-kind; ollama path runs `soup deploy
ollama --gguf ... --name ...` if a *.gguf is present in the can.
- cleanup_extract_dir: tmp-or-cwd-only safety guard around shutil.rmtree.
cans/publish.py — HF Hub publish:
- publish_can(can_path, repo_id, token, private, commit_message)
validates can-path containment, repo_id via utils/hf.validate_repo_id,
resolves token via utils/hf.resolve_token (env > cache files), uploads
to repo_type='dataset' with commit-message first-line + 200-char cap
(matches v0.29.0 push.py / data push policy). Tags as can-format-v1.
CLI: soup_cli/commands/can.py
- New `soup can run <path> [--yes] [--deploy] [--extract-dir]
[--env-capture]` — confirmation panel mandatory without --yes.
- New `soup can publish <path> --hf-hub <user/repo> [--private]
[--message]`.
Tests: +28 in tests/test_part_a_wave2.py covering schema bump (v1/v2/v3),
DeployTarget validation (path traversal, null bytes, kind enum),
capture_env (smoke + pip-failure tolerance), run_can (containment +
confirmation gate + train-argv shape via mocked subprocess), publish_can
(repo_id validation, token resolution, commit-message sanitization, HF
upload via mocked HfApi), and CLI smoke (confirmation panel, missing
file). One existing test_cans test relaxed (v1 == v1 -> v in {1,2}).
Known follow-ups (not blocking release):
- soup can run does NOT yet auto-fetch data_ref.kind=hf|url. Embedded
config must reference a local data path. Filed mentally as
v0.33.x follow-up.
- registry_snapshot.json lineage export deferred — pack already embeds
base_hash which is enough to query the source registry post-extract.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#32, #35. (#34 soup can run/publish deferred to Part A wave 2.)
#32 Live model scoring for `soup eval gate` + `soup eval quant-check`:
- gate.run_gate now dispatches judge / benchmark / custom task types,
wrapping each scorer in try/except so a backend failure produces
score=None + error=str(exc) instead of a silent score=1.0 pass.
- New _parse_judge_url splits ollama:// / http(s):// judge_model URLs
into (provider, model, api_base) for JudgeEvaluator.
- New _run_judge_task / _run_benchmark_task plug into existing
eval/judge.py and eval/forgetting.py runners.
- New quant_check.make_model_generator(model_path) wraps transformers
AutoTokenizer + AutoModelForCausalLM into a generate_fn callable;
greedy by default for reproducible scores; lazy-imported.
- gate_cmd / quant_check_cmd build live generators when --model is
given; fall back to deterministic stub on load failure so CI without
GPUs still runs the orchestration layer.
- GateTaskResult.score is now Optional[float] with new error: Optional[str].
- _print_gate_result renders ERROR + reason cleanly.
#35 Registry attach hooks:
- registry/store.py _VALID_KINDS extended with eval_results, tensorrt.
- New registry/attach.py: attach_artifact, write_eval_json
(cwd-containment via realpath+commonpath), lookup_entry_by_output_dir.
- `soup eval custom` gains --attach-to-registry + --output (paired);
on success writes JSON results and adds eval_results artifact row.
- `soup export` gains --registry-id with auto-match by source --model
output dir; auto-attaches the produced GGUF artifact. Failures here
are warnings, not hard exits — export already succeeded.
Tests: +19 in tests/test_part_a_wave1.py covering URL parser, error
propagation across all 3 task types, score=None semantics, generator
factory bounds + transformers mocking, registry attach helpers
(containment + missing entry), and CLI integration.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#21, #22.
#21 RLVR code_exec_reward: add OS-level isolation strategy detection.
- New _get_isolation_strategy / _compute_isolation_strategy with cache.
- Linux: best-effort os.unshare(CLONE_NEWUSER|CLONE_NEWNET|CLONE_NEWPID)
in preexec_fn (Python 3.12+). Silent fallback on EPERM/ENOSYS for
hosts where unprivileged user namespaces are disabled.
- macOS: prefix subprocess argv with sandbox-exec + inline default-deny
profile (deny network*, deny writes outside /tmp).
- Windows + restricted Linux: existing RLIMIT + socket-patch + ephemeral
cwd guards continue to apply (best-effort baseline).
#22 prune_checkpoints: TOCTOU-safe symlink handling.
- Top-level entries: explicit os.lstat + stat.S_ISLNK check (intent-clear)
instead of Path.is_symlink.
- shutil.rmtree now passes onerror=_abort_on_symlink to abort recursive
walk if any symlink is encountered mid-walk (defence-in-depth).
- OSError mid-prune is caught per-checkpoint so one bad dir does not
abort the whole prune pass.
Tests: +13 in tests/test_part_f_hardening.py covering strategy detection
on linux/darwin/win32, sandbox profile shape, code_exec smoke tests, and
TOCTOU-resistant prune behaviour.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Rich/Typer emits per-character ANSI escapes in CI on Linux runners
("--\x1b[m-find\x1b[m-lr") so the raw substring assertion `--find-lr`
in result.output fails on ubuntu-latest x py3.12 even though it
passes on Windows where Rich auto-disables colour.
Strip ANSI before comparing — same pattern already used by
test_hf_integration.py and test_eval_platform.py.
Tests-only commit, no version bump.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Expand the recipe catalog from 46 to 80 entries — every popular open-weight
model family now has a validated Soup recipe.
Part A — Vision (6 recipes): Llama-3.2-Vision-90B, Pixtral-12B, Qwen2-VL
(7B + 72B), InternVL 2.5, MiniCPM-V 2.6
Part B — Audio (3 recipes): Qwen2-Audio, SeamlessM4T v2, Whisper-large-v3
Part C — Reasoning (7 recipes): completes the 6 DeepSeek-R1-Distill sizes,
plus Qwen3-Coder, Qwen3-30B-A3B reasoning, Phi-4 reasoning
Part D — Edge (8 recipes): SmolLM2 (135M / 360M / 1.7B), Qwen2.5
(0.5B / 1.5B / 3B), Gemma 2 2B, Phi-3.5-mini
Part E — Domain (8 recipes): BioMistral, Meditron, CodeLlama (13B / 70B),
Magicoder, Mathstral, Nemotron-4 340B, Llama-2-13b-finance
Part F — Multimodal reasoning (2 recipes): Llama-3.2-Vision GRPO, Pixtral DPO
Part G — Recipe-validation CI workflow on every PR touching recipe / config /
data code (.github/workflows/recipe-validation.yml)
Part H — 750 parametrized tests covering catalog-wide invariants:
model-id safety (no `..`/`://`/null bytes), lora.target_modules non-empty,
max_length within schema bounds, GRPO recipes wire reward_fn +
num_generations >= 2, vision recipes set image_dir, audio recipes set
audio_dir, default data path is non-empty + relative
Live 100-step per-recipe smoke train (requires GPU runner) deferred to v0.31.1.
Tests: 2886 → 3607 (+721). Catalog: 46 → 80 (target met).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The ubuntu-latest py3.11 matrix cell doesn't install the [serve] extra,
so `soup serve` exits early with a FastAPI-missing message before
reaching --structured-output / --auto-quant / --json-schema validation.
Add `pytest.importorskip("fastapi")` to the three tests that exercise
those CLI-level validation paths.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CI failures on macOS/Windows:
1. `test_endpoint_rejects_null_byte` — ``monkeypatch.setenv("HF_ENDPOINT",
"...\x00")`` raises ``ValueError: embedded null byte`` at the C-level
setenv call on macOS/Windows before ``resolve_endpoint`` can reject it.
Linux's setenv swallows it. Replace with ``monkeypatch.setattr`` on
``os.environ`` dict so the null-byte string reaches ``resolve_endpoint``
on every platform.
2. Help-text substring tests (``test_train_shows_push_as_flag_in_help``,
``test_push_shows_collection_flag_in_help``, ``test_push_subcommand_exists``,
``test_hf_space_help_shows_flags``, ``test_train_help_shows_hf_resume``,
``test_hf_space_command_registered``) — Typer injects ANSI escape codes
on macOS/Windows pytest runs, splitting tokens like ``--push-as`` into
``-`` + ``-push-as`` across escape groups. Add ``_plain()`` helper that
strips ANSI via regex and use it in every help-text assertion.
Same pattern as 899ad8e (test_eval_gate.py) did after the v0.26.0 CI
Windows failure.
Full suite still passes locally: 100 HF integration tests in 3.25s.
- Split the 22-assert config-values test into 3 focused tests
(task+data, training hyperparams, LoRA config) so a deliberate
example change surfaces in one targeted test, not a wall of asserts
- Add module docstring explaining why these tests lock the example state
- Add `from __future__ import annotations` (defensive; matches 14 other
test modules in the project)
- Rename `f` -> `fh` in _load_jsonl to avoid shadowing short name
- Drop asserts on secondary fields (warmup_ratio, weight_decay, scheduler,
logging_steps, etc.) -- they're tweakable knobs, not the example's
teaching points; test brittleness > coverage here
Add a working DPO (Direct Preference Optimization) example using the
current Pydantic config schema with Llama 3.1 8B Instruct and QLoRA.
- examples/configs/dpo_example.yaml: DPO config with all core training
and LoRA parameters, plus commented-out advanced options
- examples/data/dpo_sample.jsonl: 8 preference pairs in DPO format
with ShareGPT-style message lists for chosen/rejected
- tests/test_dpo_example.py: 7 tests validating config loading, field
values, data format detection, and data validation
- examples/README.md: document the new DPO with QLoRA example
- Narrow 'except Exception: pass' in _get_dataset_size to specific
exceptions (OSError, ValueError, KeyError, ImportError)
- _get_dataset_size returns (size, is_estimated) so the caller can
warn when falling back to the 10k default (silent fallbacks are
misleading on a $-estimating command)
- Add -> None return type annotation on cost() (project convention)
- Add variance disclaimer: 'estimates are approximate; +/- 30%'
- Document pricing cadence in GPU_PRICING comment (last updated 2026-04)
- Use highlight=False on json.dumps output
- Fix misleading 'mock data' test comment (there is no mock)
- Add 2 tests: dataset-unreadable warning, variance disclaimer rendering
* feat(cli): implement 'soup cost' command to estimate cloud GPU training costs
* feat(cli): implement 'soup cost' command to estimate cloud GPU training costs
* test(cost): add unit tests for 'soup cost' command and output formatting
* docs(readme): add usage documentation for the new 'soup cost' command
`test_train_gate_flag_accepted` asserted `"--gate" in result.output`, but
Typer/Click under CI emits ANSI color codes that split the flag name into
non-contiguous chars: `\x1b[1;36m-\x1b[0m\x1b[1;36m-gate\x1b[0m`. The literal
"--gate" substring is never present. All 9 OS × Python combos failed on the
v0.26.0 Parts B-E push.
Fix: strip ANSI via regex before checking. Also assert on "eval-gated" from
the option description to double-check the flag is wired to its help text.
CI-only / tests-only: no soup_cli/ changes, no version bump needed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Foundation of v0.26.0 "Red and Blue Ocean" — every fine-tune is now
tracked with lineage, config, eval baseline, and shippable artifacts.
New module soup_cli/registry/:
- hashing.py: deterministic SHA-256 of config (canonical JSON) + data
(streamed) + base model; used as the entry_hash identity
- store.py: SQLite store (~/.soup/registry.db) with registry_entries,
registry_artifacts, registry_lineage, registry_tags. Context-manager
API, cycle-safe BFS walks, AmbiguousRefError on prefix collision,
LIKE-wildcard-escaped search + resolve, FK ON DELETE CASCADE.
- diff.py: flat-walk ConfigChange diff + per-benchmark eval delta.
New CLI commands:
- soup registry push/list/show/search/diff/promote/delete
- soup history <name> — lineage DAG tree viewer
Security hardening (v0.26.0):
- name/tag validation: alphanumeric + _-. only, null-byte rejected,
name ≤128, tag ≤64
- artifact path containment via os.path.realpath + commonpath
(Windows 8.3 short-name safe); enforce_cwd=True default
- SQL parameterised; LIKE wildcards %/_ escaped with ESCAPE '\'
- DB 600 perms on POSIX; SOUP_REGISTRY_DB_PATH env override
- indirect-cycle detection in add_lineage via BFS ancestor walk
- Rich markup escaped in all CLI output
- resolve() raises AmbiguousRefError instead of silent None
Tests: 92 new tests in tests/test_registry.py (hashing, validation,
CRUD, artifacts, lineage + cycle, diff, CLI, history, security,
auto-register integration with ExperimentTracker). Full suite:
2409 passed (was 2313).
All review findings addressed (4 agents: python, code, security, tdd):
HIGH: context manager + try/finally cleanup, FK cascade (removed
manual cascade), cycle detection, LIKE wildcard escaping.
MEDIUM: ambiguous resolve raises, exit 0 on user cancel, cwd
captured at construction, enforce_cwd=True default, Windows
ASCII-safe error messages.
Deferred to v0.26.1: soup eval --attach-to-registry flag and
soup export auto-artifact registration.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- happy_path: assert mocked VRAM value (4.00 GB) renders in table
- happy_path: assert 'Benchmarking Configuration' panel rendered
- happy_path: assert mock_generate call count (1 warmup + 3 prompts)
- cpu_warning: assert 'N/A' appears in VRAM column when no CUDA
- Improve docstrings to describe what each test verifies
- Use full exception repr in exit_code asserts for CI debugging
- Narrow broad 'except Exception' to specific exceptions
(OSError, UnicodeDecodeError, json.JSONDecodeError) + `raise ... from`
- Rename file handle `f` -> `fh` to avoid shadowing (ruff-friendly)
- Clarify comment on --num-prompts ignored-when-file semantics
- Strengthen test: assert actual prompts were passed to _generate
(not just exit code + output substring)
- Add PEP 8 second blank line between test functions
* feat(bench): add --prompts-file option with path traversal security
* test(bench): add unit tests for custom prompts and path traversal
* docs(bench): document --prompts-file usage in README.md
* feat(bench): add --prompts-file support with path validation
* test(bench): add unit tests for custom prompts and security checks
* style: remove trailing whitespace to pass ruff linting
* test: fix mock patch targets for local imports in bench command
* refactor(bench): simplify prompts-file logic and clean up comment
* test(bench): update assertions to match new prompts-file semantics
* feat(cli): create 'soup bench' command for inference speed and VRAM measurement
* register 'bench' command into the main CLI router
* add test case for handling missing model paths gracefully
* add 'Inference Benchmarking' section explaining the 'soup bench' tool
* Added soup.yaml
* style: fix linting (unused imports, inconsistent spacing)
* style: sort imports in bench and test_bench to satisfy ruff
* style: final import sort and grouping fix for CI
* Update gitignore
test_writes_config fails on windows-latest / Python 3.9 with exit code 1
because the path-traversal check in soup_cli/commands/autopilot.py was:
data_path = Path(data).resolve()
data_path.relative_to(Path.cwd().resolve())
On Windows + Python 3.9, Path.resolve() occasionally leaves 8.3 short
names (e.g. "C:\Users\RUNNER~1") in one of the two sides but not the
other, so relative_to raises ValueError even when both paths point to
the same location. GitHub Actions runner home dirs frequently trigger
this (the runneradmin account is created as "runneradmin" but short
names get generated as "RUNNER~1").
Fix: introduce _is_under_cwd(path) helper in soup_cli/commands/autopilot.py
that uses os.path.realpath on both sides (handles 8.3 expansion
consistently) plus os.path.commonpath for the containment check, with
case-insensitive comparison on NT. Apply it to both the --data and
--output path guards. The data_path / output_path locals are then
rebuilt from the realpath result so downstream logic sees the
canonical long-name path.
Also enriches the test assertion to print result.output and
result.exception on failure so future CI breaks are easier to diagnose
without needing to push a debug commit first.
Local verification: all 38 tests in tests/test_autopilot.py pass on
Python 3.10 Windows, full suite 2313 passed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two tests in tests/test_bugfixes.py::TestGRPOCPUMinNewTokens fail on
windows-latest / Python 3.11 when importing trl.trainer.grpo_trainer:
RuntimeError: Failed to import trl.trainer.grpo_trainer because of
the following error:
'charmap' codec can't decode byte 0x90 in position 6555: character
maps to <undefined>
Root cause: upstream trl reads an auxiliary file without an explicit
encoding, so Python uses the system default. On Windows that is cp1252
('charmap'), which chokes on non-ASCII bytes present in the file. This
is an upstream issue but Soup needs a green CI.
Two-layer fix:
1. .github/workflows/ci.yml — set PYTHONUTF8=1 and PYTHONIOENCODING=utf-8
as job-level env. Python's UTF-8 mode makes all file I/O default to
UTF-8 regardless of locale, which is the correct global fix for this
class of bug.
2. tests/test_bugfixes.py — add a _trl_grpo_importable() helper that
returns False on UnicodeDecodeError / ImportError / RuntimeError, and
use it as a belt-and-braces skip in both TestGRPOCPUMinNewTokens
tests. Ensures the tests skip cleanly instead of erroring out if a
future CI change accidentally drops PYTHONUTF8.
Local verification: both tests pass with 'pytest tests/test_bugfixes.py::
TestGRPOCPUMinNewTokens -v' (Python 3.10, Windows).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Ships v0.25.0 with eight new capabilities (Parts A–H) that close every
competitive gap vs LLaMA-Factory/Axolotl/Unsloth and add unique differentiators:
Part A — 9 new model recipes: Llama 4 Scout (sft/dpo/grpo), Qwen 3 14B/32B/8B-grpo,
Gemma 3 12B/27B-dpo, DeepSeek V3 (MoE LoRA).
Part B — Tool-calling / agentic fine-tuning: new "tool-calling" data format with
detection + normalization, synth data template, init template, eval scoring
(tool_call_match / tool_call_name_match / tool_call_args_subset), plus
qwen3-8b-tools and llama4-scout-tools recipes.
Part C — RLVR (RL from Verifiable Rewards): reward_fn=verifiable routing to
math_verify_reward (regex-only, no eval), code_exec_reward (subprocess sandbox
with RLIMIT_AS/RLIMIT_CPU on POSIX, ephemeral tempdir cwd, concurrency cap,
one-time warning panel), and json_schema_reward. verifiable_domain Literal
validated via model_validator.
Part D — VeRA + OLoRA PEFT methods: LoraConfig.use_vera / use_olora with
mutual-exclusion validator and a unified peft_builder helper that returns
either LoraConfig or VeraConfig with the right init kwargs.
Part E — Apple Silicon MLX backend: detection + hardware profiling in utils/mlx,
MLXSFTTrainerWrapper via mlx-lm, scaffolding DPO/GRPO wrappers rejected at
config load time by SoupConfig._validate_mlx_task_support, lazy trainer
registry, doctor integration, 3 MLX SFT recipes, [mlx] extra in pyproject.
Part F — Data augmentation: soup data augment with rephrase / translate / style
strategies, path-traversal-protected input/output, count capped 1-10, lang/styles
lists bounded (10 entries × 32 chars), rate limiting, and optional --dedup.
Part G — Training intelligence: forgetting detection (ForgettingDetector with
3 built-in mini benchmarks and warning levels) and checkpoint intelligence
(CheckpointTracker with composite metric, early-stop on regression, safe
top-N pruning refusing symlinks and non-checkpoint dirs). SQLite schema
extended with checkpoint_quality + forgetting_eval tables.
Part H — Autopilot: soup autopilot command with dataset/model/hardware
profilers, decision engine (task/quant/peft/batch/lr/epochs/max_length/perf
flags), YAML generator, and full CLI with dry-run + --yes + path-traversal
protection + goal whitelist + gpu_budget bounds [1GB, 1TB]. Bakes forgetting
detection + checkpoint intelligence + early-stop into the generated config.
Totals:
- 2313 tests passing (183 new, up from 2130)
- 86 test files (8 new)
- 43 ready-made recipes (14 new)
- 16 built-in templates (tool-calling added)
- Review findings: all CRITICAL/HIGH/MEDIUM/LOW addressed (3 documented
design limitations: code_exec best-effort sandbox, prune_checkpoints TOCTOU,
MLX training integration test requires real hardware)
Docs: CLAUDE.md, README.md, SECURITY.md, CONTRIBUTING.md updated.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(cli): add 'soup runs clean' intelligent checkpoint cleanup to reclaim disk space
* update README
* feat(cli): add 'soup runs clean' intelligent checkpoint cleanup to reclaim disk space
* fixed whitespace trails
* style(cli): fix lints (line length and spacing) in runs.py
* style: fix all E501 line length lint errors
* fix test mismatch, improve deletion warnings, add path validation, and enforce argument exclusivity
* fix: break long message into multiple lines for Ruff compliance
* test: update runs clean test to use CWD-based output directory for security compliance
* feat(doctor): add RAM and disk space checks to soup doctor command with tests and updated docs
* fix(doctor): resolve subprocess type checker error by manually validating macOS RAM query return code
* add --json flag to version command for machine-readable output in CI/scripts and include tests
* docs: update README with soup version --json flag examples
- Replace Windows-only paths (C:/Windows/...) with tempdir-based
paths that work on Linux/macOS CI runners
- Use monkeypatch.chdir for path traversal test isolation
- Fixes test_path_outside_cwd_raises failure on Ubuntu CI
- Replace non-ASCII symbols (checkmarks, arrows, bullets, em-dashes)
with ASCII equivalents in Rich console output to prevent
UnicodeEncodeError on Windows without PYTHONIOENCODING=utf-8
- Add _validate_output_path() for AWQ/GPTQ export — output path
traversal is now checked before import check (previously unreachable
when autoawq/auto-gptq not installed)
- 4 new tests for output path validation (2065 total, 0 failures)
- Update SECURITY.md with v0.22.0–v0.24.1 hardening history
Part A: HuggingFace Dataset browser
- soup data search: search HF Hub for datasets (sort by downloads/likes)
- soup data preview: preview remote dataset metadata, splits, features
- soup data download: stream HF dataset to local JSONL (with format conversion)
- Security: trust_remote_code=False, path traversal protection, samples cap at 1M
Part B: Freeze training (like LLaMA-Factory finetuning_type: freeze)
- freeze_layers / freeze_ratio config fields
- soup_cli/utils/freeze.py: detect layers, freeze bottom N
- Wired into SFT trainer before LoRA application
- Supports LLaMA (layers.N) and GPT-2 (h.N) naming
Part C: Loss watchdog (like Axolotl loss_watchdog_threshold)
- loss_watchdog, loss_watchdog_threshold, loss_watchdog_patience config
- Implemented in SoupTrainerCallback with patience counter
- Rich warning panel (stops Live display first), fires only once
- Wired into all 11 trainers via callback kwargs
Part D: Dataset info registry
- soup data register/unregister/registry commands
- ~/.soup/datasets.json local name→path+format mapping
- Name validation, path traversal protection, Rich markup escaping
82 new tests (2061 total), 74 test files.
Rich/Typer truncates help panel on narrow terminals (macOS CI), causing
--bits and --group-size flags to not appear in rendered help text. Switch
to inspecting the function signature directly for cross-platform reliability.
- AWQ export (`soup export --format awq`) via autoawq, with --bits, --group-size, --calibration-data
- GPTQ export (`soup export --format gptq`) via auto-gptq, with calibration data support
- Sample packing (`packing: true`) for SFT/Pretrain trainers via TRL's native packing
- `soup data split` — train/val/test splitting with random and stratified strategies
- Curriculum learning (`curriculum: true`) — sort dataset by difficulty for staged training
- New utility: soup_cli/utils/curriculum.py (sort_by_length, create_buckets)
- Security: calibration data path traversal protection, bits validation (4/8 only)
- 1970 tests across 70 test files
- Extract _parse_json_array into soup_cli/data/providers/_utils.py to
avoid circular imports between generate.py and provider modules.
- Narrow bare except Exception in detect_ollama to httpx.HTTPError/OSError
with debug logging instead of silent swallow.
Replace simple '..' check with resolve() + relative_to(cwd) for output
path. Add same confinement guard to --seed, --dedup-with, and --context
file paths. Add _path_within_cwd helper. 4 new security tests.
Rich markup wraps --model-a with ANSI codes on macOS, breaking the
substring check. Strip ANSI codes before asserting, matching the
existing pattern in test_speculative_decoding.py and test_deploy_ollama.py.
Rich markup in Typer help output inserts ANSI escape codes around
--flag names on macOS, breaking exact string matches. Check for
lowercase words instead of --prefixed flags.
Rich inserts color codes between flag name parts (e.g. --speculative
becomes \x1b[1;36m-\x1b[0m\x1b[1;36m-speculative\x1b[0m), so plain
substring match fails in CI. Strip ANSI before asserting.
- Rename --spec-tokens to --num-speculative-tokens to avoid prefix
collision with --speculative-decoding in Typer help rendering
- Add pytest.skip for _create_app tests when fastapi is not installed
Typer returns exit code 2 (not 0) when no_args_is_help=True and no
arguments are provided. Fix test_no_args_shows_help and
test_data_no_args_shows_help to accept both 0 and 2.
- Check stdout encoding instead of type to detect non-UTF-8 consoles
- Redirect stdout to UTF-8 TextIOWrapper before plotext renders
- Add unit test that simulates cp1251 stdout with plotext
- 1022 tests total
- soup data validate: default --format changed from 'alpaca' to 'auto',
uses detect_format() to auto-detect dataset format
- soup data stats: force UTF-8 stdout on Windows for plotext histograms
- soup ui: add --show-token flag, document auth token in --help
- 7 new tests (BUG-013/014/015), 1021 tests total
Add continued pre-training task and Mixture of Experts model support:
- `task: pretrain` for continued pre-training on raw text data
- `plaintext` data format ({"text": "..."} JSONL or .txt files)
- MoE model detection (Mixtral, Qwen3 MoE, DeepSeek V3, DBRX, OLMoE)
- ScatterMoE LoRA (`moe_lora: true`) targets expert FFN + attention layers
- `moe_aux_loss_coeff` for router load-balancing loss
- Templates: `soup init --template pretrain` and `--template moe`
- 85 new tests across test_pretrain.py and test_moe.py (1002 total)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Test _load_model exit paths: adapter without base model, corrupt JSON
- Test _generate branches: greedy (temp=0), sampling (temp>0), no
chat_template fallback, token count from tensor shape, role formatting
- Test max_tokens bounds: 0 and 99999 rejected by CLI
- Test tensorboard happy path: flag accepted when tensorboard installed
- Fix import-failure test: avoid builtins.__import__ recursion
- 917 tests, 44 test files, 57.92% coverage
- Fix test_tensorboard_in_train_help: strip ANSI escape codes before
asserting --tensorboard in help output (Rich splits flag across
escape sequences on Python 3.11)
- Fix TensorBoard import check: use `import tensorboard` directly
- Stream JSONL output during inference (crash-safe for large files)
- Return accurate token count from _generate via tensor shape
- Replace shallow tests with real trainer integration tests
- Cap max_tokens at 16384 + trust_remote_code warning
- Fix TensorBoard import check: use `import tensorboard` directly
(not torch.utils.tensorboard shim) for accurate availability check
- Stream JSONL output during inference instead of buffering in memory
(crash-safe, handles large prompt files)
- Return accurate token count from _generate via tensor shape instead
of re-encoding decoded text
- Replace shallow tests with real trainer integration tests that
verify report_to='tensorboard' is accepted by all trainer wrappers
Security fixes across all HTTP surfaces:
- Web UI: Bearer token auth on mutating endpoints, CORS restricted to served origin,
path traversal protection on /api/data/inspect, config validated before training,
removed user-controlled config_path from API
- Serve/vLLM: max_tokens capped at 16384, generic error messages (no stack traces)
- Generate: SSRF protection (--api-base blocks non-HTTPS for remote URLs),
--api-key deprecated in favor of OPENAI_API_KEY env var
- Export: llama.cpp pinned to tag b5270 (supply-chain safety)
- Push: --token deprecated in favor of HF_TOKEN env var
- Rewards: warning before executing custom .py reward files
- Tests: all 40 UI tests updated with auth headers, 666 tests pass
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
GRPO:
- Set default chat_template on tokenizer when missing (fixes ValueError
from trl's apply_chat_template on models without chat support)
- Ensure batch_size >= num_generations (trl 0.28 requirement)
- Verified end-to-end GRPO training on CPU succeeds
PPO:
- Tokenize dataset via .map() before passing to PPOTrainer (adds input_ids
and attention_mask columns required by trl experimental API)
Tests: 666 passed, 5 new tests for chat_template/tokenization fixes,
3 existing mock tests updated for new tokenization step.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- PPO: Skip resume_from_checkpoint when experimental PPOTrainer.train()
doesn't accept it (inspect signature at runtime, warn and proceed).
- CPU: Use device_map="cpu" instead of "auto" on CPU across all trainers
(SFT, DPO, GRPO, PPO, RewardModel) to prevent meta tensor errors.
- Add 12 new tests for both fixes (661 total passing).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- PPO: Support trl >=0.28 experimental API (ref_model, reward_model,
train_dataset, value_model positional args). Auto-import from
trl.experimental.ppo with fallback. Create reward/value models when needed.
- GRPO: Fix CPU empty generation tensor mismatch by passing
generation_kwargs={"min_new_tokens": 1} on CPU devices.
- Add 6 new tests for both fixes (649 total passing).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Use create=True in mock.patch so tests work when trl has moved
PPOTrainer to trl.experimental and it's not in the trl namespace.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
PPOTrainer.__init__() no longer accepts dataset= in newer trl versions.
Now checks via inspect.signature whether train_dataset or dataset is
accepted; if neither, sets dataset on trainer before .train() call.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- PPO: detect trl API via inspect — args= (>=0.28) vs config= (<0.28)
- PPO: split train into _train_builtin (trl >=0.28) and _train_manual
- GRPO: update error message to mention GRPO/PPO CPU limitation
- 2 new tests for PPO API detection (639 total, all passing)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- PPO: add use_cpu=True to PPOConfig when running on CPU
- GRPO: add CPU warning + use_cpu flag via inspect (trl bug workaround)
- Add use_cpu error pattern to friendly error map
- 7 new tests for CPU fixes (637 total, all passing)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>