mirror of https://github.com/razor-ai/soup.git
276 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
e27bf81eb7 |
test(ci): assert governor-db null-byte rejection on validator directly (POSIX-safe)
POSIX os.putenv forbids null bytes in env values, so monkeypatch.setenv(SOUP_EDIT_GOVERNOR_DB, 'x\x00.db') raised ValueError at setenv time on ubuntu/macos before the code under test ran (windows tolerated it). Assert _validate_governor_db_override rejects the null byte directly; the validated-None fallback branch is already covered cross-platform by test_env_override_out_of_bounds_falls_back. |
|
|
|
96c339f184 |
feat(edit): live ROME/MEMIT/AlphaEdit + GRACE + NPO/SimNPO/RMU unlearn (v0.71.9)
Closes #193, #194, #196, #197, #203. - #194 utils/edit_kernels.py: covariance-free rank-1 ROME/MEMIT/AlphaEdit; apply_edit live (load -> optimise residual -> rank-1 update -> save); edit diff live before/after generation. - #196 EditGovernorStore SQLite persistence + cross-process lock. - #197 apply_edit consults the governor (check_can_edit before, record after). - #203 GraceCodebook + apply_grace_edit + install_grace_hook + Registry kinds. - #193 utils/unlearn_kernels.py (NPO/SimNPO/RMU) + live UnlearnTrainerWrapper + soup train --task unlearn. Validated on SmolLM2-135M: ROME 0.0016->0.96, NPO/SimNPO forget loss down. +81 tests (tests/test_v0719.py). 2 review waves, all findings fixed. |
|
|
|
eb7655f81e |
test(ci): ANSI-strip help/error asserts in test_v0718 (FORCE_COLOR-robust)
CI (FORCE_COLOR) makes Rich/Typer split flag tokens at colorized hyphens (--auto-download -> -auto -download) and auto-highlight `=` in error text (name=path), so contiguous-substring asserts fail. Add the _clean_help helper (strip ANSI + all whitespace, matching the v0.71.1 / test_v0717 pattern) and apply it to the sae-diff / train / sleeper / interference --help asserts plus the bad-adapter-spec name=path error assert. Reproduced + verified with FORCE_COLOR=1 locally. No source change; test count unchanged. |
|
|
|
823456c1a5 |
feat(probe): real probe weights, SAE auto-download, truth/harm, interference --measure, capture-activations (v0.71.8)
Closes #216, #217, #218, #219. Partial #215 (calibrated vectors upstream-gated). - #215 probe_kernel.py: compute_contrast_probe + load_probe_weights (.npz/.npy/.safetensors, O_NOFOLLOW, allow_pickle=False, cwd-contained); soup probe sleeper --weights. Synthetic seed fallback retained. - #216 hubs.snapshot_download (SSRF-hardened, home/cwd/tmp cache, TOFU gate) + sae_diff.download_sae (allowlist-before-network + symlink-escape guard); soup probe sae-diff --auto-download. - #217 truth_probe.py + harm_probe.py over probe_kernel; soup probe truth/harm; probe pack ships truth+harm per base. - #218 interference_live.measure_interference_losses (live PEFT multi-adapter, add_weighted_adapter cat off-diagonal); soup probe interference --measure. - #219 live_eval.extract_layer_activations + resolve_layer_module PEFT-fallback; soup train --capture-activations writes <output>/activations/activations.json. Test count 12771 -> 12917 (+146 in tests/test_v0718.py). Step-6 smoke on SmolLM2-135M (RTX 3050) green; caught + fixed a PEFT-wrapper layer-resolution bug. |
|
|
|
f0118ccffe | test(ci): ANSI-strip help-text asserts in test_v0717 (FORCE_COLOR-robust) | |
|
|
f097528ac0 |
feat(eval): live eval runners — advise/tunability/capability/behavior/diagnose (v0.71.7)
Closes #161, #162, #208, #211, #212, #165. New utils/live_eval.py shared model-loading layer (lazy torch/transformers/peft): load_model_and_tokenizer, make_generator/make_multi_generator, compute_eval_loss, lora_probe, measure_logit_agreement, token_f1. - #161 soup advise --probe-model: live zero/few-shot token-F1 + LoRA probe - #162 base_model_proximity via held-out logit agreement - #208 soup tunability --live: per-candidate LoRA probe - #211 soup eval capability --live --model: lm-eval-harness per task (per-task isolation) - #212 soup eval behavior --base-model: live pre/post battery diff - #165 soup diagnose --base-model: utils/diagnose/live.py runs all 6 probes live Heuristic/neutral paths preserved when no model is supplied. Both new JSONL readers open with O_NOFOLLOW after cwd-containment (TOCTOU close). +68 tests (12703 -> 12771). Smoked end-to-end on SmolLM2-135M (RTX 3050). |
|
|
|
a1463bf716 |
feat(v0.71.6): live build runner + Magpie generator + 2PL/3PL IRT + augment fix
Lift the v0.69.0 deferred stubs to live + extend IRT + fix a real bug: - #231 soup build materialises (5 built-in transforms, table/view/incremental with SQLite-tracked config-fingerprint cache key, atomic JSONL, --output-dir) - #232 soup data gen-magpie live (ollama/vllm raw-completion harvest; anthropic rejected; optional --quality-filter; dedup-before-response) - #167 tokenizer-aware memorization probe (sub-word/BPE overlap, library-only) - #213 soup eval irt-subset --model 2pl|3pl (joint coordinate-ascent MLE) - #75 fix soup data augment --provider ollama|vllm ImportError + QA log Security: validate_ollama_url/validate_vllm_url reject 0.0.0.0; augment output containment+symlink reject; magpie response-body cap. Tests 12581 -> 12703 (+122 in tests/test_v0716.py). Full suite green, ruff clean. |
|
|
|
1f63393421 |
feat(v0.71.5): ingest/data/prompt/drift polish
Closes #157, #205, #207, #149, #164, #163. Defers #204 (live SaaS pull — paid accounts, infra-blocked, kept open). - #164: get_metric_series falls back to eval_results when metrics is empty - #163: build_verdict confidence biased by advise_history (same project+choice, >=3 precedents); decision never changes - #207: shared utils/webhooks.py (SSRF-hardened) + --slack-url/--discord-url on ingest/prune-prompt/ab/active-sample; ab fires only on a decision - #205: soup prune-prompt --tokenizer (token-prefix detect + decode remainder, boundary-safe) - #149: DynamicCurriculumCallback buckets by loss/perplexity percentile; length keeps round-robin - #157: soup data push/forge --hub modelscope|modelers (data score N/A) 107 new tests in tests/test_v0715.py (12474 -> 12581). ruff clean. |
|
|
|
76fdd848cb |
test(ci): ANSI-strip help-text asserts in test_v0714 (FORCE_COLOR-robust)
Rich splits `--pre-wired` / `--pack-cans` / `--push` with ANSI escapes under CI FORCE_COLOR; _clean_help() strips them before the substring check (same fix family as v0.71.1/v0.71.3). No src change. |
|
|
|
5652215d4a |
feat(adapters,loop): v0.71.4 — live canary verdict + cmaes merge + PR push + pre-wired loop + can lineage + branch↔registry
Closes #172, #173, #176, #177, #220, #223. - #172 soup adapters merge --canary/--strict-verdict: live OK/MINOR/MAJOR verdict (was UNKNOWN stub) - #220 soup adapters merge --strategy cmaes: live merge→score→write-best loop (was plan-only) - #223 soup adapters pr --push owner/repo#N: post PR comment via gh api - #176 soup loop --pre-wired: real traces→DPO→eval-gate→canary stages - #177 soup loop --pack-cans / replay --extract: iterations as Soup Cans + Registry lineage DAG - #173 soup adapters branch --from-registry / --attach-to-registry Security: backdoor-scan + license gates now run for ALL merge strategies (incl cmaes); loop canary deploy restricted to loopback/RFC1918; gh child env from allowlist; canary read uses O_NOFOLLOW+fstat (TOCTOU); pack-entry failure rolls back registry entry. Tests: 12342 → 12474 (+130 in tests/test_v0714.py). |
|
|
|
22d5c4f226 |
test(ci): ANSI-robust help asserts + POSIX-safe audit test (v0.71.3)
CI (FORCE_COLOR) renders --track-energy / --no-audit-log as split ANSI colour segments, and monkeypatch.setenv with a null byte raises at setup on POSIX (Windows tolerated both). Strip ANSI via a shared `_plain()` helper for every --help substring assert, and rewrite the never-raises audit test to monkeypatch append_audit_event to throw instead of injecting a null-byte env path. |
|
|
|
21a2bf8e8c |
feat(governance,energy): v0.71.3 — annex PDF, audit auto-log, energy hook, can v3, airgap receipt
Closes #180 #181 #182 #183 #184 #188. - #180 EnergyTracker (codecarbon offline) + `soup train --track-energy`; [carbon] extra - #181 PDF Annex XI/XII (reportlab) + paths.atomic_write_bytes; [pdf] extra - #182 Soup Can manifest v3 + attestations field + `can pack --attest` - #183 per-command audit-log auto-instrumentation (--no-audit-log / SOUP_NO_AUDIT_LOG) - #184 auto-populate Annex top_domains from the training JSONL - #188 embed repro-receipt into `soup airgap-bundle` New [pdf]+[carbon] extras (reportlab also in [dev]). +83 tests (12259 -> 12342), 79.10% coverage. Reviewed (security/code/python/tdd): security M1 raw-size gate on --attest before parse, python H1/H2 type hints, +22 negative tests. |
|
|
|
5d7828d40b |
test(ci): ANSI-robust help-text asserts in test_v0712 (v0.71.2)
CI installs [dev] with FORCE_COLOR, so Rich colorizes Typer --help and splits an option name like --key into ANSI-wrapped segments (\x1b[1;36m-\x1b[0m\x1b[1;36m-key\x1b[0m). The 4 raw-substring help asserts passed locally (no color) but failed on all 9 CI test jobs. Add a module-level _strip_ansi() helper and route the sign/verify/merge/ attest-emit --help substring checks through it (mirrors the v0.71.1 test_serve --record-thumbs fix). Confirmed locally under FORCE_COLOR=1: all 4 pass; ANSI-strip alone is sufficient (no flag line-wraps). Test-only change on the unreleased v0.71.2 — no version bump. |
|
|
|
9300ee3412 |
feat(governance,sign): v0.71.2 — ed25519 signing + supply-chain gates
ed25519 signing (#179/#185 ed25519 half): new utils/signing.py + [sign] extra. `soup adapters sign/verify --backend ed25519` (--key/--generate-key/ --public-key) and `soup attest emit --sign ed25519` + new `soup attest verify`. Sigstore keyless stays infra-blocked (OIDC/Fulcio/Rekor network). #186 namespace-pin gate wired into download_repo (anti-AI-Jacking; fail-open) #187 adapters merge auto-detects each adapter's license #190 license-override reason recorded to the audit log #191 NamespacePinStore WAL + busy_timeout + cross-process lock #192 merge refuses scan-FAIL inputs unless --allow-unscanned +106 tests (12153 -> 12259); full suite 79.07% cov. cryptography added to [dev] for CI. 5 review waves, all findings fixed. |
|
|
|
0ee5f78986 |
fix(ci): ANSI-robust serve help test + restore 77% coverage gate (v0.71.1)
The v0.71.1 release commit (
|
|
|
|
514761c89a |
feat(env,lock,serve,eval): v0.71.1 — quick wins + wiring (7 closures)
Closes #195 #210 #214 #224 #230 #233 #209. - soup env fix: print-only install-plan renderer from soup-env.lock (uv-pip / requirements; non-pip entries surfaced as comments). (#209) - soup lock write --env-lock: auto-derive --env-hash from soup-env.lock via new compute_env_hash (excludes created_at). (#224) - soup serve --record-thumbs <db>: capture thumbs-up/down into the local-RL SQLite + POST /v1/thumbs (transformers backend). (#230) - Judge-calibration persistence: JudgeCalibrationReport.to_dict + write/load_judge_calibration + judge_calibration registry kind; load re-validates the frozen dataclass with cwd/symlink containment. (#214) - soup completions: introspect a base model's real LoRA target modules (config-only AutoConfig, local_files_only, never networks/raises). (#210) - Bundled MUSE + WMDP unlearning eval fixtures; WMDP forget rows ship REDACTED (Soup never bundles verbatim hazardous content). (#195) - build_dag.validate_build_source: cwd-containment + symlink rejection. (#233) Review-fix hardening (consolidated python+code+security+tdd, 0 CRIT/0 HIGH): load_judge_calibration containment + friendly missing-field ValueError; serve thumbs success-print escape; env_fix --output Optional[str]; empty --env-hash auto-derives; render_install_plan PEP 440 docstring note. Tests: 12071 -> 12134 (12044 passed, 90 skipped, 2 deselected). |
|
|
|
894cb632dd |
chore: release v0.71.0 — split heavy deps into [train] extra
Heavy training stack (torch, transformers, peft, trl, datasets, bitsandbytes, accelerate) moves out of the core install into a new [train] optional-dependency extra. `pip install soup-cli` is now a light CLI + data-tools install with no PyTorch; `pip install 'soup-cli[train]'` adds the training stack. - pyproject: new [train] + [all] extras; [dev] self-references [train] so CI (`pip install -e ".[dev]"`) still gets torch. Pins unchanged. - errors.py: missing torch/transformers/peft/trl/datasets/bitsandbytes/ accelerate now surface a single 'install soup-cli[train]' fix. - Dockerfile: install soup-cli[train,serve,data,eval] so the GPU image can still fine-tune. - README + docs/models.md: split install into light core vs [train]. - CHANGELOG: cut [0.71.0]; bump version 0.70.0 -> 0.71.0. |
|
|
|
06d8ea7cc5 |
chore: migrate to src-layout
Move soup_cli/ -> src/soup_cli/ (history preserved via git mv). src-layout
forces the test suite to import the installed package instead of the
repo-root source tree, surfacing packaging bugs that flat-layout masks —
e.g. the v0.53.8 double-shipped-fixtures regression, invisible because
`pytest tests/` imports ./soup_cli directly and never from the wheel.
- pyproject: packages = ["src/soup_cli"]; artifacts globs -> src/soup_cli/...
The import name is unchanged, so the `soup` entry point, --cov=soup_cli,
and report_to/module-path strings stay `soup_cli`.
- CI / ownership: ruff lint path (ci.yml), recipe-validation `paths:` filters,
CODEOWNERS patterns, and the PR-template checklist all repointed to
src/soup_cli/.
- tests: source-grep regression tests that read package files by repo-relative
path repointed to src/soup_cli/ (64 files; 170 path literals). Lines pushed
over 100 chars by the prefix were wrapped to keep ruff E501 clean. Module
references (`import soup_cli`, `-m soup_cli`, mock.patch("soup_cli.x")) and
the `--cov=soup_cli` coverage target are deliberately unchanged.
- docs: AGENTS.md + CONTRIBUTING.md structure tree and lint commands.
Verified locally: ruff clean (src/soup_cli + tests); `import soup_cli`
resolves to src/soup_cli/__init__.py; built wheel ships
soup_cli/data/_fixtures/*.jsonl (10 files, no duplicates, no src/ prefix);
3174 tests across every touched test file pass. Packaging-only — no version bump.
|
|
|
|
6ec9db3ca7 |
fix(tests): strip ANSI escapes before --lang substring assertion
Rich on narrow Windows columns splits `--lang` across colour-cycle ANSI escapes (`\x1b[..m-\x1b[..m-lang`), so the literal substring check fails on CI even though the rendered help renders correctly for humans. CI was red on every commit landing after #234 hit a runner with that exact column width + Python 3.9 + Rich combination. Mirrors the `_ANSI_RE` strip pattern in tests/test_auto_tuning.py — flat regex over the captured output before the `in` check. |
|
|
|
148cb0c125 |
fix(active-sample): variance-based diversity score for K>2 reward models (#206)
v0.63.0 `score_uncertainty` raised on K>2 and `_row_uncertainty` fell back to a monotone-broken `max(scores) - min(scores)`. Now generalises to K<=32 via population variance scaled by 4 — adding a fresh RM score equal to the running mean strictly decreases uncertainty (the new contribution to the sum-of-squares is zero while the denominator grows), so consensus on redundant evidence can never spike the score. K=1 max-entropy and K=2 disagreement formulas preserved verbatim (existing operator dashboards depend on the |s1 - s2| value). Cap stays at K=32 for DoS defence. _row_uncertainty K>2 path now routes through score_uncertainty inside an isolated try/except — bad rows return 0.0 instead of crashing the batch. PEP 585 modernisation: collections.abc imports + list[...] annotations (safe because `from __future__ import annotations` is in scope). math.fsum used for the variance accumulation to keep rounding error sub-ULP at K=32. Tests: +32 net (25 in new tests/test_v0631_206.py + 7 TDD review-fix followups). Full suite 11941 -> 11973 pass. Closes #206. |
|
|
|
525a0e1114 |
feat(brain-rot): per-language low-effort + clickbait bundles for es/fr/de/ru (#234)
Extends v0.69.0 Part E score_triviality + score_popularity_signal to non-English corpora. New utils/brain_rot_lang.py ships a MappingProxyType registry of frozen BrainRotLangBundle for en/es/fr/de/ru. Every public scorer accepts an optional lang kwarg (default None preserves v0.69.0 English behaviour). The "auto" sentinel routes through the v0.53.10 [data-pro] langdetect helper with silent fallback to English on missing-package / detector-exception / unsupported-code. soup data brain-rot gains --lang en|es|fr|de|ru|auto, strictly validated at the CLI boundary (exit 2 on typos). Per-row resolution backed by eager _validate_lang_arg on dataset scorers so empty rows cannot bypass shape checks. Closes #234. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
95f9a7116d |
fix(diagnose): multilingual refusal patterns — en/es/fr/de/ru (#166)
Extends soup_cli/utils/diagnose/refusal.py from English-only to en/es/fr/de/ru via a MappingProxyType-wrapped _REFUSAL_PATTERNS_BY_LANG registry + public SUPPORTED_REFUSAL_LANGS frozenset. Adds a `lang: str = "en"` keyword on `looks_like_refusal` and `score_refusal` with strict validator (bool / null-byte / oversize / unknown / case-insensitive normalisation), resolved ONCE per `score_refusal` invocation via a new `_apply_pattern` hot-path helper so the per-prompt path skips redundant dict lookups (~8k saved on a 2k-prompt run). Closes #166 — v0.56.0 Known Limitations bullet #3 (English-only refusal heuristic). Review fixes applied (12 total): - python-reviewer (4): frozenset[str] subscript, hot-path refactor, intentional-internal-access comment on _MAX_REFUSAL_SCAN, dropped redundant forward-reference quotes. - tdd-guide (8): TestApplyPattern direct coverage, generator-type guard with default lang, cross-language dispatch matrix, evidence string lang assertion, renamed misleading test, empty-prompts matrix, whitespace-in-lang, source-grep regression guards. 112 new tests in tests/test_refusal_multilingual.py. Pre-existing v0.56.0 test_v0560.py::TestRefusal block unchanged (back-compat verified end-to-end). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
4e95d4c71f
|
feat(echo-trap): add tokenizer-aware repetition scoring (#242)
Closes #241. Adds opt-in tokenizer-aware n-gram path for the v0.70.0 Part F echo-trap detector. The existing whitespace `score_echo_signal` is unchanged; callers opt in via the new `score_trajectory_repetition_tokenized` / `score_echo_signal_tokenized` helpers or the `--echo-trap-tokenizer-aware` train flag. Acceptance criterion from #241 verified: synthetic case where decoded strings differ by punctuation but token-id sequence repeats — whitespace path returns OK, tokenizer-aware path returns TRAP. |
|
|
|
74edac95d1 |
feat(v0.70.0): Loop Hardening — reward-hacking + ULD + MiniLLM + RL ckpt + iterative DPO + echo-trap
Six-part schema-only release shipping the axis-3 + axis-13 training-loop hardening. Every live trainer-callback / math kernel is deferred to v0.70.1 per the project's established stub-then-live cadence (matches v0.50.0 / v0.62.0 / v0.69.0). Part A — Reward-hacking detector (soup_cli/utils/reward_hacking.py): InfoRM Cluster-Separation Index (Wang et al. 2024 arXiv:2402.09345) + RM-ensemble pairwise variance + OK/WARN/HACK taxonomy at 0.10/0.30 thresholds. TrainingConfig.reward_hack_detector / reward_hack_halt; SoupConfig task-gate (grpo/ppo only, mlx rejected, halt requires detector). Part B — Cross-tokenizer ULD (soup_cli/utils/uld.py): Universal Logit Distillation (Boizard et al. 2024 arXiv:2402.12030). wasserstein + topk_align allowlist; ULDConfig frozen with topk/strategy cross-validators; vocab-size cap 262144. Schema-gated to task='distill'. Part C — MiniLLM reverse-KL on-policy distillation (soup_cli/utils/minillm.py): Gu et al. 2024 (arXiv:2306.08543) — bundles teacher-mixed sampling + length-norm + pretrain-loss anchor stability tricks. Anchor weight↔path mutual-requirement cross-validators reject silent no-op combos. Part D — Mid-epoch RL checkpoint (soup_cli/utils/rl_checkpoint.py): Optimizer-state serialization TorchTune explicitly punts. RLCheckpointConfig + RLCheckpointState frozen + JSON-serialisable manifest. RL-task gate. Part E — Iterative DPO loop driver (soup_cli/utils/iterative_dpo.py + commands/iterative_dpo.py): sample → RM-score → re-pair → retrain over N rounds. IterativeDPOPlan with consecutive-round_index invariant. New `soup iterative-dpo` CLI; --plan-only live, runner deferred. Part F — RAGEN echo-trap detector (soup_cli/utils/echo_trap.py): Zhu et al. 2025 (arXiv:2504.14437) — n-gram trajectory-repetition kernels with DoS caps (max 32 ngram_n, 1M tokens, 100k trajectories). OK/WARN/TRAP at 0.30/0.60. Composes with v0.53.11 #127 GRPOStabilityCallback. Cross-cutting hardening: - 6 new util modules + 1 new top-level CLI + 11 new TrainingConfig fields + 6 new SoupConfig cross-validators + 3 new field validators - Closed allowlists (frozenset) + MappingProxyType registries everywhere - Frozen dataclasses with post-init validation on every public record - Bool-as-int rejection on every numeric (matches v0.30.0 / v0.41.0 policy) - math.isfinite NaN/Inf rejection on every float - Null-byte rejection + per-field length caps on every string - No top-level torch imports (4 source-grep regression tests) - Deferred-live stubs validate inputs FIRST then raise NotImplementedError with explicit v0.70.1 marker - CLI exit codes split: 2 = validation rejection, 3 = deferred-live Test count: 11487 → 11824 (+337 net). 12-invariant self-review against the full project checklist (closed allowlists, frozen dataclasses, MappingProxyType, bool-as-int rejection, finite check, null-byte, length caps, no top-level torch, TypeError/ValueError split, deferred-live, tuples-not-lists, CLI exit codes) all green across all 6 Parts. Manual CPU smokes (Step 6): every CLI happy + failure path exercised — `soup iterative-dpo --plan-only` 3-round plan rendered end-to-end with per-round artifacts; 5 happy-path YAML loads + 5 failure-mode rejections across reward_hack / uld / minillm / rl_checkpoint / echo_trap. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
49943a5af6 |
feat(v0.69.0): Data Engineering Pro — soup build + expect + gen-magpie + persona-mix + brain-rot
5 parts shipping axis-2 (dbt-for-SFT) + axis-13 (data ops): - soup build — dbt-for-SFT DAG with refs / incremental materialization / content-hash row-diff kernel (run_build live runner deferred → v0.69.1) - soup expect <data> <suite> — LIVE expectations suite: PII / token-length / refusal / chosen-vs-rejected judge; exit 3 on suite failure - soup data gen-magpie — Magpie synthetic generator plan (live → v0.69.1) - soup data persona-mix — Persona-Hub × style sampler with bundled 12×5 set, atomic JSONL write (LIVE) - soup data brain-rot — arXiv 2510.13928 detector with --strict CI gate, worst-signal composite (LIVE) Centralised TOCTOU defence behind utils/paths.enforce_under_cwd_and_no_symlink in build_dag / expectations / expect.py (code-review CRIT — replaces 3 duplicate os.lstat + S_ISLNK + realpath + is_under_cwd blocks). DoS caps on every new JSONL loader (brain-rot 1 GiB + 1M rows; persona-mix 100 MiB + 100k entries). persona-mix --output TOCTOU symlink rejection. magpie quality_filter validator + expectations._dispatch_expectation raw-args pass-through (no int/float coercion bypass). BuildModel seed/derived cross-validator rejects ambiguous shapes at schema load. Review-fix coverage across 4 waves (security + code + python + TDD): 1 CRITICAL + 4 HIGH + 5 MEDIUM + 4 LOW. Test count: 11225 → 11487 (+262 net across 5 new files). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
aa71658f50 |
feat(v0.68.0): Anti-trend Insurance — compile (DSPy/GEPA) + distill-prompt + compile-tools + apple-adapter + local-rl
5 commands that hedge Soup against paradigm shifts. If 1M-context kills FT, `soup compile` (DSPy + GEPA + TextGrad prompt-program compilation) takes its place. If teams hit prompt-cost walls, `soup distill-prompt` bridges to small FT. If only Apple Foundation Models win on-device, `soup apple-adapter` ships the converter+signing surface. If personal-LLM flywheels become the shape, `soup local-rl` captures thumbs into SQLite and emits DPO pairs. - Part A: `soup compile <program.py> --eval <suite> [--optimizer mipro|gepa|...]` - Part B: `soup distill-prompt --traces <jsonl> --teacher --student --strategy` - Part C: `soup compile-tools <spec.json|yaml> --eval <jsonl>` - Part D: `soup apple-adapter <source-dir> --direction hf-to-mlx|... --output` - Part E: `soup local-rl init/status/record/harvest/train` (LIVE except train) Schema + path containment + symlink rejection + atomic-write surface ship now; live runners for Parts A/B/C/D + Part E nightly scheduler deferred to v0.68.1 (stub-then-live, mirrors v0.50.0 / v0.61.0 / v0.62.0 / v0.67.0 cadence). Test count: 11021 -> 11225 (+204). Review-fix: 0 CRIT + 4 HIGH + 10 MED + 4 LOW. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
32145097ec |
feat(v0.67.0): Adapter Lifecycle Finish — CMA-ES merge + VeRA bank + MoLE + PRs + soup.lock + bisect
Six surfaces close v0.57: - Part A: pure-Python rank-mu CMA-ES evolutionary merge (cmaes_merge.py) + soup adapters merge --strategy cmaes --eval <s> --budget 1h - Part B: VeRA / VB-LoRA vector-bank schema + atomic JSON I/O (vector_bank.py) - Part C: MoLE per-token routing schema + new task='moe_lora_routing' (mole_routing.py) - Part D: GitHub-shaped adapter PR renderer (adapter_pr.py) + soup adapters pr <title> --base-sha --adapter --eval --samples - Part E: soup.lock shared run lockfile (soup_lock.py + commands/lock.py) + soup lock write/show/check (exit 3 on drift) - Part F: training-history binary search (adapter_bisect.py) + soup adapters bisect <ckpts> --eval-command "..." Live wiring deferred to v0.67.1: CMA-ES eval-suite auto-bind, VeRA serving, MoLE gating kernel. +185 tests (10836 -> 11021) across 7 new test files. Review-fix coverage from 2 sequential waves (security + tdd-guide). All step-6 smokes green. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
a015ccc812 |
feat(v0.66.0): Post-train X-rays — SAE diff + live blame + sleeper probe + interference matrix + probe pack
Extends `soup diagnose` from 6 failure modes to 10. Closes v0.57 #171 — live blame runner replaces the NotImplementedError stub. - `soup probe sae-diff`: SAE feature attribution (pure-numpy; HF_HUB_ALLOWLIST) - `soup adapters blame --top-k 50`: live DataInf influence runner (closes #171; replaces v0.57 stub with cos(grad_row, grad_probe) × |grad_row|) - `soup probe sleeper`: calibrated defection probe (6 bundled bases; OK/MINOR/MAJOR at 1%/5%; exit 2 on MAJOR) - `soup probe interference`: pairwise N×N matrix (OK/MINOR/MAJOR at 5%/20%; exit 2 on MAJOR worst-pair) - `soup probe pack`: per-base probe manifest assembler Review-fix coverage across 3 sequential waves: 0 CRITICAL + 9 HIGH + 14 MEDIUM + 5 LOW. Notable hardening: - TOCTOU O_NOFOLLOW probe-open in load_sae_weights + _count_dataset_rows - hashlib.sha256 replaces process-salted hash() for CI reproducibility - Rich-markup escape on adapter / verdict / description / layer - TypeError on bool/non-str verdict before membership check - Non-numeric loss rejection in `probe interference` CLI - 10M-row hard reject (no silent truncate); 100k synthetic-probe cap - _LOWER_INDEX MappingProxyType for O(1) case-insensitive lookup - Mapping from collections.abc (PEP 585); frozenset[str] type params - Frozen dataclasses + FrozenInstanceError regression tests Note: Windows cp1251 print on stdout-capturing Python wrappers can crash on Rich's '→' arrow output; the soup CLI itself uses force_utf8_stdio. Test count: 10577 → 10836 (+259 net across test_v0660_part_{a-e}.py, test_v0660_cli.py, test_v0660_followups.py). Full suite green (10836 passed, 81 skipped); ruff clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
799f5e8522 |
fix(tests): floor-check version assertion in test_version_bumped_to_0640
Was `assert soup_cli.__version__ == "0.64.0"` — broke on v0.65.0 bump on CI across all 9 platform×Python matrix jobs. The test intent is "v0.64.0 has shipped at least once"; switch to a tuple floor check matching the v0.51 / v0.54 / v0.57 / v0.60 idiom. Caught by CI red on v0.65.0 push to main; local pytest passed because we ran the v0.65 test files in isolation per the Release Checklist Step 4 ``pytest --no-cov`` invocation. Lesson: include the full suite in step 4 or grep for ``soup_cli.__version__ ==`` exact-equality assertions in any future version bump. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
1e822af461 |
feat(v0.65.0): Eval Depth — judge calibration + behaviour battery + capability suite + CheckList DSL + IRT subset
5 LIVE parts closing axis 4 of the roadmap. Evals as first-class surface, not afterthought:
- Judge calibration: SCOPE/CJE-style bidirectional pairwise judging in eval/calibrate.py
with PairwiseJudgement / fit_position_bias / conformal_threshold +
ensure_judge_calibrated production gate that refuses to score with an uncalibrated
judge (RuntimeError on None / calibrated=False / low agreement / extreme bias).
- Behaviour battery (soup eval behavior): closed allowlist of XSTest / HarmBench /
JailbreakBench / ELEPHANT / SycEval with 5 tiny bundled redacted probe sets under
soup_cli/data/_fixtures/behavior/. Word-boundary regex agreement rejects
"safe" in "unsafe" false positives. Pre/post diff with OK/MINOR/MAJOR verdict
(matches v0.26 / v0.56 taxonomy).
- Capability auto-suite (soup eval capability): MMLU-Pro / GPQA / BBEH / AIME /
MATH-500 / HumanEval+ / SWE-bench-Verified with full / fast / math / code profile
selector. Emits (benchmark, lm-eval task) manifest for downstream
soup eval benchmark chaining.
- CheckList DSL (soup eval checklist): Ribeiro et al. 2020 MFT / INV / DIR test kinds
rendered from YAML. Word-boundary matching prevents "and" matching "sand".
Per-test pass/fail + OK/MINOR/MAJOR overall verdict.
- IRT eval-cost optimizer (soup eval irt-subset): 1PL Rasch closed-form fit on
per-item correctness signals + high-info subset selector (full / small / tiny
profiles). 5-10x cut in eval bills without losing ranking power.
Cross-cutting hardening (review-fix coverage across 2 review waves):
- TOCTOU defence: every new read path uses O_NOFOLLOW + os.fstat on SAME fd
(load_checklist_spec, load_response_rows, _read_evidence_json). Earlier
double-lstat-on-path was a race the attacker could win by swapping the file
between calls.
- Namespace-package safety: load_battery_probes uses importlib.resources.files
Traversable / op + as_file (was Path(os.path.join(str(pkg_root), ...)) which
silently fails is_file() on MultiplexedPath installs).
- Word-boundary regex agreement in behavior_battery + checklist_dsl.
- CLI _validate_run_id gate; 16 MiB --evidence cap with O_NOFOLLOW;
_MAX_ROWS=1_000_000 cap counts skipped lines toward total in load_response_rows.
- INV empty-string normalisation no longer spuriously passes.
- _write_json_output / _read_evidence_json / _validate_run_id dedup helpers in
commands/_eval_v0650.py.
Review fixes: 0 CRITICAL + 6 HIGH + 9 MEDIUM + 7 LOW resolved across 2 waves.
Test count: 10306 -> 10577 (+271 net across test_v0650_part_{a,b,c,d,e}.py +
test_v0650_followups.py). Full suite 10577/10577 passing. 0 regressions.
Step 6 smoke: every new CLI command + 3 failure modes exercised end-to-end
(behavior with --evidence happy + MAJOR exit 2; capability fast with output;
checklist with real YAML; irt-subset on 600-row synthetic data; unknown
battery / size / kind all exit 2; outside-cwd evidence rejected).
Known limitations:
1. Live lm-eval-harness invocation deferred — soup eval capability emits the
manifest for downstream soup eval benchmark chaining (Typer commands aren't
safe to re-enter; matches v0.46.0 / v0.44.0 design).
2. Live model-driven soup eval behavior deferred — without --evidence, emits a
neutral OK report (v0.65.1).
3. Behaviour battery probe sets ship as tiny redacted placeholders — operators
pull real harmful prompts from upstream papers.
4. IRT model is 1PL Rasch only (2PL / 3PL deferred to v0.65.x).
Step 6 quirk worth noting: --evidence containment rejects /tmp/ on Windows
WSL bash; operators must run from cwd or pass cwd-contained paths.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
8b5991674b |
feat(v0.64.0): Pre-flight & Tooling — tunability, plan/apply, env, hardware-fit, completions, license-advisor
Six new top-level commands close axis 1 + 11 of the roadmap: pick the right base, lock the env, refuse OOMs before launch, and clear license-clean deploys. - soup tunability: probe-train 8 candidate bases (Qwen3-0.6/1.7B, Llama-3.2-1/3B, Gemma-3-E2B, Phi-4-mini, SmolLM3, Qwen2.5-1.5B) -> Pareto frontier over (delta x cost x license). Live LoRA probe -> v0.64.1. - soup plan / soup apply: Terraform-shape lock-and-execute. `apply` refuses on drift between soup.yaml and soup.tfstate (exit 3). - soup env lock / status / check: hermetic env lockfile via importlib.metadata across 15 ABI-sensitive packages + Python + CUDA. `env check` exits 3 on drift. - Hardware-fit calculator: static analytical 5-bucket VRAM predictor with 10% safety margin + actionable hint on OOM. - soup completions bash|zsh|fish: sourceable shell completion scripts; recipe names auto-complete from the 115-recipe catalogue. - soup license-advisor: per-deploy-target license matrix (b2c/defense/embedded) + Llama community + 700M MAU gate (exit 3). Composes with v0.60 license-conflict matrix. Tests: 10035 -> 10306 (+271 net in 7 new files). Review-fix coverage: 0 CRITICAL + 6 HIGH + 8 MEDIUM + 4 LOW across consolidated code+security+TDD review wave. Every HIGH lands a regression test in tests/test_v0640_followups.py (POSIX-skipped symlink rejection, containment-before-existence ordering, drift-refusal exit-3 end-to-end). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
40bd6251a2 |
feat(v0.63.0): Production Trace Ecosystem — soup ingest + prune-prompt + active-sample + ab + drift-alarm
5 new top-level commands close axis 7 of the roadmap. Every Part LIVE on
day one (no deferred stubs):
- soup ingest: universal trace importer (Langfuse / LangSmith / Helicone /
OpenPipe / OTel / OpenAI Stored Completions). 6 adapters
+ frozen TraceRecord with MappingProxyType-wrapped metadata.
Zero credential-handling threat surface — Soup parses the
JSONL export, never makes the SaaS network call.
- soup prune-prompt: detect + strip a shared system-prompt prefix so the
FT model internalises it (OpenPipe's signature trick,
OSS). Binary-search over up to 32 templates finds the
longest threshold-meeting prefix.
- soup data active-sample: surface top-uncertainty prod traces for human
review. Max-entropy on single rm_score or
pairwise disagreement on dual rm_scores.
- soup ab: Wald sequential SPRT for the point alternative. LLR is a
martingale under H0 so Type-I error is controlled at every
stopping time per the optional stopping theorem.
- soup drift-alarm: rolling KL on whitespace-tokenised output distribution
+ SSRF-hardened Slack/Discord webhook (full parity with
v0.51.0 validate_hub_endpoint). Exit 3 on drift for
cron-friendly automation.
Test count: 9816 -> 10035 (+219 net across 6 new test files).
Review-fix coverage (code-reviewer + tdd-guide returned actionable;
python-reviewer + security-reviewer agents context-thrashed on the large
CLAUDE.md release-notes history — matches the v0.58.0 / v0.59.0 / v0.60.0
/ v0.61.0 / v0.62.0 idiom; verified manually):
- 1 CRITICAL: mSPRT log-likelihood-ratio sign error drove Type-I error
to 1.0 as n grew. Replaced with Wald's classic point-
alternative SPRT (martingale under H0).
- 2 HIGH: detect_common_prefix early-exit on 100% match returned the
shortest qualifying prefix instead of the longest;
_MAX_SCAN_ROWS DoS cap used 'pass' instead of 'break'.
- 3 MEDIUM: TraceRecord.metadata now MappingProxyType-wrapped post-init
(frozen-dataclass mutation hazard); _AUTH_ENV table
deduplicated; drift_alarm precedence parens on SSRF gate.
- 2 LOW: pooled_se dead-branch refactor; mean_uncertainty NaN guard.
- 8 follow-up tests: msprt zero-variance, partial-majority binary-search
activation, score_uncertainty exact boundaries, rolling_kl identical
+ disjoint, validate_budget + validate_threshold exact endpoints,
_signal_from_thumbs boundaries, no-heavy-top-level-imports source-grep
guard across all 5 new util modules.
Step 6 smoke verified for all 5 commands + 6 failure-mode rejection
paths.
CRLF gotcha note for future maintainers: PowerShell wrote the smoke
fixtures with a UTF-8 BOM on Windows during Step 6 — switched to
inline Python for the fixture write. Production CLI input handling is
already BOM-tolerant (utf-8-sig in JSONL loaders via v0.40.1 Part E).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
0d6f95181a |
feat(v0.62.0): RAG & Activation Steering — RAFT + RA-DIT + soup steer + citation-faithful + GRACE codebook
5 Parts shipping wedge 14 of the roadmap (RAG-aware fine-tuning +
activation steering + lifelong edit codebook). Schema-only release;
live training loops + decode-hook intervention + codebook lookup all
land in v0.62.1 (mirrors v0.50.0 / v0.52.0 / v0.61.0 stub-then-live).
Part A — RAFT data format: new data.format='raft' schema +
_convert_raft validator (64 KiB per-field cap, 64-distractor cap,
null-byte rejection on every field) + raft-llama3-8b recipe.
Part B — RA-DIT two-stage: TrainingConfig.ra_dit_stage Literal
{retriever, generator} + ra_dit_retriever_model field + closed
allowlist + cross-validator enforcing stage to base-task pairing
(retriever to embedding, generator to sft) + 2 recipes.
Part C — soup steer (CAA / ITI / RepE): closed-allowlist control-vector
methods + validate_steering_method/name/strength + Typer subcommands
train/apply/list + soup serve --steer/--steer-strength flags +
steering_vector Registry artifact kind. apply_steering +
build_steering_vector deferred-live stubs raise NotImplementedError
with v0.62.1 marker after validating inputs.
Part D — Citation-faithful FT: score_citations precision/recall/F1
kernel + extract_citation_ids public API + citation_faithful /
citation_style / citation_recall_threshold schema. Cross-validator:
citation_faithful=true requires data.format='raft' AND task in
{sft, pretrain} (silent-no-op footgun rejection mirroring v0.52.0
distill / classifier task-gate policy).
Part E — GRACE codebook: GraceCodebookConfig + bounded size [1, 100k]
+ bounded dim [1, 16384]. Extends v0.61.0 SUPPORTED_EDIT_METHODS
allowlist with 'grace'; apply_edit routes grace plans to v0.62.1
marker while legacy rome/memit/alphaedit retain v0.61.1 marker
(regression-guarded via TestEditMarkerRegressionGuard).
Test count: 9571 -> 9786 (+215 net). 4 review-agent waves resolved
0 CRITICAL + 0 HIGH + 4 MEDIUM + 11 LOW (broken list_steers registry
context-manager + dict-key access; missing version bump;
citation_faithful task-gate; shared TOCTOU helper delegation; Rich
markup escape on --steer exception messages; --base length cap +
null-byte rejection; typing.Iterable -> collections.abc.Iterable
migration; except Exception -> except ImportError narrowing).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
4f54179ea4 |
fix(security): backfill Rich markup escape in legacy adapters commands
Closes #174. Picked up after PR #175 (dreamer0129) went quiet — the list/info path was clean, but compare() escaped only inside the highlight branch, leaving a shared crafted base_model = "[link=evil] click[/]" un-escaped on equal-value rows. Fix follows the "escape always at the value layer, decoration wraps after" pattern mirroring v0.57.0 `adapters diff` / `info`: - list_adapters: wrap base / lora_r / peft_type / rel_path with rich.markup.escape() before table.add_row(); also escape adapter_path in the JSONDecodeError fallback. - info: wrap base_model / peft_type / task_type / lora_r / lora_alpha / lora_dropout / modules_str inside the Rich Panel f-string; also escape adapter_path.name in the Panel title. - compare: escape val1_str / val2_str unconditionally; [yellow] highlight wraps already-escaped values when they differ. Equal-value rows now also escape (was the v0.57.0 known-limitation gap). +4 regression tests in tests/test_adapters.py::TestAdaptersMarkupEscape: - test_list_escapes_crafted_base_model — asserts no ANSI hyperlink sequence (\x1b]8;) leaks from a crafted [link=http://evil/...] payload. - test_info_escapes_crafted_base_model — same assertion for Panel. - test_compare_escapes_equal_crafted_values — the specific regression for the PR #175 review gap (identical crafted values on both sides must NOT smuggle live markup through the equal-branch). - test_compare_escapes_differing_crafted_values — highlight branch also escapes. Closes v0.57.0 Known Limitation (9). Verified locally: - ruff check soup_cli/commands/adapters.py tests/test_adapters.py -> clean - pytest tests/test_adapters.py --no-cov -> 20 passed |
|
|
|
6ddaeb30d1 |
polish(v0.60.0 strict-safetensors): tighten header cap + input guards
Follow-up to PR #198 (issue #189): - _MAX_SAFETENSORS_HEADER_BYTES: 1 GiB -> 100 MiB. Real safetensors headers are <10 MiB even for 70B-parameter models; 100 MiB is a generous defence-in-depth ceiling that still rejects an adversary's "header_len = 999 MiB" allocation attempt before fh.read() commits. - is_safetensors_magic: input-shape guards (non-string / empty / null-byte path return False, never raise). Matches project policy for detection-style helpers (mirrors v0.30.0 Candidate, v0.41.0 lr_groups, v0.53.3 is_known_vlm_base). - +2 regression tests in tests/test_v0600_part_c.py: - test_is_safetensors_magic_rejects_invalid_input (5 bad inputs) - test_max_safetensors_header_bytes_tightened (guards against re-widening to 1 GiB in a future patch) Module docstring updated to reference PR #198 / issue #189. Verified locally: - ruff check soup_cli/utils/strict_safetensors.py tests/test_v0600_part_c.py -> clean - pytest tests/test_v0600_part_c.py --no-cov -> 21 passed, 1 POSIX-skipped on Windows Closes v0.60.0 Known Limitation (6) — full magic-byte + JSON-header shape verification with hardened input surface. |
|
|
|
914a299965
|
Check safetensors magic bytes (#198)
Co-authored-by: Sumit Dhawan <sumitdhawan@Sumits-MacBook-Air.local> |
|
|
|
740832e1b4 |
feat(unlearn/edit): v0.61.0 — Unlearning & Knowledge Edit (NPO/SimNPO/RMU + ROME/MEMIT/AlphaEdit)
5 Parts shipping schema + CLI surface for two of the most under-served axes in fine-tuning: GDPR right-to-be-forgotten unlearning (the legal-liability axis upstream TRL avoids) and surgical knowledge editing (research-coded everywhere, productized nowhere). Schema-only release; live trainer + kernel wiring deferred to v0.61.1 (matches established v0.50.0 / v0.52.0 / v0.53.0 stub-then-live cadence). Part A — task='unlearn' + NPO/SimNPO/RMU allowlist + UnlearnTrainerWrapper + data.forget_set / data.retain_set + training.unlearn_method/_alpha Part B — soup eval unlearning (TOFU/MUSE/WMDP) with Forget Quality + Model Utility + PrivLeak kernels + OK/MINOR/MAJOR taxonomy; bundled TOFU mini-fixture under soup_cli/data/_fixtures/unlearning/ Part C — soup edit set (ROME/MEMIT/AlphaEdit) + EditPlan + per-method default layer; --plan-only ships live, apply_edit kernel deferred Part D — Sequential edit governor: norm-blowup detection (OK/WARN/BLOWUP), auto-switch ROME→AlphaEdit at edit#10 or BLOWUP, refuses past cap Part E — soup edit diff: cwd-contained probe loader, atomic JSONL out, shape + table renderer (live before/after generation v0.61.1) Net: +125 tests (9446 → 9571), +5 utility modules + 1 trainer wrapper + 2 commands. Review-fix coverage: 0 CRITICAL + 5 HIGH + 11 MEDIUM + 11 LOW. All ruff + pytest green; Step 6 smokes (CLI plumbing + happy paths + 5 schema rejection paths) confirmed end-to-end. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
47409df730 |
fix(tests): v0.60.0 CI green — widen version floors + strip ANSI from merge help assert
Three failures on CI run 26084542388 — all are version-pin / Rich-wrap artefacts, not real regressions in v0.60.0 functionality: - test_v0560 test_pyproject_version: regex-based >=0.56 floor check (was substring `version = "0.5`) - test_v0590 test_version_is_0_59 -> test_version_is_at_least_0_59: >=0.59 floor (matches v0.51/v0.54 floor-check idiom) - test_v0600_part_e merge_help_lists_license_flags: strip ANSI codes before substring check (Rich splits `--license` across `\x1b[1;36m` escapes in the wrapped Typer table) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
f3c40e7753 |
feat(security): v0.60.0 — Supply Chain Security wedge (adapter scan/sign/verify, strict-safetensors, namespace-pin, license-matrix, airgap-bundle)
Six controls that hosted vendors structurally can't provide: - soup adapters scan: spectral backdoor scanner (rank-1 dominance + energy concentration + NaN/Inf + Frobenius outlier via robust median+MAD); pure numpy, reuses v0.57.0 adapter_diff loader - soup adapters sign / verify: Merkle-root manifest in .soup-signature.json; recursive file enumeration (catches nested tokenizer/ tamper); sigstore + ed25519 backends stub-then-live (v0.60.1) - soup adapters check-safetensors: closed 8-entry unsafe-extension allowlist; strict exit 3 for CI gating - NamespacePinStore: TOFU SQLite anti-AI-Jacking; author + created_at fingerprint compared via datetime.fromisoformat for offset-aware order; bool opt-in rejected so --allow-namespace-shift cannot be a free-for-all - License-conflict matrix: 33 SPDX-ish ids in MappingProxyType compat table; soup adapters merge --license <id> --license-override <reason> gate - soup airgap-bundle: signed tarball with deterministic dataset labeling (sorted basename, NOT argv order); TOCTOU lstat+S_ISLNK on parent + output; atomic os.replace; tarfile.data_filter for future extractall callers Test count: 9294 -> 9446 (+152 net across 6 new test files). Review-fix coverage across 5 waves (python / security / code / tdd / smoke): 0 CRITICAL + 12 HIGH + 11 MEDIUM + 6 LOW fixed before commit. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
9563699eca |
fix(v0.59.0): rewrite null-byte env test — OS layer rejects setenv on every platform
The previous test_default_log_path_rejects_null_byte_env used monkeypatch.setenv to inject a null byte into SOUP_AUDIT_LOG_PATH and expected default_log_path to fall back gracefully. But the OS layer rejects null bytes in env vars on every platform we ship on: - POSIX (Linux/macOS): `ValueError: embedded null byte` - Windows: `ValueError: embedded null character` The setenv call itself raises, never reaching default_log_path. Split into two tests that hit the actual validation surfaces: 1. test_default_log_path_rejects_null_byte_override — calls the private _validate_log_path_override helper directly with a null-byte string and asserts it returns None (so the caller falls back to the safe default). 2. test_default_log_path_handles_env_read_value_error — monkeypatches os.environ.get to raise ValueError, exercising the defence-in-depth try/except around the env read in default_log_path(). Both tests pass on Linux + macOS + Windows. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
d0f5e35c99 |
fix(v0.59.0): macOS CI — strip ANSI from train --help assert + handle null-byte env
Three macOS-3.11 CI failures in test_v0590.py post-merge:
1+2. test_train_annex_xi_flag_present_in_help / test_train_repro_receipt_flag_present_in_help
— Typer's Rich-renderer wraps long lines and inserts ANSI colour codes
BETWEEN the two dashes of `--annex-xi` / `--repro-receipt`, so the
literal substring match fails. Strip ANSI escape codes via regex before
asserting; also accept the bare option name as a defence-in-depth
fallback against future Rich line-wrap quirks.
3. test_default_log_path_rejects_null_byte_env — POSIX `os.environ.get` raises
`ValueError("embedded null byte")` when the env value contains a NUL
character, while Windows allows the read. Wrap the env read in
`try/except ValueError` so the function falls back to the safe default
(~/.soup/audit.jsonl) on either platform.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
6d44f0f931 |
feat(governance): v0.59.0 — CycloneDX/SPDX BOM + in-toto/SLSA-3 attest + Annex XI/XII + audit-log + repro-receipt + energy schema
Six Parts ship the procurement-floor moat — every Soup run can now emit
the formats regulated orgs demand, with no SaaS structurally able to
follow. Pure orchestration on top of v0.26 Registry + v0.34 cost
tracker + v0.56 diagnose — schema + atomic-write surface only; live
Sigstore signing, CodeCarbon hook, PDF rendering deferred to v0.59.1.
Part A — `soup bom emit` CycloneDX 1.6 ML-BOM + SPDX 2.3 AI-profile
dual emitter from any RegistryEntry. SHA-256 validation on every sha
field, license-id chain, base-model component with hash, per-artifact
file components, energy properties under metadata.properties.
Part B — `soup attest emit` in-toto v1 Statement wrapping SLSA-3
provenance v1 predicate. Stage allowlist (extract/train/eval/export/
publish), subject SHA locked to 64-hex, builder_id capped, SignatureBackend
enum with UNSIGNED live + SIGSTORE/ED25519 stubs raising NotImplementedError
with explicit v0.59.1 marker.
Part C — `soup train --annex-xi` EU AI Act Annex XI Sections 1+2 +
Annex XII Article 53(1)(d) markdown auto-doc. Top-10 domain cap,
modality breakdown, FLOPs/kWh/CO2. `_md_escape` neutralises |[](){}!<>
plus newline/CR/tab in every operator-controlled field — defends
against forged-heading + Markdown-link injection in downstream PDF/HTML
renderers (mirrors v0.29.0 model-card v2 policy).
Part D — `soup audit-log tail/rotate` HIPAA/SOC2-shaped JSONL with
PII redaction across every string field via v0.40.3 _SECRET_RE policy.
POSIX O_NOFOLLOW on append + 0o600 perms + lstat-based symlink rejection
at rotation backup path (no lexists race). SOUP_AUDIT_LOG_PATH env
override containment-checked to $HOME / $CWD / $TMPDIR.
Part E — `soup train --repro-receipt` SR 11-7-style receipt: seeds
(torch/numpy/python), kernel versions (CUDA/cuDNN/NCCL via best-effort
torch probes), GPU model + driver, OS + arch, Python version. Atomic
write, cwd-contained.
Part F — CodeCarbon hook schema + electricityMap SSRF validator with
full parity to v0.51.0 hubs.validate_hub_endpoint (scheme allowlist,
loopback-only HTTP, RFC1918 / link-local / reserved / multicast IP
rejection via ipaddress.ip_address, control-char + null-byte
rejection). PUE math + attach_energy populating BomEntry.
Cross-cutting: new paths.atomic_write_text shared TOCTOU-safe helper
centralises the v0.33.0 #22 / v0.43.0 / v0.55.0 / v0.56.0 / v0.57.0
/ v0.58.0 atomic-write pattern from four separate copies into one
single-source-of-truth (mirrors v0.40.6 / v0.53.5 peft_wiring policy).
Four review waves (python-reviewer + general-purpose security/code/tdd):
0 CRITICAL + 8 HIGH + 12 MEDIUM + 4 LOW resolved before commit.
HIGH fixes: audit-log lstat-before-write TOCTOU, O_NOFOLLOW on
append, redaction extended to host_id/operator_id/command, audit-log
env override containment, bom artifact size_bytes validation,
BomEntry attach_energy type-hint fix, default_log_path public symbol,
duplicated seeds validation removed.
Test count 9193 → 9294 (+99 net in tests/test_v0590.py; 93 pass +
6 POSIX-skipped on Windows for symlink rejection branches). v0.58.0
floor-check assertions widened from exact-match in test_v0580.py.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
168ecc0100
|
Add `--nccl` flag to `soup doctor` for multi-GPU bandwidth checks (#178)
* feat(doctor): add --nccl flag to measure and validate multi-GPU bandwidth * test(doctor): add mocked CUDA tests to verify --nccl skip and success behaviors * docs(readme): document the new --nccl bandwidth check flag for the doctor command |
|
|
|
b344aa881a |
feat(loop): soup loop CLI-first data flywheel capstone (v0.58.0)
Connects 8 existing uniques into one workflow: production traces -> preference pairs -> Eval-Gated DPO -> canary deploy -> rollback, all from a single CLI with budget guardrails and per-iteration replay. Modules (live): - utils/loop_state.py: LoopState frozen + atomic .soup/loop.yaml I/O - utils/canary_router.py: deterministic SHA-256 routing + BucketStats - utils/loop_budget.py: parse_budget_string + check_budget + UTC rollover - utils/loop_iteration.py: IterationRecord + write/read/list manifests - utils/loop_daemon.py: WatchConfig + run_once + watch daemon - commands/loop.py: init / status / pause / resume / watch / canary / replay Three review waves fixed 1 CRITICAL + 7 HIGH + 9 MEDIUM + 2 LOW total: python-review wave 1 (BucketStats lock scope + TOCTOU lstat-before-write on _check_path + init_state + NUL-byte on _bucket_for_key); code-review wave 2 (watch preserves paused / budget-skip writes no manifest / canary autoroll persisted to LoopState / route() math.ceil for sub-bucket predictability / parse_budget_string usd-only friendly error / list_iterations swallows OSError / module-top replace import); security + tdd wave 3 (_check_dir TOCTOU mirrors _check_path pattern, exact-boundary tests at _MAX_STR_FIELD=512 and _MAX_FILE_BYTES=1 MiB, bool-rejection on 4 counters, empty-string rejection on 3 optional-str). verification-loop: manual CPU smoke covering init / status / pause / resume / watch --max-iterations / canary / replay end-to-end. Notes: - ASCII arrows (->) in user-facing help text (CI test_help_output_is_ascii_safe). - Source-grep tests use Path(__file__).resolve().parent.parent for cwd- independence (defends against monkeypatch.chdir side-effects from earlier tests in the suite). - Stage callbacks ship as no-op stubs; v0.26 trace-to-pref / eval-gate / v0.30 multi-adapter deploy wiring is operator-driven via WatchConfig fields. Pre-wired versions tracked for v0.58.1. Test count: 8998 -> 9193 (+195 net in tests/test_v0580.py). Lint clean. Full repo pytest green (9105 pass + 53 skipped pre-fixes). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
76dfbb363e |
test(adapters): mock os.environ.get for null-byte/CRLF env tests (CI fix)
POSIX setenv (and Windows equivalent) reject null bytes + control chars
at the syscall boundary, so monkeypatch.setenv("SOUP_BRANCHES_DIR",
"/some\x00path") raises ValueError on every CI runner before our code
ever sees the env var.
Stub os.environ.get directly so the helper's rejection branch is
exercised exactly as it would be if the env var arrived through some
other channel (subprocess env inheritance, in-process programmatic
mutation, etc).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
8577bc2800 |
test(adapters): strip ANSI before help-output substring asserts (v0.57.0 CI fix)
Rich-wrap-CI workaround — same fix pattern as v0.55.0 / v0.56.0:
CliRunner output contains ANSI color escapes that break literal
'--top-k' in output substring matches because Rich renders option
names as -\x1b[0m\x1b[1;36m-top-k.
Adds _ANSI_RE + _strip_ansi() helper to each of the 4 test files
(test_v0570_part_{a,b,c,d}.py) and routes every help-output
substring assertion through it.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
7da82d355a |
feat(adapters): v0.57.0 — git for LoRA (diff, merge, blame, branch)
Ships "soup adapters {diff,merge,blame,branch,checkout,branches}" — git-shaped
UX for LoRA adapter management. Pure-numpy math (no torch); TOCTOU defence on
every path; atomic writes throughout.
Part A — adapters diff (utils/adapter_diff.py, ~270 LOC):
Per-layer ΔW Frobenius norm + relative drift; effective-rank delta via
SVD entropy; top-K changed projections; JSON/Markdown/table output.
Part B — adapters merge (utils/adapter_merge.py, ~280 LOC):
Four strategies — linear (weighted avg), ties (Yadav et al. — trim/
elect-sign/disjoint), dare (Yu et al. — drop+rescale, deterministic),
svd (low-rank reconstruction). MergeReport.verdict='UNKNOWN' is a stub;
live canary verdict via v0.55 eval gate ships in v0.57.1.
Part C — adapters blame (utils/blame.py, ~190 LOC):
Leave-one-out plan emitter + budget tracker (parse_budget mirrors v0.48.0
data_mix idiom). Per-shard work table + feasibility check. Live ablation
runner raises NotImplementedError v0.57.1 (mirrors v0.27.0 / v0.50.0 /
v0.56.0 stub-then-live pattern).
Part D — adapters branch / checkout / branches (utils/adapter_branch.py,
~230 LOC):
SHA-256 snapshot pointers under ~/.soup/branches/ (SOUP_BRANCHES_DIR
override, $HOME/$CWD/$TMPDIR-bounded). Drift detection on checkout —
refuses restore when source SHA != snapshot SHA. CRLF/null-byte
rejection on env override (mirrors v0.51.0 hub-endpoint policy).
5-agent review-fix wave (1 CRITICAL + 9 HIGH + 11 MEDIUM + 4 LOW):
- TIES tied-sign defaults to +1 (np.sign(0)==0 would silently zero all
tied parameters)
- load_branch / delete_branch reject symlinks via os.lstat + S_ISLNK
before read/unlink
- merge output safetensors + adapter_config.json atomic writes with
symlink target rejection at output path; source config size-capped
at 256 KB
- compute_adapter_diff weights-file path symlink-rejected via lstat
BEFORE is_file() (defends against .safetensors -> /etc/passwd escape)
- _count_dataset_rows opens via realpath captured at containment check
(closes TOCTOU window)
- diff --output write is atomic (tempfile + os.replace)
- SOUP_BRANCHES_DIR rejects every C0 control char, not just null
- SUPPORTED_STRATEGIES is now frozenset (matches v0.41.0+ allowlist
policy); STRATEGY_ORDER tuple preserved for canonical iteration
- 5× pytest.raises(Exception) tightened to FrozenInstanceError
- 2 zero-assertion Part D tests converted to real assertions
- Added: bool base_model rejection, top_k boundary 1/201, density=1.0
inclusive bound, inf weight rejection, tied-sign positive default,
bool False for num_shards/budget_seconds, POSIX symlink rejections
for diff weights / merge output / load_branch / delete_branch,
no-top-level-torch source-grep guards, traversal delete_branch.
Plus v0.56.0 follow-up: test_v0560.py version-floor tests widened from
exact-match to floor-check (matches v0.51.0 / v0.54.0 idiom — every
subsequent release would otherwise edit this one line).
Test count: 8849 → 8998 (+149 net in 4 new files; 4 POSIX-only symlink
tests skipped on Windows). Full suite green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
a3810823d1 |
refactor(train): harden diagnose-gate rank guard from PR #169
PR #169 wired LOCAL_RANK==0 guard on _run_diagnose_gate so distributed launches only run the gate on one worker per machine. Two minor polish items on top of the merged version: - Wrap the int() parse in try/except ValueError. A malformed LOCAL_RANK (garbage value from a misconfigured launcher) would previously crash the post-training gate. Falling back to True is safer than silently skipping the gate -- over-running is recoverable, under-running hides failures. - Expand the docstring to explain why we use LOCAL_RANK (per-machine) rather than RANK (global): the gate reads the local output_dir, so one gate per machine is the right granularity for typical single- machine multi-GPU runs. Documents the choice for future readers. - Add a focused test (test_diagnose_gate_handles_malformed_local_rank) asserting the safe fallback path. |
|
|
|
4c2a578ac0
|
Guard diagnose gate on distributed worker ranks (#169)
Co-authored-by: mzl2233 <mzl2233@users.noreply.github.com> |
|
|
|
7d81496c69 |
test(diagnose): strip ANSI before help-output substring asserts (v0.56.0 CI fix)
Rich's CliRunner output on CI carries ANSI escape codes that split long option names like `--badge` and `--diagnose-gate` across colour-reset boundaries (`-\x1b[0m\x1b[1;36m-badge`), breaking naive `"--badge" in result.output` substring checks. Same fix pattern as v0.55.0 CI hotfix. Failures: tests/test_v0560.py::TestCli::test_diagnose_help and TestTrainDiagnoseGate::test_help_lists_flag on all 9 CI matrix cells. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
72189baba7 |
feat(diagnose): soup diagnose — post-training model report card (v0.56.0)
Six failure-mode probes (forgetting / refusal / format / mode_collapse / memorization / contamination) + FailureReport frozen dataclass + SVG badge + soup train --diagnose-gate. Same OK/MINOR/MAJOR taxonomy as v0.26.0 Quant-Lobotomy. - soup diagnose <run-id> [--evidence|--output|--badge|--attach-to-registry] - soup train --diagnose-gate <evidence.json> refuses MAJOR runs - diagnose_report added to registry._VALID_KINDS Review wave (4 agents): 4 HIGH + 8 MEDIUM + 2 LOW addressed — atomic+TOCTOU-safe badge write, typer.Exit (not sys.exit), realpath containment, evidence size cap, contamination combined-complexity cap, ReDoS probe, extras null-byte sanitisation, extract_row_text centralisation, tokenize delegates to _eval_text. Test count: 8676 -> 8849 (+123 in test_v0560.py + 50 net adjustments). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
04c504e761 |
test(eval): strip ANSI before help-output substring asserts (v0.55.0 CI fix)
3 macOS CI failures from the v0.55.0 push — Rich wraps option names with ANSI escapes when the terminal is narrow (macOS CI runners default to a smaller width than Linux/Windows), so substring searches like `"--goal" in result.output` fail because the actual output contains `\x1b[1;36m-\x1b[0m\x1b[1;36m-goal\x1b[0m`. Project precedent: v0.53.5 / v0.53.6 / v0.53.8 / v0.53.9 all hit the same pattern; tests/test_auto_tuning.py and tests/test_eval_platform.py already ship `_ANSI_RE` + `_strip_ansi` helpers. Failures fixed: tests/test_v0550.py::TestCLIPlumbing::test_eval_design_help tests/test_v0550.py::TestEvalAgainst::test_against_help tests/test_v0550_followups.py::TestEvalAgainst::test_against_cli_help_lists_flag Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
58d7d510bf |
feat(eval): soup eval design — derive evals from data (v0.55.0)
Trainer libraries help you RUN evals — none help you DEFINE them.
v0.55.0 closes that gap with 5 new subcommands:
- soup eval design <data> --goal "..." → goal-conditioned EvalDesign
(TF-IDF salience + scorer dispatch)
- soup eval discover <data> → held-out canaries + memorization probes
(farthest-first Jaccard clustering)
- soup eval lock + soup eval coverage → SHA-256-checksummed artifact +
gap analysis vs v0.54.0 task taxonomy
- soup eval gate-install --baseline R → pre-push regression gate
(paired-bootstrap CI, shlex.quote)
- soup eval against B --candidate C → run-vs-run paired-bootstrap CI
Heuristic / CPU-only — no GPU required. Lazy imports across all 6 new
modules so `soup --help` startup remains < 200 ms.
New registry artifact kinds: eval_suite, canaries.
New tracker accessor: ExperimentTracker.get_metric_series(run_id, metric).
Security policy (all atomic-write + read surfaces):
- cwd containment via os.path.realpath + commonpath
- unconditional os.lstat + stat.S_ISLNK rejection (TOCTOU defence)
- atomic write via tempfile.mkstemp + os.replace
- shlex.quote for shell-script generation (NO hand-rolled escape)
- MappingProxyType on every registry / metric / scorer map
- frozen dataclass on every public return type
- bool-as-int rejection on every numeric input
- DoS caps: 10k subsample for TF-IDF + clustering hot paths
Review-fix coverage across 4 agents (python / security / code / tdd):
0 CRITICAL + 7 HIGH + 11 MEDIUM + 6 LOW resolved before commit.
Tests: 8571 → 8676 (+105 net).
Lint: ruff clean.
Smoke: every CLI command + every failure mode exercised in /tmp/soup_smoke.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
9e18643ea0 |
test(advise): fix cross-platform CI failures in test_v0540
Two failures on ubuntu/macos/windows × py3.9/3.11/3.12 after v0.54.0 push: 1. test_env_null_byte_falls_back: monkeypatch.setenv can't set raw NUL into the OS env layer (POSIX execve + Win32 SetEnv both refuse). Switched to a temporary `advise_history.os.environ` swap so the helper's defence-in-depth NUL guard is still exercised without going through the C env layer. 2. test_default_missing_data: Click 8.0–8.1 returns rc=0 on `no_args_is_help=True` invocations; Click 8.2+ returns rc=2 (the "missing command" convention). CI runners had the newer Click; dev box had the older. Accept both renderings. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
600686cd70 |
feat(advise): soup advise — pre-flight decision (v0.54.0)
`soup advise <data.jsonl> --goal "..."` returns one of PROMPT_ENG / RAG / SFT / DPO / GRPO with a confidence, reason, and reverse-when criterion BEFORE the user spends 8 hours on a GPU. Layer above autopilot — autopilot picks hyperparams AFTER the training decision; advise picks the training decision itself. Three Parts: - Part A: Verdict engine — TASK_CATEGORIES + CHOICES allowlists, frozen Verdict / DatasetProfile / ROIEstimate dataclasses, pure- Python classify_task + compute_dataset_profile + build_verdict rubric (DPO / GRPO floor 500 / PROMPT_ENG floor 50 / RAG / SFT). - Part B: Probe runner — synth_probe_baselines + synth_probe_lora_delta heuristic stubs with forward-compat model/device/lr/timeout_seconds kwargs (v0.54.1 lifts to live model loading per stub-then-live cadence used by v0.27.0 MII / v0.37.0 multipack / v0.50.0 GRPO Plus). - Part C: Cross-project learning — ~/.soup/advise_history.jsonl with cross-process file locking (fcntl on POSIX, sidecar <path>.lock + msvcrt on Windows). `soup advise compare` reads history; env override SOUP_ADVISE_HISTORY_PATH containment-checked to $HOME / $CWD / tempdir (mirrors v0.36.0 SOUP_BATCH_CACHE_PATH policy). CLI: Typer subcommand group `run` / `explain` / `compare` plus argv preprocessor in cli.py that maps `soup advise data.jsonl` → `soup advise run data.jsonl`. Scoped to argv[1] == "advise" only (code-review HIGH fix — defends against rewrites when an unrelated arg contains the literal string "advise"). Schema: AdviseConfig (goal / probe / record) field on SoupConfig honors the plan's cross-cutting bullet. Security: cwd-containment + os.lstat + S_ISLNK symlink reject on every path input; atomic writes via tempfile.mkstemp + os.replace on scratch + history; per-line 64 KB cap + 16 MiB file cap on history reads; bool / finite / NUL / oversize guards on every public input; Rich markup escape on user-controlled output. Reviewed by python / code / security / tdd / architect agents — every finding fixed before commit (0 CRITICAL + 5 HIGH + 7 MEDIUM + 4 LOW). Test count: 8400 → 8571 (+136 in tests/test_v0540.py, +35 net adjustments to v0.53.x version-pin assertions to forward-compat >=). Note: Windows CRLF / LF warnings during stage are .gitattributes- governed and benign. CI runs on ubuntu-latest / windows-latest / macos-latest × Python 3.9 / 3.11 / 3.12. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
cfdabf2b3b |
feat(v0.53.11): GRPO Plus finish + preference live
Closes v0.50.1 (#123, #126, #127), v0.49.1 (#119), v0.40.1 (#68). #123 — live math kernels for 6 GRPO variants (gspo/dapo/dr_grpo/bnpo/ two_sided/rft) + `_GRPOTrainerVariant` HF Trainer subclass via `make_grpo_trainer_variant` factory. Variant compute_loss reads kernel inputs FIRST (no double-forward); falls back to super() only on missing attrs. Case-insensitive variant normalisation before lru_cache. #126 — PRMTrainerWrapper + `_PRMTrainer` HF Trainer subclass with real compute_loss (gather hidden states at step_positions -> reward_head -> MSE via compute_prm_loss). Dataset wrapped in datasets.Dataset.from_list for HF Trainer compatibility. Bool-before-isinstance guard on batch_size. #127 — GRPOStabilityCallback inherits transformers.TrainerCallback (lazy), live EMA ref-model update in on_step_end with strict=True + fallback-to-strict=False-with-WARNING on key mismatch (silent corruption defence). math.isfinite guard on alpha. #119 — LongLoRA forward override via LongLoRAForwardOverride context manager with idempotent install (_soup_longlora_patched marker prevents re-entry double-wrap), 256-char class name cap on regex match, restore on __exit__ AND on exception. #68 — true per-batch weighted-sum preference combine reading policy/ref logps from TRL inputs + each compute_*_term kernel + combine_losses. Explicit None checks on trainer attrs (no `or` on possibly-tensor), DEBUG log on per-term skip. Review fixes from 4 agents (python/code/security/tdd): 10 HIGH + 8 MEDIUM + 7 LOW — see CLAUDE.md v0.53.11 entry for the full list. Test count: 8330 -> 8400 (+75 in test_v05311.py: 54 initial + 21 review-fix coverage gaps). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
76f033a6ff |
feat(v0.53.10): Quick wins + packaging + UX wiring
7 issues closed: - #150 [mix] pyproject extra bundles scikit-optimize so `soup data mix --optimize` runs the Bayesian loop instead of the v0.48.0 Dirichlet fallback; new describe_default_optimizer() helper labels the active backend without paying skopt's import cost. - #113 [data-pro] extras (langdetect + presidio-analyzer) with lazy fall-through helpers in utils/data_score (broader language coverage + Presidio entity recognition on top of the v0.47.0 regex baseline). Llama-Guard-3-1B documented as a manual recipe (license + size). - #154 SOUP_POSTHOG_KEY / SOUP_POSTHOG_ENDPOINT env override via sentinel-based explicit-vs-env precedence; HTTPS-only + RFC1918/link-local rejection on the endpoint; null-byte / control-char / >256-char rejection on the key. - #152 --hub flag plumbed on chat / serve / infer / merge / export / push via shared utils/hubs.apply_hub_to_cli_model + prefetch_model_from_hub helpers; push uses upload_repo (skips HF-specific Collections + model-card auto-render on non-HF hubs). - #153 `soup data download --hub modelscope|modelers` live SDK (lifts the v0.53.8 advisory-only path); friendly ImportError advisory when the SDK is missing. - #155 Web UI Tool Outputs panel — `loadToolOutputs` polls /api/tool-outputs every 3s; XSS-safe DOM-built table (textContent per cell, no innerHTML for user-controlled fields); Bearer token threaded via the v0.53.9 window._authToken bootstrap. - #156 SoupTrainerCallback.on_step_end records tool_calls counts from kwargs['inputs'] into the global tool buffer. Best-effort (# noqa: BLE001 per project policy — training must never crash). 13 review-fixes applied (4 HIGH / 5 MEDIUM / 4 LOW): - HIGH PostHog explicit-endpoint precedence sentinel - HIGH absolute path leak in local_path advisory reduced to relpath - HIGH Rich markup escape on base / local_path / cache_dir - HIGH callback # noqa: BLE001 per project policy - MED `import time` moved out of try block - MED oversize key + explicit-empty key rejection tests - MED source-grep regression guards (advisory-removal, helper imports across 5 non-push commands) - MED `prefetch_model_from_hub` outside-cwd cache_root rejection - LOW empty-list + bool-True tool_calls no-op tests - LOW push.py uses upload_repo + validate_hub_name regression guard Test count: 8285 -> 8330 (+45 in tests/test_v05310.py). Full suite green; ruff clean; on Win+Py3.10. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
e8422660a5 |
fix(tests): strip ANSI codes before help-output substring asserts (v0.53.9)
CI's Rich pipeline emits styled output that splits `--vocab-size` into multiple ANSI-bracketed spans (e.g. `\x1b[36m-\x1b[0m\x1b[36m-vocab\x1b[0m\x1b[36m-size\x1b[0m`), breaking naive `"--vocab-size" in result.output` checks. Locally Rich auto-detects non-TTY and skips the codes, so the regression only shows on CI (ubuntu/macos/windows × 3.9/3.11/3.12). Fix: small `_plain()` helper using `re.sub(r"\x1b\[[0-9;]*m", "", ...)` applied to the 7 failing assertions. Same approach already used in several other v0.5x test modules. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
bde9d149a0 |
feat(v0.53.9): Live Dashboard + UX + Bench + Standalone CLIs
Eight features that close out the v0.44.x live-monitoring deferrals plus a long tail of standalone CLI wins: - #94 /api/train/stream async SSE with per-subscriber cursor + JS EventSource consumer; SoupTrainerCallback pushes TrainEvent on each on_log. - #95 soup ui --public derives LAN IP via SOCK_DGRAM connect-trick, prints scannable QR; --auth-token override; SPA bootstrap hydrates window._authToken from ?token= + sessionStorage and cleans the URL via history.replaceState; CORS regex auto-widens to loopback + RFC1918 in public mode; set_auth_token rotation race fixed via threading.Lock. - #98 soup serve --reasoning-parser strips <think>...</think> (and OpenThinker tags); pre-compiled regex with marker-token fast-path + 1 MiB cap + leading-newline-only strip. - #100 ToolOutputsBuffer global singleton + /api/tool-outputs JSON endpoint; best-effort observation hook in callback.on_log. - #15 soup tokenizer train: BPE training CLI with raw-path lstat symlink rejection, 50 MiB total / 8 KiB per-line caps, post-mkdir output-dir re-check, --special-token NUL/oversize dedup, vocab bounds [256, 200000]. - #26 soup bench --p50 --p95 renders extra per-prompt tail-latency Rich table; --prompts-file gains symlink rejection. - #28 soup bench --backend auto: MLX weights.npz probe (per-entry lstat) -> config.json model_type keyword -> transformers fallback; SOUP_BENCH_BACKEND env hint. - #12 examples/synthetic_workflow.{md,yaml} end-to-end walkthrough. Review fixes: 0 CRITICAL + 11 HIGH + 14 MEDIUM + 9 LOW across the python / code / security / tdd review agents. Notable HIGH: - QR token now consumed by SPA (was unreachable previously). - set_auth_token rotation lock-protected, 8-thread stress tested. - Tokenizer input + output symlink TOCTOU defence on raw path. - SSE generator switched to async (asyncio.sleep) for non-blocking multi-subscriber operation. - CORS regex for --public LAN mode (the old fixed allowlist of http://0.0.0.0:port never matched a real Origin header). - _has_mlx_weights per-entry lstat so a symlinked weights.npz can't trigger MLX dispatch. Test count: 8257 -> 8285 (+57 in tests/test_v0539.py, minus the relaxed v0.53.8 version-pin asserts in tests/test_v0538.py). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
53bb82afeb |
fix(v0.53.8.1): hatch artifacts directive — fix PyPI duplicate-filename 400
v0.53.8 PyPI publish failed with: 400 Invalid distribution file. ZIP archive not accepted: Duplicate filename in local headers Root cause: `[tool.hatch.build.targets.wheel.force-include]` shipped `soup_cli/data/_fixtures/` AND `packages = ["soup_cli"]` recursed into the same path, so both the wheel and sdist contained each JSONL twice. Fix: switch from force-include to `artifacts = [...]` which adds non-Python files to the existing package tree exactly once. Standard hatchling pattern for shipping data files inside an already-packaged directory. Version bumped to v0.53.8.1 (patch) — same code surface, just a build config fix. v0.53.8 GitHub release remains as the feature changelog; PyPI ships under v0.53.8.1. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
7b98982ab6 |
fix(test): v0.53.8 CI hardening — strip ANSI + use _repo_root() helper
Five v0.53.8 CI failures (ubuntu/macos × py3.9/3.11/3.12):
1. test_help_lists_hub_flag — Typer's Rich-rendered help wraps long
option help across ANSI box-drawing lines; "--hub" appears as
"│ --\nhub" in the CI terminal renderer. Strip ANSI + collapse
whitespace before asserting.
2-5. test_pyproject_version / test_*_extra_present / test_force_include
— used `Path("pyproject.toml")` (relative to cwd). CI invokes pytest
from a different cwd than the repo root on at least one matrix
entry. Switched to a `_repo_root()` helper that derives from
`__file__` (matches v0.43.0 Part D demo_bundles approach).
Local re-run: 66/66 v0.53.8 tests pass after the fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
6d2170c4f3 |
feat(v0.53.8): Remote data + Hubs + Trackers (wave 2) — 6 features
- #85 fsspec live loaders — data/loader.py routes the v0.42.0 fsspec scheme allowlist (s3:// / gs:// / gcs:// / az:// / abfs:// / abfss:// / oci://) through fsspec.open with validate_remote_uri containment BEFORE connection. Friendly Rich panel names the pip install advisory when the backend SDK is missing. Threads data.streaming + data.buffer_size. Row count capped at 1M. - #130 Hub dispatcher live — utils/hubs.download_repo() and upload_repo() lazy-import per backend (huggingface_hub / modelscope / openmind_hub). Shared _validate_repo_id_shape (bool / null-byte / leading-slash / .. / control-char / oversize) + cwd containment on local_dir / folder_path. commands/train.py pre-fetches non-HF base into .soup_hub_cache/ (sanitised slug, idempotent on resume, cfg.base updated via model_copy). soup data download --hub flag plumbed. Multi-command rollout for chat / serve / infer / merge / export / push tracked for v0.53.9. - #89 [trackers] pyproject extra bundles mlflow / swanlab / trackio; tracker_missing_dep_message surfaces a friendly pip install advisory via importlib.util.find_spec (non-executing probe). - #90 utils/trackers.send_telemetry_payload — opt-IN via SOUP_TELEMETRY=1; lazy httpx; 1s hard timeout; HTTPS-only with SSRF re-validation (mirrors v0.51.0 hub endpoint policy); silent-fail on every exception. - #93 Fixtures migrated to soup_cli/data/_fixtures/ — zipapp / namespace-package safe via [tool.hatch.build.targets.wheel.force-include]; _bundle_source_path falls back to examples/data/ for editable installs. - #69 utils/hf_space.detect_space_sdk(requirements_text) — picks "streamlit" / "gradio" from the rendered requirements.txt; closes the v0.40.2 known limitation that custom Spaces always defaulted to gradio. Wired into commands/deploy.py. Review pass: python-review + code-review + security-review ran in parallel; 16 findings fixed (3 HIGH + 8 MEDIUM + 5 LOW). Highlights: cwd-containment on local_dir/folder_path, Windows ..\ traversal defence on .soup_hub_cache slug, Pydantic model_copy(update=...) instead of attribute mutation, idempotent pre-fetch via cache probe, 1M-row cap on remote materialisation, SSRF re-validation on telemetry endpoint override, 256 KB cap on detect_space_sdk input, modelscope.push_model commit_message kwarg removed (would TypeError at runtime), find_spec instead of __import__ to avoid swanlab side-effects. Test count: 8162 -> 8257 (+66 in tests/test_v0538.py + 29 net adjustments). Lint clean. CPU smoke: version, --help, load_config_from_string with hub: modelscope passes; mlx + non-HF rejected; data download --hub modelscope advisory rendered; detect_space_sdk live on real requirements.txt bodies; package-data fixtures resolve from soup_cli/data/_fixtures/. v0.53.7 known limitation #1 (bash 501 marker) bumped to v0.53.9. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
a26511f868 |
fix(test): v0.53.7 — symlink test accepts either rejection branch
test_load_jsonl_rows_rejects_symlink: when the symlink target is outside
cwd, is_under_cwd's realpath resolution catches it before the lstat check
fires. Both rejections are valid security guards; broaden the regex to
match either error message ("symlink" or "under cwd").
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
9c9f962676 |
fix(test): v0.53.7 CI hardening — importorskip(fastapi) + Arrow target setup
CI run 25805598433 failed across all 9 OS×Python cells: - TestToolEndpointsLive / TestAnthropicMessagesStreaming / TestReviewFixesVllmAnthropicLive ModuleNotFoundError: fastapi (CI does not install [serve] extra) → autouse fixture pytest.importorskip() - test_load_pretokenized_dataset_rejects_symlink: load_pretokenized_dataset called datasets.load_from_disk on the symlink target before the lstat check ran → moved the lstat + S_ISLNK check to the entry of the helper so symlinks reject before any load attempt - test_redact_exc_message_handles_windows_paths: hardened _redact_exc_message to strip both POSIX absolute paths and Windows-style paths regardless of host platform Local pytest tests/test_v0537.py: 112 passed, 7 skipped. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
18d8b36114 |
feat(v0.53.7): Data Forge + Pipeline live (wave 1) — 11 items
Closes v0.47.0 deferrals (#111, #112), community QA (#75), v0.42.1 wave 1 (#87, #86, #88), and 5 v0.53.6 stub-to-live carry-overs (#102 vLLM parity + SSE streaming, #103 tool HTTP endpoints, #105 instantiate_trainer_plugins, #106 run_recipe DAG runner). - #88 markdown ingest heading split (split_markdown_by_headings) - #112 soup data decontaminate --benchmark-file (cwd-contained operator corpus + symlink rejection) - #87 prompt_strategy live resolver (resolve_prompt_strategy + lru_cache, importlib-based; per-row hook in sft_format.py) - #86 soup data preprocess AOT tokenize (atomic Arrow shard write + cache metadata sidecar; SFT + Pretrain wrappers short-circuit on format='pre_tokenized' + tokenized_path with cache-hash gate) - #111 forge --judge-provider {ollama,anthropic,vllm} live (lazy v0.20.0 providers; SSRF parity) - #75 QA log entry for synth-data provider manual smoke - #106 run_recipe LIVE for 6 NODE_KINDS (seed / llm_text / code / judge / validator / sampler); atomic checkpoint via tempfile.mkstemp; resume rehydrates predecessor outputs from per-node sidecar JSONL; lstat-on- raw-path symlink rejection (v0.33.0 #22 TOCTOU parity); failed_reason path-redacted - #105 instantiate_trainer_plugins LIVE for cce_plugin / grokfast / spectrum / llmcompressor / sonicmoe / math_verify (lazy imports, friendly pip-install advisory on missing dep) - #103 POST /v1/tools/python + /v1/tools/web_search LIVE (Bearer auth gate, deny-by-default domain allowlist, 5s timeout, 5-result cap). bash reverted to HTTP 501 — security review caught /bin/sh -c child escapes RLVR sandbox's OS-level isolation; deferred to v0.53.8. - #102 vLLM /v1/messages parity LIVE on both backends; CORS loopback-only - #102 Anthropic-shape SSE streaming on /v1/messages LIVE; Cache-Control: no-store; CRLF/NUL/oversize strip on model+msg_id (header injection) Review fixes (1 CRITICAL + 11 HIGH + 17 MEDIUM + 10 LOW from python-reviewer + code-reviewer + security-reviewer + tdd-guide) all addressed in this commit. Test count: 8051 → 8162 (+111 in tests/test_v0537.py). Test files: 189 → 190. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
c9e32f4e19 |
fix(test): v0.53.6 CI hardening — importorskip(fastapi) + strip ANSI
CI failures on `tests/test_v0536.py`:
1. `ModuleNotFoundError: No module named 'fastapi'` on Ubuntu/macOS runners
— fastapi is in the [serve] extra, not [dev]. Added
`pytest.importorskip("fastapi")` to all 7 tests that use TestClient
(matches the pattern in `_build_app`).
2. `--execute` / `--output` / `v0.53.7` substring checks failed because
CI terminals render Typer/Rich help text with style spans (`-` and
`-execute` end up in separate `\x1b[...]m` runs). Added `_strip_ansi`
helper + wrapped 5 substring checks. Same fix pattern as the v0.53.5
`--live` CI fix on test_v0535.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
70326f1aa7 |
feat(v0.53.6): plugin callback + Anthropic /v1/messages + n-gram spec + 3 stubs
Ships v0.53.6 "Plugin + Agent + Anthropic API" — 6 features, 3 live and 3 stub-then-live with v0.53.7 markers. LIVE - #101 SoupPluginCallback fans HF Trainer events to enabled Soup plugin hooks (pre_train/post_train/pre_step/post_step). Hook exceptions swallowed at WARNING. Hook snapshot collected once and passed to ctor (race-free per code-review fix). Wired into all 13 transformer-backend trainers via utils/peft_wiring.attach_plugin_callback. - #102 POST /v1/messages on transformers backend reuses the v0.45.0 anthropic_messages converter + existing chat handler. Streaming -> 501. Validation errors -> generic 'Invalid request' body, details at DEBUG (security-review redaction fix). - #104 n-gram speculative decoding: NgramSpecConfig.num_draft_tokens threaded through model.generate(prompt_lookup_num_tokens=N). Mutually exclusive with assistant_model. STUB-THEN-LIVE (v0.53.7) - #103 /v1/tools/{python,bash,web_search} return HTTP 501. - #105 instantiate_trainer_plugins validates then NotImplementedError. - #106 run_recipe + 'soup data recipe --execute --output <dir>' with CLI-side is_under_cwd containment before the live runner. Tests: 7998 -> 8051 (+53 in tests/test_v0536.py). Lint clean. Three review agents (python / code / security); every HIGH/MEDIUM fixed. TDD coverage gaps closed (ngram-None regression, max_tokens cap, run_recipe boundaries, console-print failure swallow). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
8bd67f697f |
fix(test): strip ANSI before --live substring check (v0.53.5 CI fix)
Rich splits the `--live` token across ANSI colour codes on narrow CI terminals (`-\x1b[0m\x1b[1;36m-live`), so the raw-output substring assertion failed on macOS/Windows runners but passed locally on a wide terminal. Strip ANSI escapes before the check — matches how earlier test files (e.g. test_v0402_part_b) handle Rich-coloured `--help` output. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
bc1f060da9 |
feat(adaptive): v0.53.5 — Adaptive Training (BETA → stable)
Closes #114, #115, #116, #117, #118, #17 — lifts every v0.48.0 BETA deferral and adds the deepseek-v3-reasoning recipe. - #114 DynamicCurriculumCallback live (TrainerCallback + all_reduce + JSONL) - #115 curriculum_dynamic schema gate widened to 13 transformer trainers - #116 soup data mix --live runs real proxy soup train subprocesses - #117 skopt.Optimizer(GP) wrapped behind OptimizerProtocol - #118 MixOptimizationReport.elapsed_seconds excludes failed candidates - #17 deepseek-v3-reasoning GRPO recipe Test count: 7935 → 7998 (+63). 4 review agents (python/code/security/tdd) ran sequentially; every CRITICAL→LOW finding fixed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
89b06efb8c |
feat(longctx): v0.53.4 — Long Context + Architecture
Six closes lifting the v0.49.0 LongLoRA hardening + v0.41.0 LLaMA Pro deferred stubs, plus a UX upgrade to the CUDA-OOM friendly message: - #11 utils/errors.py: OOM hint now names --batch-size / --grad-accum - #122 flash_attn.is_flash_attn_v3_available() + LongLoRA+FA3 schema reject - #120 LongLoRA arch allowlist expansion (Mistral / Qwen / Phi); Mixtral intentionally excluded (regex matches the bare 'mistral' token only) - #121 apply_long_context_config auto-detects 'llama3' when caller passes rope_scaling_type=None and the model config carries a Llama 3.1 rope_scaling block - #83 block_expansion.expand_model_blocks LIVE (deepcopy last-N blocks, zero-init residual projections, append, bump num_hidden_layers) + apply_llama_pro_freeze + shared apply_block_expansion_if_configured helper wired into SFT + Pretrain (mirrors v0.40.6 peft_wiring centralisation policy so SFT and Pretrain stay in lock-step) - #74 HF push surface QA — test plan recorded in tests/qa/v053_qa.md; live execution against a private HF repo deferred to a credentialed contributor Review pipeline (python / code / security / tdd agents) ran; every CRITICAL -> LOW finding addressed: - bool-first guards in _check_model_name (defends against int subclass) - is_supported_longlora_arch defensive non-string surface (returns False, never raises) matching v0.53.3 is_known_vlm_base policy - _truncate_for_message(value, limit=64) bounds the base echo in LongLoRA error messages (security MEDIUM, mirrors v0.34.0 crash.py) - null-byte + non-string TypeError guards on validate_longlora_compat task / backend params (matches v0.50.0 validate_long_context_grpo_compat) - _get_layers_module uses explicit `is None` not falsy shortcut (defends against nn.Module.__bool__ overrides on subclasses) - _zero_init_block_residual returns bool + warnings.warn when neither standard projection matches the cloned block (non-Llama-shaped arches still train but lose the LLaMA Pro identity-init guarantee) Test count: 7879 -> 7935 (+56 net; +49 in new tests/test_v0534.py). Lint clean. CPU smoke verified on a real transformers.LlamaForCausalLM: 4 -> 6 layers, down_proj + o_proj actually zeroed on PyTorch tensors, old blocks frozen + new blocks trainable, forward pass finite. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
29f875f6ef |
feat(grpo): v0.53.3 — grpo_fp16 routing + vision-GRPO VLM base probe
Two surgical fixes from the v0.50.0 GRPO Plus deferred-stub family land: - #128 grpo_fp16 routing: GRPOTrainerWrapper._build_precision_kwargs returns {fp16, bf16} per (device, grpo_fp16) matrix (CPU/MPS/XPU → both False, CUDA + grpo_fp16=True → fp16/!bf16, default CUDA → legacy bf16). SoupConfig._validate_grpo_fp16_amp_exclusive rejects the silent-mutex combo with auto_mixed_precision=True; short-circuits when task != 'grpo' so the v0.50.0 task-gate diagnosis fires first. - #129 vision-GRPO base probe: KNOWN_VLM_REGEX covers 10 VLM families (Qwen2-VL/Qwen2.5-VL/QVQ/Pixtral/InternVL/Llama-3.2-Vision/LLaVA/ MiniCPM-V/Idefics/ShareGPT4V/Fuyu) with word-boundary anchors; is_known_vlm_base returns False (never raises) on bad input; validate_vision_grpo_compat now accepts optional base kwarg with 64-char error-message truncation. YAML pairing vision_grpo: true with a non-VLM base is rejected at schema load with a friendly families listing instead of a cryptic runtime AttributeError. Scope: 4 larger v0.53.3 items (#127 stability callback, #123 GRPO variant losses, #126 PRMTrainerWrapper, #68 multi-objective preference live combine) are scope-deferred to v0.53.4 — each warrants its own focused release per the v0.40.x stub-then-live cadence. Tests: 7842 -> 7879 (+37 in tests/test_v0533.py). Four review agents (python/code/security/tdd) ran; every HIGH/MEDIUM/LOW finding fixed (task-gate priority short-circuit, MPS branch documented, 64-char error truncation, QVQ regex coverage, 512-byte boundary test). Two pre-existing v0.50.0 Part E tests migrated `base: test-llama` -> `base: Qwen/Qwen2-VL-7B-Instruct` to clear the new probe. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
2292e81c3f |
feat(modality): v0.53.2 — lift Modality II stubs (distill + classifier + EBFT/GDPO + reasoning_effort)
Closes #132, #133, #135, #137. Records #71 ONNX QA (partial — tiny-gpt2 PASS, TinyLlama-1.1B blocked by host RAM during onnx.load post-process). New trainer wrappers: - DistillTrainerWrapper (soup_cli/trainer/distill.py) — student + frozen teacher, KL/JS divergence kernels scaled by T**2, device-bridge for HF Trainer auto-CUDA promotion, DataCollatorForSeq2Seq for variable-length loss-masked rows, separate trust_remote_code resolution per model. - ClassifierTrainerWrapper (soup_cli/trainer/classifier.py) — single/multi label sequence classification, 1024-entry multi-label cap, label_names string-to-int resolution. Routes classifier / reranker / cross_encoder. Live loss kernels: - apply_ebft_loss (structured / strided) + attach_ebft_compute_loss (SFT) - apply_gdpo_loss (standard / length_normalized / margin) + attach_gdpo_compute_loss (DPO). Both attach hooks idempotent. Prompt-format wiring: - apply_reasoning_effort_prefix injects gpt-oss <|reasoning_effort|>{low,medium,high}<|/reasoning_effort|> header. - build_assistant_only_labels(train_on_eot=True) keeps EOT/EOS unmasked. Bugs surfaced + fixed during Wave 3 CPU smoke (regression guards in tests): - Distill collator did not pad pre-tokenised labels (variable-length crash) - Distill compute_loss device-mismatch when HF Trainer auto-promoted student to CUDA while teacher stayed on CPU. Tests: 7722 -> 7842 (+120 in test_v0532.py). 5 review agents run; every CRITICAL/HIGH/MEDIUM/LOW finding fixed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
2ea26df5ea |
fix(v0.53.1): CI green — inspect click params directly instead of asserting on Rich --help output
The three CLI tests added in the previous commit (test_help_lists_measure_flag, test_save_format_help_lists_flag, test_torchao_help_lists_quant_config) assumed the literal option name (e.g. `--measure`, `--save-format`, `--quant-config`) would appear contiguously in CliRunner-captured output. On CI runners the terminal defaults to 80-col and Rich wraps long option names across lines, splitting the literal string. Fix: walk `typer.main.get_command(app).params` and collect every `opt` + `secondary_opt` into a set, then assert membership. The test now verifies what we actually care about (the option is registered) without depending on Rich's wrapping behaviour. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
725696b1da |
feat(v0.53.1): Quant Menu II + Export pipeline live
Lift six v0.53.0 deferred stubs from NotImplementedError to live wiring: - #82 autopilot pre-quantized base detection utils name regex over gptq/awq/aqlm/eetq/fp8/mxfp4 with word-boundary anchoring + HQQ Nbit extraction + config.json quantization_config probe (cwd-contained + symlink-rejected). decide_quantization() short-circuits the VRAM heuristic when prequantized is set. autopilot pipeline auto- applies so TheBloke/Llama-2-7B-Chat-GPTQ is recommended gptq instead of 4bit-on-top-of-quantized. - #142 merge_4bit + export_torchao live writers soup merge --save-format {fp16|4bit|4bit_forced}: single BNB-4bit merged checkpoint without the dequant->merge->requant cycle (fixes wrong-name llm_int8_skip_modules to bnb_4bit_skip_modules per code- review). soup export --format torchao --quant-config <yaml>: torchao .quantize_ + save_pretrained with per-scheme closed kwarg allowlist (Int4WeightOnly accepts {group_size, inner_k_tiles}, NVFP4 accepts nothing extra; dunder + unknown keys rejected per security-review H1). load_quant_config enforces yaml.safe_load + 256 KB cap + extension allowlist + cwd containment + S_ISLNK rejection. - #139 export_advanced_gguf via llama.cpp imatrix 3-stage pipeline: convert_hf_to_gguf.py -> optional imatrix -> quantize. argv-list subprocess (no shell), 30-min timeout, realpath- verified convert script stays inside llama_cpp_dir (security-review M5). _prepare_calibration_text accepts JSONL with text/prompt/content field aliases + raw text fallback; strips null bytes, collapses newlines, 8 KB per-line + 50 MB total cap (security-review M1); POSIX O_NOFOLLOW closes the TOCTOU window between dispatch-time check and open() (security-review M3). UD- prefix stripped before passing to llama-quantize. _safe_stderr Rich-escapes subprocess stderr before embedding in RuntimeError (security-review L4). - #109 soup deploy autopilot --measure Live Quant-Lobotomy scorecard: classifies each candidate quant OK / MINOR / MAJOR (thresholds 2% / 5% mirror v0.26.0 Part D). Results cached at ~/.soup/deploy_autopilot_cache.json (atomic write, 0o600 perms on POSIX, S_ISLNK rejection on BOTH load and save). pick_best soft-fallback now picks max-by-delta (was max-by-after) matching the v0.33.0 #54 design intent. _DEPLOY_MEASURE_BEFORE_GEN / _AFTER_FACTORY module-level hooks act as the stop-gap escape hatch until v0.46.1 ships first-party transformers / vLLM generator factories. - #70/#72 manual QA log scripted at tests/qa/v053_qa.md with exact reproduction recipes + acceptance criteria for the CUDA + llama.cpp smokes that can't run on the CI runners. Shared cleanup: - soup_cli/utils/paths.enforce_under_cwd_and_no_symlink consolidates the v0.33.0 #22 TOCTOU pattern previously copy-pasted in save_formats.py and gguf_quant.py (code-review HIGH fix). Reviews ran: python / code / security / tdd. Every CRITICAL / HIGH / MEDIUM / LOW finding fixed or documented. Test count: 7610 -> 7722 (+112 across 4 new files). Known limitations: live GPU + bitsandbytes / torchao smokes for the new merge / export paths remain pending (recipes in QA log); injected- generator escape hatch is non-public until v0.46.1; cache key truncates base_sha to 16 hex (1-in-2^32 collision floor). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
07e7214ed3 |
feat(v0.53.0): Quant Menu II — UD GGUFs + KV cache + NVFP4 + LF parity + save formats
Schema-only release. Live wiring deferred to v0.53.1 (mirrors v0.50.0 /
v0.51.0 / v0.52.0 stub-then-live pattern).
- Part A — Unsloth Dynamic 2.0 GGUF ladder (14 entries: UD-Q8_K_XL ... UD-IQ1_M)
+ validate_calibration_data_path shape validator.
- Part B — IQ (12) + Apple/ARM (10) GGUF flavours in utils/gguf_quant.py;
O(1) _LOWER_INDEX MappingProxyType for case-insensitive lookup.
- Part C — training.kv_cache_type: q8_0 | bf16 | f16 | fp8 (fp8 Hopper-only;
MLX rejected). requires_hopper reads from spec metadata (single source).
- Part D — fp8_attention (requires quantization_aware='fp8') + nvfp4 (Blackwell)
+ native unsloth_bnb_4bit bool flags with cross-validators.
- Part E — bnb_4bit_use_double_quant + llm_int8 (explicit 8bit assertion,
distinct from v0.41.0 load_in_8bit aliasing) + quantize_ref_model
(extends ref-task set with grpo/kto/ppo) + quantize_reward_model.
- Part F — soup merge --save-format {fp16, 4bit, 4bit_forced} + soup export
--format torchao with closed PTQ scheme allowlist (Int4WeightOnly,
Int8DynActInt4, Float8DynActFloat8, NVFP4 — case-sensitive PyTorch names).
Test count: 7453 → 7610 (+157 net new across 154 tests in test_v0530.py).
ruff check soup_cli/ tests/ — clean.
5 review agents ran (python / code / security / tdd / verification);
every CRITICAL / HIGH / MEDIUM / LOW finding fixed or documented:
- O(N) gguf walk → O(1) _LOWER_INDEX MappingProxyType
- ref_tasks extended with grpo + kto + ppo (silent-no-op footgun)
- _validate_v053_bool_fields no longer coerces None → False
- requires_hopper delegates to _KV_CACHE_METADATA spec
- fp8_attention validator order: quantization_aware before MLX
- validate_calibration_data_path + validate_quant_config_path docstrings
name the exact controls v0.53.1 CLI dispatch MUST add (TOCTOU contract)
- validate_torchao_scheme case-sensitivity documented at validator
- tautological `result == result` test replaced with allowlist invariant
- bool guards added on backend/modality/quantization across all Part D
validators
- exact 4096/4097 boundary tests for path shape validators
Docs updated: CLAUDE.md (test counts + utils list + changelog + test-table),
README.md (What's New replaced + 5 new dedicated sections), SECURITY.md
(support window + v0.53.0 hardening entry), CONTRIBUTING.md (test count
+ utils list + test-table row), .claude/plan.md (heading + boxes + banner —
gitignored, local only).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
df7f49feda |
feat(v0.52.0): Modality II — TTS + Distillation + BitNet + EBFT-GDPO + MoE quant + reasoning_effort
7 schema-only Parts; live trainer / loss / export wiring deferred to v0.52.1
(mirrors v0.27.0 MII / v0.50.0 GRPO Plus / v0.51.0 hubs stub-then-live pattern).
- Part A: task='tts' + modality='audio_out' + 5 families (Orpheus/Sesame-CSM/
Llasa/Spark/Oute) + per-family emotion allowlist (Orpheus 8 / Oute 6) +
5 recipes (orpheus-tts-sft, sesame-csm-tts, llasa-tts, spark-tts, oute-tts).
- Part B: classifier / reranker / cross_encoder tasks + num_labels (bool-
before-int guard) + classifier_kind + label_names (dedup + null-byte + cap).
- Part C: task='distill' + teacher_model + distill_divergence (kl alias
canonicalises to forward_kl; Literal excludes alias) + distill_temperature
(math.isfinite + [0.05, 100.0] bounds).
- Part D: quantization='bitnet_1.58' (gated to non-MLX + text + task in
{sft, pretrain, dpo}) + Falcon-E BitNet recipe + soup export
--format bitnet/tq1_0 CLI stubs (yellow deferred-advisory panel, exit 0).
- Part E: EBFT (structured/strided + bounded ebft_temperature; SFT-only)
+ GDPO (standard/length_normalized/margin; DPO-family-only).
- Part F: moe_expert_quant (nf4/int8_rowwise) + train_router_only — both
require moe_lora=true (silent-no-op footgun rejection).
- Part G: reasoning_effort (low/medium/high) + train_on_eot — both gated to
the SFT-family task set (sft/pretrain/distill/classifier/reranker/
cross_encoder); rejected on DPO/GRPO/PPO/etc. with named offenders.
Review fixes (5 agents: python, security, code, tdd, verification-loop-manual):
- num_labels bool-before-int field_validator (security HIGH)
- reasoning_effort + train_on_eot SoupConfig task-gate (code HIGH)
- _validate_classifier_compat lazy-import early-return (code HIGH)
- Oute emotion allowlist via data-driven _FAMILY_EMOTIONS (python+code MED)
- _MAX_LEN -> _MAX_REASONING_EFFORT_LEN (python MED)
- validate_reasoning_effort wired via field_validator (security MED)
- distill_divergence Literal excludes "kl" alias (code MED)
- DIVERGENCES derived from _DIVERGENCE_ALIASES drift guard (python LOW)
- is_bitnet_model comment fixed to match impl (python LOW)
- sister-fn bool guards on every compat helper (python+security MED)
- TDD coverage gaps closed: EBFT variant oversize, GDPO full rejection matrix,
ebft_temperature explicit-exc table, TTS compat input guards, recipe
model-id null/whitespace, full reasoning_effort task-matrix.
Drift fixes:
- tests/test_onnx_tensorrt_export.py + tests/test_awq_gptq_export.py
SUPPORTED_FORMATS count bumped 5 -> 7 (bitnet + tq1_0 stubs).
- tests/test_recipes.py catalog_size assertion 106 -> 112 (5 TTS + Falcon-E).
Test count: 7184 -> 7456 (+272). 230 new tests in tests/test_v0520.py;
remainder from drift-fix parametrize expansions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
1e1abacb44 |
feat(catalog): v0.51.0 — Model Catalog Expansion + Alternative Hubs
26 new ready-made recipes (catalog 80 → 106) covering 25 model families: GPT-OSS 20B/120B, GLM 4.6/5, Kimi K2 / K2-Thinking GRPO, MiniMax-M2, QwQ-32B GRPO, QVQ-72B, Granite 4, Liquid LFM2, Cogito v2, Mistral Small 3 / Medium 3.5, Magistral / Devstral / Ministral, MedGemma, EmbeddingGemma, LLaVA-Next, InternVL 3.5, Voxtral, Baichuan 2, Qwen-Image, DeepSeek-OCR, Paddle-OCR-VL. Part D: MULTIPACK_ARCHITECTURES extended 18 → 38 (Granite, GLM, Kimi, MiniMax, QwQ, QVQ, GPT-OSS, Magistral, Devstral, Ministral, MedGemma, LFM2, Cogito, Hunyuan, Ernie, Yi, Baichuan, ChatGLM). Part E: alternative model hubs. New soup_cli/utils/hubs.py with closed allowlist (hf/modelscope/modelers), SSRF-hardened endpoint validators mirroring v0.29.0 HF_ENDPOINT policy (scheme allowlist, loopback-only HTTP, RFC1918/link-local rejection, control-char/CRLF rejection, IPv6-mapped private rejection). TrainingConfig.hub Literal field with case-insensitive _normalize_hub field_validator; SoupConfig _validate_hub_supported rejects backend=mlx + hub != hf. Live downloader / uploader wiring deferred to v0.51.1 (matches the v0.27.0 MII / v0.37.0 multipack stub-then-live pattern). +455 tests (6729 → 7184), 0 regressions. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
33c60b4c1f |
feat(grpo): v0.50.0 — GRPO Plus (unsloth + axolotl RL parity)
22 features across 5 internal Parts shipped as schema-only — closed allowlists, Pydantic validators, NotImplementedError stubs for live wiring deferred to v0.50.1 (mirrors v0.27.0 MII / v0.37.0 multipack / v0.41.0 LLaMA Pro / v0.45.0 plugins / v0.48.0 curriculum / v0.49.0 LongLoRA stub-then-live pattern). Part A — 7 GRPO objective variants (gspo / dapo / dr_grpo / bnpo / two_sided / rft / standard) with `validate_grpo_variant` + frozen `GRPOVariantSpec` metadata + `MappingProxyType`-wrapped registry. `validate_grpo_delta` is bool-first / math.isfinite / (0, 1] bounded. `apply_variant_loss` raises NotImplementedError with v0.50.1 marker for the 6 new variants and is a no-op for standard. Part B — long_context_grpo + vllm_sleep_mode schema gates with compat validators (null-byte rejection on task + backend, bool guard on use_ring_attention). vllm_sleep_mode requires task='grpo' AND a transformers/unsloth backend (code-review HIGH fix — sleep is a between-rollouts feature). Part C — 4 multi-turn rollout backends (art / ruler / nemo_gym / openenv) with closed allowlist + per-entry required_package mapping. Part D — 7 stability/efficiency knobs (ref_model_ema_alpha, replay_buffer_size, async_grpo_prefetch, tis_threshold, mask_truncated_completions, defer_rerolling, skip_zero_advantage, off_policy_mask_threshold). Every numeric field rejects bool via a shared `_reject_bool_on_grpo_numerics` field_validator (tdd-guide HIGH fix — Pydantic v2 coerces True→1 by default). The `mask_truncated_completions` + `tis_threshold` pairing is enforced by a cross-validator (matches v0.32.0 spike-recovery+watchdog policy). Part E — top-level task='prm' Literal addition (Process Reward Model / stepwise-supervised, paired with data.format='prm' from v0.42.0) + `vision_grpo: bool` flag for VLM-RL on Qwen2-VL / Pixtral / InternVL. Compat helpers gate task / modality / backend. Review-round fixes applied (5 sequential reviews per CLAUDE.md): - python-review: list_variants annotation, frozenset[str] params, Optional[str] → str | None, module-level math import, D401 imperative docstrings, dropped *args/**kwargs on stubs. - code-review: grpo_fp16 added to GRPO-only task-gate; vllm_sleep_mode requires task='grpo'. - security-review: explicit field_validator for grpo_delta NaN/Inf rejection (Pydantic le=1.0 incidentally rejects NaN, made explicit); null-byte rejection on backend/task in grpo_long_context helpers; use_ring_attention bool guard. - tdd-guide: bool-rejecting field_validator on all Part D numeric fields + grpo_delta; missing bool-rejection tests added on validate_grpo_variant / validate_rollout_backend; null-byte test on validate_vllm_sleep_mode_compat; required_rollout_package rejection path; RolloutBackendSpec.live_wired; PPO+vision_grpo round-trip; _DEFERRED_LIVE invariant. Test count: 6490 → 6729 (+239 across 5 new test files). Notes for future maintainers: - v0.50.0 has zero new CLI commands and zero new trainer wirings; every step 6d/6e is intentionally n/a. Step 6 smoke runs schema happy + every documented cross-validator rejection. - All `task='grpo'` gates use `if self.task != 'grpo'` literal comparisons; do NOT switch to a set membership check until Part D knobs are wired into PPO/preference trainers in v0.50.x. - Multi-modal Vision RL does not yet verify the base model is actually a VLM — upstream trainer surfaces the error loudly when it fails to load the vision tower. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
161c81ee7f |
feat(long-context): v0.49.0 — YaRN, Dynamic NTK, LongLoRA S², Llama 3.1 NTK
- Part A: YaRN RoPE scaling — math kernels (yarn_find_correction_dim /
yarn_find_correction_range / yarn_linear_ramp_mask / yarn_get_mscale) +
4 yarn_* schema fields + cross-validator rejecting yarn fields outside
rope_scaling_type=yarn. Pure-Python implementations of the upstream YaRN
paper §3.4/§3.5 with bool/NaN/Inf rejection on every numeric input.
- Part B: Dynamic NTK — existing path verified, explicit test coverage.
- Part C: LongLoRA S² shifted-sparse attention (schema-only) — new
soup_cli/utils/longlora.py with is_llama_model (word-boundary regex
mirroring v0.39.0 is_gemma4_model policy) + validate_longlora_compat.
TrainingConfig.use_longlora + SoupConfig._validate_longlora_compat.
Live LlamaAttention.forward override deferred to v0.49.1 (stub-then-live
pattern, mirrors v0.27.0 MII / v0.37.0 multipack).
- Part D: Llama 3.1 NTK-aware (full impl) — scale_inv_freq_llama3
smooth-transition kernel + detect_llama3_rope_in_config HF-config probe
+ "llama3" added to rope_scaling_type Literal. LLAMA3_DEFAULT_*
constants per Unsloth models/llama.py:1853.
Public-boundary input validation on get_rope_scaling_config (bool/NaN/Inf
rejection on target_length / original_length / yarn_factor) per security
review — prevents direct callers from emitting {factor: NaN} into HF
model configs when bypassing the Pydantic schema.
Reviews: code-reviewer (2 HIGH + 1 MEDIUM + 1 LOW), security-reviewer
(1 MEDIUM + 2 LOW), python-reviewer (4 findings), tdd-guide (8 coverage
gaps) — all findings fixed. verification-loop done as manual equivalent
(version + --help + happy/failure YAML smoke).
+80 net new tests (6410 → 6490). Full suite green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
f789953d46 |
feat(data): Data Mixing Optimizer + v0.48.0 release (v0.48.0 Part B, BETA)
BETA. New `soup data mix --optimize --budget 1h --datasets a,b,c` runs N short proxy-training runs over candidate mixture weights and writes a canonical recipe YAML you can splice into `soup.yaml` under `data.interleave`. Per-candidate proxy failures are isolated (DEBUG-log + sentinel `_MAX_LOSS` + `continue`); `KeyboardInterrupt` / `SystemExit` re-raised; budget cap surfaces `MixOptimizationReport.partial=True`. `soup data mix --apply <recipe.yaml>` re-loads + prints the recipe's canonical interleave block. Both modes enforce `is_under_cwd` containment + TOCTOU symlink rejection (`os.lstat + S_ISLNK`) + 256 KB file cap; YAML key injection defended at the renderer (rejects newlines / null bytes / oversize dataset paths). Synthetic offline proxy ships in v0.48.0; live `soup train` proxy + scikit-optimize backend wiring deferred to v0.48.1 via `OptimizerProtocol` ducktype (default fallback: deterministic Dirichlet sampler). Review fixes: - `validate_datasets` early `len(raw) < 2` check (code-review MEDIUM) — prevents the less-actionable error after realpath resolution. - `run_mix_optimizer` proxy exceptions now isolated per-candidate (code-review MEDIUM) — first-cut raised RuntimeError on the first proxy failure, breaking the documented `partial=True` contract. - `load_mix_recipe` `os.lstat` wrapped in `try/except OSError` (security HIGH) — closes a TOCTOU race where path disappearance between `lexists` and `lstat` would raise an unhandled OSError. Release bundle (v0.48.0): - version bump → 0.48.0 in pyproject.toml + soup_cli/__init__.py - README.md: replaced "What's New" + 2 new dedicated `##` sections - SECURITY.md: supported-versions window + per-version notes for v0.48.0 - CONTRIBUTING.md: test counts (6242 → 6410) + 2 new test-table rows +94 tests. Net release total: 6242 → 6410 (+168 tests, +2 test files). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
fe9fe06b68 |
feat(training): Curriculum-Aware Trainer — dynamic bucket re-weighting (v0.48.0 Part A, BETA)
BETA. Adds `training.curriculum_dynamic: true` schema flag with online uncertainty estimation: every N steps, aggregate per-sample loss + grad-norm into per-bucket softmax weights, water-filled to enforce a minimum `curriculum_dynamic_floor`. DDP/grad-accum safety via `validate_distributed_curriculum` cross-validator that rejects un-coordinated multi-rank runs upfront — the well-known footgun where divergent per-rank stats silently desynchronise the sampler. New `soup runs curriculum-curve <run_id>` visualiser with TOCTOU (`os.lstat + S_ISLNK`) + 50 MB file-size cap + 100k-line streaming cap on the history file. Schema gated to sft/pretrain on transformers backend; mlx + non-SFT rejected with distinct messages. Live HF Trainer callback wiring deferred to v0.48.1 (stub-then-live pattern; mirrors v0.27.0 MII / v0.37.0 multipack / v0.41.0 LLaMA Pro). Review fixes: - water-fill design fix (code-review HIGH): removed trailing renorm that could push elements below `floor` when accumulated float error left sum slightly > 1.0. Softmax already sums to 1.0, so water-fill output also sums to 1.0 (drift bounded by nb*eps). - DoS caps on `render_curve` + `parse_history_jsonl` (`_MAX_HISTORY_ROWS=100_000`) — without these an attacker-controlled JSONL with 10M rows would OOM the process. - `curriculum-curve` CLI: symlink rejection + 50 MB + 100k-line caps, null-byte rejection on tracker-supplied `output_dir`. +74 tests. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
f64cee569e |
feat(v0.47.0): Data Forge — synthetic data pipeline + data quality moat
Part A — Synthetic Data Forge (utils/data_forge.py + commands/data_forge.py):
- soup data forge --docs <dir> --task sft|preference|tool with full provenance manifest
- ForgePlan / ProvenanceRecord / ForgeRow frozen dataclasses
- chunk_document + score_uncertainty pure-function kernel
- discover_documents (cwd-contained, symlink-rejecting, .txt/.md/.json/.jsonl allowlist)
- synthesise_forge_rows with judge-exception swallow at DEBUG
- Atomic JSONL + provenance writes via staged-tempfile + os.replace
Part B — Data Quality Moat (utils/data_score.py + commands/data_score.py):
- soup data score / decontaminate / toxicity / langdetect / pii / educational
- ReDoS-hardened PII regexes (phone + credit-card rewritten, 50 KB pre-cap)
- Containment-based n-gram decontamination (docstring corrected from "Jaccard")
- 6-language stopword heuristic for langdetect
- compute_scorecard with per-row DEBUG logging on swallowed errors
Security review fixes applied:
- math.isfinite guard on _validate_float_unit (NaN/Inf rejected before bounds)
- ReDoS: phone regex flattened (no nested optional quantifiers); credit_card
rewritten from {13,19}-loop to anchored 4-4-4-N; 50 KB pre-cap before finditer
- is_under_cwd moved inside discover_documents (no longer relies on caller)
- os.lstat + S_ISLNK rejection on every write target + tempfile staging
- _require_str rejects null bytes (consistency with data_forge._validate_str)
- compute_scorecard try/except blocks log at DEBUG (no silent swallow)
- decontaminate_texts: Optional[...] = None (no more type: ignore)
- _read_rows / _write_rows have full type annotations
- ngram_overlap_ratio docstring renamed to "containment ratio"
- import math + import tempfile moved to module top (lazy-import policy
applies to heavy ML deps only, not stdlib)
- Duplicate discover_documents call in CLI collapsed (TOCTOU window closed)
Live judge providers, [data-pro] extras (Llama-Guard / FineWeb-Edu / Presidio /
fastText / langdetect), and operator-supplied benchmark corpora are
stub-then-live and ship in v0.47.1 (mirrors v0.27.0 MII / v0.37.0 multipack
precedent).
Tests: 6126 -> 6242 (+116 net new). All v0.47.0 + full suite green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
986a4c00e3 |
feat(agent): Agent Forge — OpenAPI/MCP/GraphQL spec → tool-calling SFT dataset (v0.46.0 Part B)
Bumps to v0.46.0 and ships the Agent Forge: parse OpenAPI 3.x, MCP server
manifests, or GraphQL introspection JSON straight into a tool-calling SFT
dataset where each row is `{messages: [user, assistant_with_tool_call],
tool, source_endpoint}`. No more hand-rolled jsonl scaffolding for
function-calling fine-tunes.
* soup_cli/utils/agent_forge.py — `Endpoint` / `SynthRow` / `SpecReport`
frozen dataclasses. `parse_openapi` / `parse_mcp` / `parse_graphql`
parsers leave `$ref` strings opaque (no external resolution — defends
against file-read SSRF). All synthesised path strings routed through
`_validate_path` (rejects newline-in-name across all three parsers).
`load_spec_file`: `is_under_cwd` + `os.lstat + S_ISLNK` BEFORE realpath
(corrects v0.46.0 first-cut ordering caught by security review) + 5MiB
cap + yaml.safe_load only. `write_dataset` atomic via mkstemp +
os.replace (mid-stream TypeError never leaves partial file; mirrors
v0.43.0 Part D `copy_bundle_to` policy) + symlink rejection at target.
Caps: `_MAX_ENDPOINTS=10_000`, `_MAX_SPEC_BYTES=5MiB`,
`_MAX_ROWS_PER_ENDPOINT=32`.
* soup_cli/commands/agent.py — `soup agent synth/train/eval` Typer
subcommands. `synth` table cells pass through `rich.markup.escape`.
`train` rejects NUL/newline/oversize in `--base` and `--output-dir`
BEFORE embedding into rendered YAML recipe (CRITICAL security fix —
defends against YAML key injection where `--base $'evil\ntraining:
{ epochs: 9999 }'` would smuggle in injected training keys). `eval`
enforces predictions `is_under_cwd` + symlink rejection +
`_MAX_PRED_LINES=1_000_000` DoS cap.
* soup_cli/cli.py — registers `agent` Typer group; help string uses
ASCII-safe `->` (`test_help_output_is_ascii_safe` regression test caught
a Unicode `→` on first try).
* tests/test_v0460_part_b.py — 71 tests covering every parser kind,
failure modes (cycle / cap / null-byte / oversize / outside-cwd /
symlink), atomic-write partial-failure invariant, every CLI surface.
* Docs: README ## What's New replaced + dedicated `## Deploy Autopilot`
and `## Agent Forge` sections added; SECURITY.md supported-window
shifted (v0.46→full, v0.41→drop) + v0.46.0 fix-notes entry;
CONTRIBUTING.md test counts 165→167 / 5989→6126.
Test suite: 5989 → 6126 (+137 net new) green on Windows.
Known limitations (live runtime deferred to v0.46.1):
- Quant-Lobotomy auto-measure for deploy autopilot
- RLVR `code_exec` sandbox scoring in `agent eval`
- In-process `soup train` re-entry in `agent train` (Typer commands aren't
safe to re-enter — matches v0.44.0 `soup quantize` design)
- ExecuTorch packaging for iphone-16 / pixel-9 (lands in v0.54.0)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
479931282c |
feat(deploy): On-Device Deploy Autopilot — 10-profile catalog + recipe/script writer (v0.46.0 Part A)
Closes the v0.45.0 Known Limitations gap that flagged the 15-entry external integrations catalog as descriptive-only. This release ships the live autopilot CLI that picks PEFT + quantisation + speculative-decoding for the target hardware and emits a ready-to-train recipe. * soup_cli/utils/deploy_autopilot.py — `DeployProfile` frozen dataclass + 10-entry MappingProxyType-wrapped catalog: mac-m3 / mac-m4-pro / rtx-3060-12gb / rtx-4090-24gb / iphone-16 / pixel-9 / ollama-local / lm-studio / runpod-a100 / hf-jobs-h100. `_make` factory rejects bool recommended_max_length, non-kebab-case names, runtime/quant/peft outside closed allowlists. `render_recipe_yaml` validates `base` (≤200 chars, no NUL/newline) + `output_dir` (≤4096 chars). `render_deploy_script` uses `shlex.quote` on model_path. `write_recipe` / `write_deploy_script` enforce `is_under_cwd` + `os.lstat + S_ISLNK` TOCTOU rejection at the write target (matches v0.33.0 #22 / v0.43.0 Part C / v0.44.0 Part B / v0.45.0 Part E policy). * soup_cli/commands/deploy.py — new `autopilot` Typer subcommand with `--target`, `--base`, `--recipe-out`, `--script-out`, `--list`. Every Rich-rendered profile field passes through `rich.markup.escape`. * tests/test_v0460_part_a.py — 76 tests covering catalog immutability, 10-profile presence, case-insensitive resolution, every `_make` failure mode, render-* validation matrix, write-* path-containment + symlink rejection (POSIX-only), CLI smoke (list / writes / outside-cwd reject). Live Quant-Lobotomy auto-measure (against the v0.26.0 Checker) is deferred to v0.46.1 — this release writes the canonical combo per profile, the runtime measurement step lands in the patch release. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
74301843a2 |
feat(v0.45.0): Plugin System & Ecosystem Wins — 5 Parts, +169 tests
Adds a public plugin/hook system plus the schema scaffolding for 20+ ecosystem
integrations. Live trainer-callbacks, Anthropic /v1/messages route, server-tool
HTTP endpoints, and the recipe runner ship in v0.45.1 (matches v0.27.0 MII /
v0.37.0 multipack / v0.41.0 LLaMA Pro stub-then-live pattern).
Part A — Plugin / hook system
* New soup_cli/plugins/ package: BasePlugin Protocol, PluginSpec frozen
dataclass, register_plugin / discover_hooks / enable_plugin /
disable_plugin / load_plugins. Kebab-case name regex, semver-ish version,
null-byte rejection on every string. Idempotency check covers
(version, plugin object, templates, model_groups, description) — review-fix
added description after first-cut omitted it. Per-list caps on templates
and model_groups (32 entries, 128-char per name).
* New soup plugins list/install/enable/disable Typer CLI; all user-controlled
output passes through rich.markup.escape.
Part B — API extensions (schema-only)
* utils/anthropic_messages.py: to_anthropic / from_anthropic /
validate_anthropic_payload converters. Multiple system messages join with
\n\n; tool role with structured (list) content concatenated into single
tool_result text block (review-fix MEDIUM — first-cut silently dropped).
max_tokens cap 16384, temperature [0.0, 2.0], bool rejection on numerics.
* utils/server_tools.py: closed {python, bash, web_search} allowlist,
WebSearchConfig with domain allowlist + leading-dot subdomain pattern,
rate_limit [1, 600]. is_domain_allowed strips :port suffix and rejects
IPv6 literals (review-fix MEDIUM).
* utils/ngram_spec.py: NgramSpecConfig validators with bounded n / draft
tokens / prompt-lookup-max; bool rejection on every numeric field.
Part C — External integrations catalog
* utils/integrations.py: 15-entry MappingProxyType catalog of ecosystem
targets (lm-studio, comfyui, ollama, claude-code, cursor, continue, ...).
Part D — Advanced trainer-plugin allowlist
* utils/trainer_plugins.py: 6-entry allowlist (grokfast, spectrum,
llmcompressor, sonicmoe, cce_plugin, math_verify) + validate_trainer_
plugin_list (Sequence[str], dedup, _MAX_PLUGINS_PER_RUN=8).
Part E — Data Recipe DAG
* utils/recipe_dag.py: closed NODE_KINDS frozenset, Kahn's topological
sort via collections.deque (review-fix HIGH — first-cut had O(N^2 log N)
queue.sort() inside the BFS body), cycle / self-loop / dangling-edge
rejection, _MAX_NODES=256 / _MAX_EDGES=1024 / _MAX_FILE_BYTES=1MiB.
load_recipe_yaml enforces is_under_cwd containment AND os.lstat + S_ISLNK
symlink rejection (review-fix MEDIUM — TOCTOU defence; mirrors v0.33.0 #22
/ v0.43.0 Part C / v0.44.0 Part B policy).
* New soup data recipe <path> CLI validates topology and prints planned
topo order; live runner deferred to v0.45.1.
Reviews: python-review, security-review, code-review, tdd-guide all run;
verification-loop replaced by manual smoke (CLI happy + failure paths
exercised on real fixtures).
Test count: 5820 -> 5989 (+169). Test files: 164 -> 165. Ruff clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
c4ac3da695 |
feat(v0.44.0): Live Dashboard & UX - 21 features, +192 tests
Part A - Live monitoring: soup monitor (nvidia-smi panel), EMA + p95/p99 tail-latency stats, SSE training-stream schema, phone-visible URL + ASCII-QR helper, llama-server timings parser + KV-cache bar, thread-safe ToolOutputsBuffer + ToolCallTimer. Part B - UX fixes: GracefulSaveHandler (first SIGINT saves, second stops), .checkpoint_now trigger file (cwd-contained, symlink-rejected), desktop / .command / .cmd shortcut builders, onboarding-wizard YAML renderer. Part C - UI tabs: drop-in soup_cli/ui/plugins/*.py registry with kebab-case name allowlist + 32-tab cap, API_HOST / API_PORT / API_KEY + GRADIO_HOST / GRADIO_PORT env knobs. Part D - Standalone CLIs: soup fetch (bundled examples + configs + deepspeed_configs catalog), soup quantize (ergonomic alias), soup merge-sharded-fsdp-weights, soup delinearize-llama4 (planners; live runtime in v0.44.1), soup llama <sub> (closed-allowlist proxy with filtered child env that drops HF_TOKEN / OPENAI_API_KEY / ANTHROPIC_API_KEY), soup_cli.utils.sweep_config (separate sweep.yaml loader), reasoning_parser allowlist for soup serve. Security review fixes: fetch symlink-at-target rejection + bundled-source commonpath check, write_trigger symlink rejection (TOCTOU), llama child-env secret allowlist, onboarding output cwd-containment, qr token moved from URL fragment to query string (so server actually sees it), sweep-config scalar allowlist + MappingProxyType[Tuple] immutability. Code/Python review fixes: detect_apple_silicon clean rewrite (was buggy parser-priority ternary), all frozen-dataclass List fields -> Tuple, ToolOutputsBuffer -> collections.deque(maxlen=1000), os.path.realpath over abspath, IPv6 host auto-bracketing per RFC 3986, frozenset over mutable set, type hints on __exit__/_make_proxy. Test count: 5628 -> 5820 (+192). Lint clean. Help output ASCII-safe (em-dash check enforced by test_cli_subprocess). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
82d5693b75 |
feat(eval): Tracker & Eval Pro — 18 features (v0.43.0)
Closes the observability gap with all three competitors in one release.
Part A — Trackers
* --tracker flag (mlflow/swanlab/trackio) on soup train, mutually
exclusive with --wandb/--tensorboard. Closed allowlist via
MappingProxyType. Live integrations rely on HF Trainer's report_to.
* SOUP_TELEMETRY=1 opt-IN env var; build_telemetry_payload schema is
closed-key (no model names / dataset paths / config contents). Live
PostHog network code deferred to v0.43.1.
Part B — Eval metrics
* Pure-Python BLEU + ROUGE-1/2/L + effective_tokens_per_second.
* KL-divergence calibration framework with OK/MINOR/MAJOR thresholds.
* Model Arena Elo tournament (256-model cap, MappingProxyType view,
Rich-markup metacharacter rejection on names).
* ceval / cmmlu / aider_polyglot benchmark allowlist (live Aider
runner deferred to v0.43.1).
Part C — Profiling
* memory_snapshot_context (narrow RuntimeError catch — review fix
prevents generator-already-executing on user-body RuntimeError).
* detect_anomaly_context, nccl_bandwidth_check (h100/a100/v100/rtx
reference table; live measurement CLI surface deferred).
* write_vscode_launch with TOCTOU symlink rejection at the target
path regardless of force=True.
Part D — Demo bundles
* `soup data demo` lists / copies 4 bundled JSONL fixtures
(alpaca / sharegpt / dpo / grpo) with staged-tempfile atomic
rename + 50 MB cap + symlink rejection on the staging path.
Tests: 5389 -> 5628 (+239). Ruff clean. Five sequential review waves
(python / code / security / tdd / smoke) ran; HIGH/MEDIUM/LOW findings
all fixed including: tracker name shadow in train.py, _lcs_length DP
double-buffer bug, BLEU geo-mean policy, base_dir absolute/.. escape,
demo_bundles tmp symlink TOCTOU, vscode launch symlink TOCTOU.
Note (Windows CI): line-ending warnings (LF -> CRLF) on commit are
expected; `.gitattributes` policy is unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
05093ebfdb |
feat(data): Data Pipeline Pro — 18 features, axolotl + LF parity (v0.42.0)
Schema-first surface for the data pipeline gap with Axolotl + LlamaFactory. Ships in one release: 5 new formats (prm, pre_tokenized, input_output, video, multimodal), remote URI allowlist (s3/gs/gcs/az/abfs/abfss/oci) + streaming + sharding, AOT preprocess cache + `soup data preprocess` CLI, multi-dataset interleave (concat/under/over/probs) + 8 advanced masking fields, vocab expansion (add_new_tokens / new_special_tokens / resize_vocab) + custom prompt_strategy, and document ingestion (`soup data ingest` for PDF/DOCX/MD/TXT). Live wiring for fsspec backends, AOT tokenize loop, custom prompt-strategy runtime, and PRM trainer integration is deferred to v0.42.1+ (stub-then-live pattern from v0.27.0 / v0.37.0 / v0.41.0). Schema gates fire at config load so misconfiguration fails fast. Security: full v0.42.0 hardening matrix — `_REMOTE_SCHEMES` MappingProxyType allowlist; bucket regex 1-63 chars per S3/GCS spec; userinfo / fragment / query-string rejection on remote URIs (query-string forwarded to fsspec is SSRF-adjacent); null-byte + length caps on every string-shaped input; bool-rejected-before-int on every numeric input; frozen InterleaveSpec dataclass; 10k caps on add_new_tokens; `is_under_cwd` containment on video_dir + tokenized_path schema fields and on preprocess --config / both ingest paths; `os.lstat + S_ISLNK` symlink rejection on ingest input; PRM converter type-checks completions (str) + labels (bool, not int); video field null-byte + 2KB cap; field-name threading on image-pixels validator so error messages name the actual field. 5242 tests → 5389 (+147 net). 11 review findings addressed across python-review / code-review / security-review / tdd-guide (CRITICAL→LOW). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
1491025a36 |
feat(trainer): Optimizer & PEFT Zoo (v0.41.0)
Optimizer Zoo (Part A): closed-allowlist SUPPORTED_OPTIMIZERS adds 14 new
entries — BAdam, APOLLO (apollo_adamw), Adam-mini, lomo / adalomo,
grokadamw, schedule_free_adamw / schedule_free_sgd, muon, dion,
came_pytorch, ao_adamw_{fp8,4bit,8bit}. Validates name + lower-cases
deterministically; rejects non-string / empty / null-byte / >64-char.
_OPTIMIZER_PACKAGES wrapped in MappingProxyType (matches v0.36.0 _REGISTRY).
Per-module LR (Part B): training.lr_groups accepts list-of-pairs /
list-of-dicts / {pattern: lr} mapping; canonical [{pattern, lr}, ...]
storage. Capped at MAX_LR_GROUPS=32; per-pattern non-empty string
≤256 chars + null-byte rejection + re.compile + best-effort ReDoS probe;
per-LR (0.0, 1.0] + math.isfinite (rejects NaN AND ±inf) + bool rejection
(matches v0.30.0 Candidate policy); duplicates rejected. lr_groups_from_schema
bridges canonical schema shape into runtime List[LrGroup] for
build_optimizer_param_groups (first-match-wins routing). LrGroup is
@dataclass(frozen=True). PyYAML scientific-notation (1e-4) parses as
string in YAML 1.1; _validate_lr coerces str → float so soup.yaml
round-trips work. base_lr rejects bool / non-positive (defence-in-depth).
PEFT methods (Part C): LoraConfig.init_strategy="loftq" + loftq_iter
∈ [1, 10] + loftq_bits ∈ {2, 4, 8}; cross-validator rejects loftq +
use_dora / use_vera. utils/loftq_init.py exposes validators +
build_loftq_config (lazy peft.LoftQConfig with actionable ImportError
hint). LLaMA Pro: TrainingConfig.expand_layers ∈ [1, 64] +
freeze_trainable_layers (signed, |x| ≤ 1000); cross-validator requires
the pair (LLaMA Pro freezes original layers and trains only new blocks).
field_validator(mode="before") on both rejects bool BEFORE Pydantic ge/le
silently coerces True → 1. expand_model_blocks raises NotImplementedError
with v0.41.1 marker — schema-only release (mirrors v0.27.0 / v0.37.0
stub-then-live pattern). utils/block_expansion._count_layers uses
hasattr(__len__) instead of try/except TypeError so legitimate __len__
bugs surface loudly. use_mod boolean for Mixture-of-Depths (schema only —
live patch deferred to v0.41.1).
Friendly aliases: load_in_8bit / load_in_16bit (Optional[bool]) for
LlamaFactory / Axolotl users. is True policy on both — explicit False is
"no preference", not "off"; mutually-exclusive both-True rejected; alias
combined with explicit Quant Menu format raises rather than silently
overriding. Alias-driven quantization rewrite via direct
self.quantization = ... (Pydantic v2 BaseModel non-frozen path), NOT
object.__setattr__ — code review caught that the latter would silently
bypass any future field_validator on quantization.
Five review agents (4 in parallel + manual smoke for verification-loop):
all CRITICAL/HIGH/MEDIUM/LOW findings fixed. Local smoke caught a real
bug — PyYAML parsing 1e-4 as string — fixed pre-commit with 2 added tests.
5242 tests pass (+120 net new); ruff clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
5e0872b9ea |
feat(trainer): ReLoRA + surgical PEFT non-SFT (v0.40.6 #67)
Extends the v0.39.0 ReLoRA callback (Part B) and surgical PEFT patches (Part D — Gemma4 ClippableLinear swap + 3-D fused-MoE expert dropout strip) from SFT-only to all 11 non-SFT transformer-backend trainers (DPO, GRPO, KTO, ORPO, SimPO, IPO, PPO, RewardModel, Pretrain, Embedding, BCO). - New shared helper soup_cli/utils/peft_wiring.py exposes apply_pre_lora_patches, apply_post_lora_patches, attach_relora_callback. - SFT migrated to the same helpers in the same release (centralisation invariant; no drift between SFT and non-SFT wiring). - SoupConfig._validate_relora_supported_tasks: task != "sft" rejection removed; MLX backend still rejected with distinct message. Review fixes: - attach_relora_callback uses `if relora_steps is None:` (project policy) so a schema-bypassing relora_steps=0 surfaces as a loud ReLoRAPolicy ValueError rather than a silent skip. - Direct attribute access on tcfg.relora_warmup_ratio / _reset_optimizer / _prune_ratio (Pydantic schema guarantees them); no getattr defaults. - 11 behavioural helper tests (Gemma4 happy path + exception swallow, post-LoRA strip happy + exception swallow, ReLoRA policy field forwarding, schema-bypass loud-fail). - Schema-gate matrix covers `task='preference'` dispatcher. Tests: 5061 -> 5122 (+61). Closes #67. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
697cc8dad7 |
feat(trainer): Quant Menu non-SFT multi-trainer wiring (v0.40.5 #66)
Extends the v0.38.0 train-time quantization menu (gptq / awq / hqq:Nbit /
aqlm / eetq / mxfp4 / fp8) from SFT-only to all 11 transformer-backend
trainers (DPO / GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel /
Pretrain / Embedding / BCO). Closes the v0.38.0 known gap.
Schema gate: SoupConfig._validate_quant_menu_supported_tasks removes the
`task != "sft"` rejection branch. MLX backend rejection retained with
distinct message; modality=text gate retained (vision/audio Quant Menu
deferred — modality-specific kwargs not yet threaded through the unified
loader).
Trainer wiring (11 sites): each non-SFT _setup_transformers replaces the
inline BitsAndBytesConfig(load_in_4bit=True, ...) block with a call to
build_quantization_config_for_loader(tcfg=tcfg, base=cfg.base, console=console)
— mirrors sft.py:420-440 exactly. kbit-prep tuple widened from
("4bit", "8bit") to ("4bit", "8bit", "mxfp4"). BitsAndBytesConfig import
removed from each non-SFT trainer.
Review fix — PPO reward model: _load_reward_model gains optional tcfg
kwarg; when supplied, the reward checkpoint loads with the same Quant
Menu config as the policy. Both PPO call sites forward tcfg=tcfg.
Defends against silent fp16 OOM on a GPTQ/AWQ/HQQ policy run.
Review fix — defence-in-depth: new TrainingConfig.reward_model field
validator rejects null bytes and caps length at 512 chars (matches
cfg.base policy). The Quant Menu loader's per-call null-byte check
in _check_local_marker remains as the runtime backstop.
Tests: +131 net new (4930 -> 5061). New tests/test_v0405_part_a.py
parametrizes 11 tasks x 7 quant formats; covers MLX rejection per task,
quantization_aware x Quant Menu cross-validator regression for non-SFT,
source-level invariants (no inline BNB literal, kbit tuple regex),
and a live mock-based dispatch test for _load_reward_model proving the
Quant Menu path is reachable when tcfg is supplied and skipped when
tcfg=None.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
|
|
|
c85a1017b5 |
fix(v0.40.4): strip ANSI + route correctly in TestCommandFlagsExist
CI failure root cause: Rich line-wraps `--trust-remote-code` with ANSI colour escapes between `-`, `-trust`, `-remote-code` on narrow CI terminals (mirrors v0.40.3 ANSI fix). Substring assertion missed because the ANSI escapes were embedded mid-flag. Also: original test used `[cmd, "--help"]` then fell back to `["data", cmd, "--help"]`. For diff/export/merge/infer the first invocation worked (top-level commands) but the substring miss triggered the fallback into `data` subcommand, which then errored "No such command 'X'" — masking the real ANSI issue. Switched to explicit per-command argv lists. Adds the `_strip_ansi` helper from tests/test_trust_remote_code.py and routes `data generate` directly via `["data", "generate", "--help"]`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
6fdf7e2570 |
feat(trainer): multipack live HF Trainer DataLoader override (v0.40.4 Part B)
Closes #65 (deferred from v0.40.3). make_multipack_trainer_class adds a get_train_dataloader override that builds a MultipackBatchSampler(real_batches=False) — yields a flat list[int] per packed sequence, which is the contract HF DataLoader.batch_sampler expects — and installs it via DataLoader(..., batch_sampler=sampler, collate_fn=self.data_collator, num_workers=args.dataloader_num_workers, pin_memory=args.dataloader_pin_memory). drop_last is forwarded from TrainingArguments.dataloader_drop_last. Falls back to super().get_train_dataloader() when state was never attached OR when train_dataset is unset — defence-in-depth so the subclass remains safe to instantiate even when multipack is later disabled. The state-presence guard switched from falsy (`not max_seq`) to explicit `is None` (plus `not lengths` for empty-list defence) — attach_multipack_state already rejects non-positive ints, so the falsy guard would only mask configurator bugs. _get_train_sampler override stays as a defensive no-op fallback that ALWAYS delegates to super (review-fix from v0.40.4 code-review: returning a multipack list[list[int]] from this method would cause a shape mismatch if any HF eval / prediction loop bypasses get_train_dataloader and calls _get_train_sampler directly). SFT and Pretrain trainer wrappers now invoke make_multipack_trainer_class(SFTTrainer) and attach_multipack_state(...) when multipack: true. The v0.40.3 yellow advisory + standard-sampler fallback is gone. Architecture allowlist (validate_multipack_architecture) still gates at build time. tests/test_v0403_part_b.py: TestSftAndPretrainWiringDeferred renamed to TestSftAndPretrainWiringLive; the deferred-state test (_get_train_sampler returns MultipackBatchSampler when state is set) is replaced by the live-state test (_get_train_sampler always delegates to super even with state attached). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
3ab36e2aad |
feat(security): trust_remote_code opt-in across non-SFT trainers + 5 commands (v0.40.4 Part A)
Closes the v0.36.0 #63 known gap. Every non-SFT trainer wrapper (DPO / GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel / Pretrain / Embedding / BCO + the unified Preference dispatcher) now accepts trust_remote_code: bool = False on __init__, resolves once via the v0.36.0 helper (model_requires_trust_remote_code + resolve_trust_remote_code), and stores the resolved value on self._trust_remote_code. Every from_pretrained call site reads from the resolved attribute — no remaining trust_remote_code=True literal in any trainer file (asserted by tests/test_v0404_part_a.py). Five standalone commands gain a --trust-remote-code Typer flag with the same default-deny + KNOWN_SAFE_PREFIXES allowlist behaviour as soup train: soup diff, soup export, soup merge, soup infer, soup data generate. commands/train.py removes the v0.36.0 sft_kwargs split — every trainer receives trust_remote_code from the same trainer_kwargs dict. PreferenceTrainerWrapper forwards the raw bool to the inner DPO / SimPO / ORPO / IPO / BCO wrapper kwargs at both _build_inner and _build_multi_objective sites; the resolver fires inside the inner wrapper at construction time. _load_reward_model (module-level helper in ppo.py) accepts a trust_remote_code: bool parameter and resolves internally — design intent is that the helper is independently safe to call outside PPOTrainerWrapper. _export_onnx / _export_tensorrt / _export_awq / _export_gptq and _merge_adapter helpers all gain a trust_remote_code: bool = False parameter threaded from the Typer flag. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
21453101e7 |
fix(v0.40.3): strip ANSI escapes in help-text assertions
CI failure on previous v0.40.3 hotfix: Typer renders Rich-styled help with ANSI escape codes BETWEEN flag fragments (`--trace\x1b[0m\x1b[1;36m-log`), so a whitespace-only strip still failed to find `--trace-log` in the flattened output. Strip both ANSI codes (`\x1b\[[0-9;]*m`) and whitespace in one pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
18fc5b8d02 |
fix(v0.40.3): width-independent help-text + fastapi-skip on CI
CI failures on this commit: - macOS Typer help text wraps `--judge` / `--trace-log` to two lines on narrow CI terminals; tests asserted the raw string. Strip whitespace before match (mirrors v0.40.2 width-independent fix). - `fastapi` is not in the base CI deps (only `[serve]` extra); two `_create_app` tests ImportError-ed. Skip those tests when fastapi is unavailable. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
1f050fc235 |
feat(v0.40.3): Stub-to-live wave 1 (#33, #64; #65 still deferred)
Three v0.X.0 deferred-stub features become live runtime — closes #33 (harvester judge filter + serve trace log) and #64 (live CUDA OOM probe). #65 (multipack live wiring in HF Trainer) remains deferred to v0.40.4 after the adversarial 5th-review pass surfaced a Sampler[int] vs list[list[int]] shape mismatch with HF Trainer's DataLoader; helpers (`make_multipack_trainer_class`, `attach_multipack_state`, `lengths_from_dataset`, `detect_arch_name`) ship as a stub used by unit tests, but the SFT/Pretrain wrappers print a yellow advisory and fall back to the standard sampler when `multipack: true`. Live CUDA OOM probe (#64): `make_cuda_probe_fn` builds a closure that runs ONE forward+backward+step on a synthetic batch per candidate. `model.zero_grad(set_to_none=True)` runs BEFORE forward; intermediate ids/attn/labels/outputs are del-ed before `loss.backward()` so peak VRAM reflects a realistic training step (matches v0.35.0 #45 policy). `pad_id` is bounded by `len(tokenizer)` (not `vocab_size`) so extended vocabs (Llama-3 + `<|pad|>`) don't fold pad to a random byte token. SFT-only this release. Trace-to-Preference judge filter (#33 (a)): `judge_filter_pairs` reuses v0.19.0 JudgeEvaluator backends (openai/server/ollama). Threshold rejects bool/NaN/out-of-[0,1]; `_MAX_BATCH=100_000` cap applied via lazy `itertools.islice`; per-pair backend exceptions caught and DEBUG- logged (matches v0.33.0 #47 policy); `judge_provider` validated against the allowlist at the CLI boundary BEFORE constructor with a Rich-escape error message; yellow projected-call-count warning before the loop (2× per pair). Inference Server trace log (#33 (b)): `TraceLogWriter` is thread-safe (single-process lock — multi-worker documented as known limitation); path containment via shared `is_under_cwd`; null-byte/empty/non-string path rejected; cap_mb bounds [1, 10000] with explicit bool rejection. Rotation (one backup retained) refuses symlink at the backup path via `os.lstat + stat.S_ISLNK` (matches v0.33.0 #22 TOCTOU policy). Secret redaction (`hf_*` ≥8, `sk-*` ≥16, `Bearer …` ≥8 with `.` excluded so end-of-sentence period survives) applied to prompt + response and recursively to caller-supplied `extra` dict values. Streaming SSE path also records (was a coverage gap caught in adversarial review). Behaviour change: v0.40.2 users with `auto_batch_size_strategy: probe` were silently getting the static fallback. v0.40.3 actually runs a CUDA probe on first run (~5–30s, cached per (model, max_length, quant, lora_r, gpu) tuple). Reviews: 5 agents (python, code, security, tdd, verification-loop). Verification-loop run twice — once shallow smoke (PASS), once adversarial bug-hunt which found C1/C2 (multipack live wiring crash — demoted to v0.40.4), H1 (streaming SSE missing trace log — fixed), H4 (vocab_size vs len(tokenizer) on extended vocabs — fixed), H3 (Bearer regex consumed trailing period — fixed), H2 (judge cost shock — warning added), M2 (empty lengths accepted — rejected), L1 (extra dict bypassed redaction — recursive walk added). Tests: 4756 → 4855 (+99 net new) across test_v0403_part_a/b/c.py. Lint clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
676078056d |
fix(v0.40.2): make help-text assertions width-independent
CI on Linux/macOS runners has narrower terminals than the local Windows shell. Rich wraps long option names like ``--template-dir`` across two lines (``-\n-template\x1b...-dir``) which makes a substring check on the raw output string fail. Updated `_plain` helper in both v0.40.2 test files to strip whitespace in addition to ANSI escapes — matches the option name regardless of terminal width. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
b0fc586706 |
feat(v0.40.2): Quick polish + v0.40.1 carry-overs (#36, #50, #51 + 7 papercuts)
Closes 3 originally-scheduled GitHub issues plus 7 v0.40.1 long-tail UX papercuts. No new schema fields, no new trainers — pure polish. Originally scheduled: - #36 format_gate_row helper for the eval-gate dashboard row (pure formatter in soup_cli/monitoring/display.py; passed=is True so missing field renders neutral; supports stop/warn action suffixes; multi-task " | " join). - #50 prepare_hf_resume now skips snapshot_download when local checkpoint-N is greater-or-equal to the remote highest-N. New _find_highest_local_checkpoint helper handles missing dirs / OSError / non-directories cleanly. - #51 soup deploy hf-space --template-dir <path> via new soup_cli/utils/hf_space.py:render_custom_template_dir. Containment via is_under_cwd; validate_repo_id BEFORE substitution; per-file 256 KB cap; symlinks + non-regular files rejected (TOCTOU defence per v0.33.0 #22). v0.40.1 carry-overs: - H2: data filter --min-coherence alias; data split --train no-op; data register/unregister positional <name> <path> + Optional --name/--path with conflict detection. - H3: soup quickstart --output DIR (containment-checked) routes data, config, run dir under the chosen directory. - N1/G2: apply_logging_level pushes parsed --log-level tier into the root logger so transformers / peft / trl actually respect QUIET / DEBUG. - N7: shared _resolve_model_source in commands/infer.py (used by bench.py too) — path-like-but-missing raises FileNotFoundError; non-path-like values fall through to HF download via from_pretrained. - G13: verified ONNX/AWQ/GPTQ/TensorRT install hints already correct. - M4: verified data dedup --threshold already exposed. - M5: soup runs --cwd-only + _filter_runs_by_cwd helper using os.path.realpath + commonpath (Windows 8.3 + cross-drive safe). Review-fix follow-ups landed in the same release: - soup_cli/commands/infer.py: from __future__ import annotations (Py3.9 PEP 604 fix); --output containment via is_under_cwd, late-evaluated to preserve pre-existing test contracts. - soup_cli/commands/data.py register_data + soup_cli/commands/bench.py prompts file: Path.resolve()+relative_to() → is_under_cwd (project rule for Windows 8.3 short-name safety). - soup_cli/commands/runs.py: typed _filter_runs_by_cwd, removed redundant inner import os. - soup_cli/commands/deploy.py: confirmation panel now shows --template-dir path when set, not the unused --template default. Tests: 4720 → 4756 (+36) across two new files (test_v0402_part_a.py, test_v0402_part_b.py). 5 review agents (python / code / security / tdd / verification) all clean after fixes. Known limitations: - Custom HF Space templates always create the Space with sdk=gradio regardless of the supplied app.py. Use --template streamlit-chat with the inline registry for Streamlit. Tracked for v0.40.3+. - _resolve_model_source returns ("hf", repo_id) without validate_repo_id; transformers.from_pretrained will raise loudly on malformed ids. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
56bea56c08 |
fix(v0.40.1): QA Hardening — UTF-8 bootstrap, schema strictness, multi-objective preference runtime, CLI UX
Closes the QA findings from the Windows + RTX 3050 4 GB pass (2026-05-07): - Part A: UTF-8 stdio bootstrap on Windows (closes C1/C4/H1/N5/N8/G5) - Part B: root-level `lora:` migrates into training.lora (no more silent init_strategy bypass); multi-objective preference loss runtime no longer raises NotImplementedError (primary-loss approximation; full per-batch weighted combination deferred to v0.40.2) - Part C: autopilot 7B → 1B fallback + safetensors cache probe; transformers <5.0.0 cap with INCOMPATIBLE flag in `soup doctor`; quickstart auto-switches to SmolLM2-135M on ≤6 GB VRAM; --find-lr load_local → load_raw_data import fix - Part D (subset): dynamic --template help (H4); init --force (M2); migrate JSONL friendly error (N2); eval custom -o independent of attach-to-registry + loop-shadow bug fix (G10); history suggests dataset registry (N6); doctor importlib.metadata fallback (M1) + GPU diagnostic distinguishes CPU build (N3) + dual-Python detector (N4) - Part E: recipe fuzzy-match suggestions (M3); sample filename embeds strategy (no overwrite); JSONL BOM auto-strip Net +64 tests (4656 → 4720). 4 review agents clean (python/code/security/tdd). Long-tail UX papercuts (H2/H3/N7/M4/M5 + #36/#50/#51) deferred to v0.40.2. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |