Alpamys
dca58c4107
docs(train): v0.71.26 release — closed-loop reward-hacking mitigation
...
Version 0.71.25 -> 0.71.26 (pyproject + __init__). CHANGELOG [0.71.26] entry
(feature + security). README What's New slot. docs/training.md mitigation
section + docs/commands.md flag. CONTRIBUTING + examples/README. Also folds in
the already-merged qwen2.5-coder-7b-sft recipe (#285 ) that rides this release.
2026-07-01 16:55:59 +05:00
Alpamys
fcf4b33394
feat(train): native Spectrum targeted training — soup spectrum scan + training.unfrozen_parameters (v0.71.23)
...
Closes #266 . `soup spectrum scan` streams safetensors per-tensor (no model
load, CPU-friendly) and computes a singular-value SNR per weight matrix
(Marchenko-Pastur, arXiv:2406.06623), emitting a ready-to-paste
training.unfrozen_parameters patch. The SFT trainer freezes all params then
unfreezes the matched set (full FT, LoRA off).
- utils/spectrum_scan.py: pure-numpy transpose-invariant SNR kernel +
per-tensor safetensors streaming (2^31 SVD cap, symlink skip) + cache
(~/.soup/spectrum, SOUP_SPECTRUM_CACHE_DIR containment) + hardened
hubs.snapshot_download.
- commands/spectrum.py: soup spectrum scan (SNR table + YAML patch).
- schema: training.unfrozen_parameters (caps/NUL/invalid-regex/ReDoS reject)
+ gates (sft/transformers/text/quantization=none; mutually exclusive with
LoRA features / freeze_layers / freeze_ratio / train_router_only /
expand_layers).
- trainer/sft.py: full-FT branch via apply_unfrozen_parameters +
enable_input_require_grads (fixes grad-checkpointing through frozen
embeddings).
Existing spectrum trainer-plugin untouched (back-compat); LISA -> #267 .
Live-validated on Windows + RTX 3050: CPU scan of SmolLM2-135M + top-25%
unfrozen full-FT train (loss 3.455 -> 0.719).
Note: after the version bump the editable install metadata was stale
(0.71.17); pip install -e . --force-reinstall --no-deps re-synced it so
test_cli_subprocess::test_version passes.
+94 tests in tests/test_v07123.py (14184 -> 14278).
2026-06-12 17:40:46 +05:00
Alpamys
8b7d63f944
docs: clarify Orpheus live-codec TTS is live in training.md (v0.71.22)
...
The live-codec block claimed the entire data.format=audio path was "not
validated on the maintainer's box" and only surfaced a RuntimeError. v0.71.22
made the Orpheus SNAC encode live + validated; note that while the other four
families stay dependency-gated. Docs-only, no version bump.
2026-06-10 22:02:50 +05:00
Alpamys
ed5fc3a8b3
feat(precision,rollout): live fp8/nvfp4 + vLLM sleep + openenv rollout + apple-adapter + delinearize-llama4 (v0.71.21)
...
Closes #141 , #124 , #125 , #228 , #97 .
- #141 : apply_fp8_attention (torchao float8 on attention projections, Hopper
gate) + apply_nvfp4 (NVFP4Config, Blackwell gate); partial-conversion honesty;
wired into the v0.28 speed/memory pipeline with yellow-advisory degrade.
- #124 : vllm_sleep_mode live - create_vllm_engine(sleep_mode=True) +
vllm_sleep_cycle ctx (wake in finally) + TRL GRPOConfig hook probe.
- #125 : openenv rollout fully live via training.rollout_func module:fn
resolver; rows replace the prompt dataset; art/ruler/nemo_gym honest dep
gates + _EXTERNAL_ROLLOUT_RUNNERS seam. Real GRPO train on SmolLM2-135M.
- #228 : convert_apple_adapter live - PEFT LoRA <-> mlx-lm (both matrices
transpose, bf16 upcast, adapters.safetensors + num_layers, npz legacy read,
np.ascontiguousarray fix for safetensors non-contiguous mangling);
*-to-apple upstream-gated exit 3.
- #97 : delinearize-llama4 live - [E*din,dout] -> [E,din,dout] per shard,
config.json expert-count probe + --num-experts, sidecar copy, atomic writes.
Review waves: 3 HIGH + ~8 MEDIUM + ~12 LOW fixed.
Tests: 13874 -> 14084 (+210 in tests/test_v07121.py).
Full suite: 13967 passed, 117 skipped. ruff clean.
2026-06-10 16:29:52 +05:00
Alpamys
a4dfbb308c
feat(trainer): live TTS / BitNet / MoE-expert-quant trainers (v0.71.20)
...
Lift three v0.52.0 schema-only NotImplementedError stubs to real code.
- #131 TTS: TTSTrainerWrapper(SFTTrainerWrapper) — TTS fine-tune = next-token
CE over [text][audio-codec-token] chat; per-family emotion templating
(Orpheus/Oute) + codec special-token registration. Pre-encoded chat path
live-validated on SmolLM2-135M-Instruct; live-codec (data.format=audio)
hardware-gated per family.
- #134 BitNet: BitNetTrainerWrapper gated on onebitllms; export --format
bitnet|tq1_0 runs real llama.cpp TQ1_0 ternary GGUF export.
- #136 MoE: apply_moe_expert_quant swaps fused-MoE experts to bnb Linear4bit/
Linear8bitLt (pre-LoRA); train_router_only freezes experts (post-LoRA).
Live-validated on RTX 3050 (dequant err 0.0155).
Review fixes: H1 explicit Params4bit/Int8Params weight-carry; H2 quant
pre-LoRA / freeze post-LoRA + skip PEFT-wrapped modules; M4 device-aware
placement.
Tests 13807 -> 13874 (+69 in test_v07120.py, -2 lifted stubs in test_v0520.py).
2026-06-10 12:37:14 +05:00
Alpamys
70fd5ee9f3
feat(distill,agent,cloud): on-policy MiniLLM + aligned ULD + agent sandbox eval + Modal cloud (v0.71.18)
...
Closes #257 , #258 , #110 , #16 .
- #257 MiniLLM true on-policy rollout: minillm_on_policy_rollout (Gu et al. §3.1
autoregressive teacher-mixed rollout, reverse-KL on full distributions,
grad-to-student-only) + on_policy_term + training.minillm_on_policy /
minillm_rollout_length, wired into DistillTrainer.compute_loss.
- #258 cross-tokenizer ULD wasserstein_aligned: align_token_sequences (difflib
char-span) + aggregate_aligned_logits + uld_aligned_loss for fully-disjoint
tokenizers, wired into DistillTrainer.
- #110 soup agent eval --sandbox: build_eval_stub (base64-embed-as-data) +
run_eval_in_sandbox (v0.25 RLVR isolation + SANDBOX_NETWORK_GUARD) +
classify_sandbox_outcome (ok/tool_error/timeout/arg_error).
- #16 soup train --cloud modal: render a Modal app from soup.yaml (config
base64-embedded), plan-only default, --cloud-submit token-gated; [modal] extra.
+114 tests (tests/test_v07118.py); 13656 -> 13770. Step-6 smoke on real input
(Windows + RTX 3050): Modal stub render, real subprocess sandbox scorecard,
on-policy distill (tiny-gpt2), cross-tokenizer aligned ULD (GPT-2 + Llama).
2026-06-08 23:49:14 +05:00
Alpamys
0d55ba9e46
feat(serve): serve-time MoLE + per-request vector banks + epoch RAFT shuffle (v0.71.17)
...
Closes #259 (soup serve --mole: load base + N frozen task LoRAs + mole_gate.pt,
blend per-token at decode; train writes mole_manifest.json).
Closes #260 (soup serve --bank active user per-request via contextvars.ContextVar,
no cross-request leak; streaming path re-selects in-context).
Closes #253 (data.raft_epoch_shuffle: re-permute golden/distractor docs each epoch;
epoch=0 == legacy order).
Closes #254 (soup diagnose --citation-style / --shuffle-seed into the live probe).
Fix: MoLE train() returns initial_loss/final_loss/total_steps/duration_secs so
task=moe_lora_routing completes cleanly (surfaced by the #259 smoke).
Validated live on SmolLM2-135M (RTX 3050). 13595 -> 13656 tests.
2026-06-08 20:06:43 +05:00
Alpamys
a66e4b9ebe
feat: architecture + distill + adapter-train live wiring (v0.71.12)
...
Lift seven schema-only stubs to live, validated on tiny models:
- #145 distill_mode token|sequence — sequence-level teacher-continuation KD
- #146 classifier LoRA — frozen encoder + adapter (classifier/reranker/cross_encoder)
- #148 LLaMA Pro block expansion — per-arch (Llama/Qwen/Mistral) zero-init blocks
- #158 LongLoRA S2 — shifted-sparse attention on Q/K projections (Llama/Mistral/Qwen/Phi)
- #84 Mixture-of-Depths — per-layer top-k token router (use_mod; Llama/Qwen/Mistral)
- #221 VeRA/VB-LoRA serving — soup serve --bank, per-user delta via X-User-Id header
- #222 MoLE — task=moe_lora_routing, per-token gate over N frozen task LoRAs (gate-only)
Tests 13203 -> 13329 (+126; tests/test_v07112.py). ruff clean, coverage 78.49%.
2026-06-04 19:36:04 +05:00
Alpamys
f316d334bc
feat(rl): live GRPO/RL callbacks — reward-hack, echo-trap, RL ckpt, ULD, MiniLLM, iterative-DPO (v0.71.11)
...
Lifts the v0.70.0 schema-only build_*_callback / build_uld_projection /
run_iterative_dpo stubs. Validated end-to-end on SmolLM2-135M.
Closes #235 , #236 , #237 , #238 , #239 , #240 , #159 , #160
- #235 RewardHackCallback: info_rm cluster-sep / rm_ensemble divergence,
OK/WARN/HACK, halt on HACK. Shared thread-safe RLSignalBuffer captures
per-step rewards by wrapping the reward fns (no TRL monkeypatching).
- #236 ULD: Wasserstein-1 / top-k aligned distill loss in DistillTrainer.
- #237 MiniLLM: teacher-mixed length-normalised reverse-KL + pretrain anchor.
- #238 RLCheckpointCallback: adapter + optimizer.pt + manifest + keep_last prune.
- #239 run_iterative_dpo: sample -> RM-score -> build-pairs -> DPO-train per round.
- #240 EchoTrapCallback: n-gram repetition OK/WARN/TRAP, halt on TRAP.
- #159 one-shot WARNING when a GRPO variant compute_loss falls back to super().
- #160 in-place ref-model EMA (no state_dict round-trip) + 0-overlap warning.
Tests 13142 -> 13203 (+62 in tests/test_v07111.py).
2026-06-04 17:14:08 +05:00
Alpamys
a2287dd6a5
feat(rag): RAFT span-mask trainer + RA-DIT auto-link + live steering + eval citation (v0.71.10)
...
Lifts the v0.62.0 RAG-family schema-only stubs to live, validated on SmolLM2-135M:
- #199 RAFT: data.format=raft trains answer-only (prompt span masked to -100,
[doc-N] citation ids, deterministic doc shuffle by raft_shuffle_seed); rows
whose prompt fills max_length are dropped with a warning. New utils/raft.py +
trainer/raft.py (RaftDataCollator + weighted-CE _RaftTrainer).
- #200 soup ra-dit: one-shot two-stage orchestrator (train retriever -> record
it as the generator's paired retriever -> train generator); a generator-stage
`soup train` with no retriever set auto-links the latest RA-DIT retriever from
the Registry. New utils/ra_dit_run.py + commands/ra_dit.py.
- #201 soup steer train/apply + soup serve --steer: live CAA/ITI/RepE fit from
{positive, negative} pairs + decode-time forward hook (transformers backend).
Lifts the steering.py apply_steering/build_steering_vector stubs.
- #202 soup eval citation + citation-span per-token loss boost + 7th `citation`
failure mode in soup diagnose. New commands/_eval_v07110.py +
diagnose/citation.py.
Review fixes (3 agents, all CRITICAL->LOW): markup-escaped autolink advisory;
shared enforce_under_cwd_and_no_symlink + O_NOFOLLOW on every new file read;
steering-artifact containment; honest RA-DIT docs (records pairing, no weight
fusion); public validate_ra_dit_config_path + render_raft_prompt; repe/iti
require >=2 pairs; eval citation --shuffle-seed.
Full suite: 13034 passed, 106 skipped (13142 collected). ruff clean.
2026-06-03 19:28:45 +05:00
Alpamys
96c339f184
feat(edit): live ROME/MEMIT/AlphaEdit + GRACE + NPO/SimNPO/RMU unlearn (v0.71.9)
...
Closes #193 , #194 , #196 , #197 , #203 .
- #194 utils/edit_kernels.py: covariance-free rank-1 ROME/MEMIT/AlphaEdit;
apply_edit live (load -> optimise residual -> rank-1 update -> save);
edit diff live before/after generation.
- #196 EditGovernorStore SQLite persistence + cross-process lock.
- #197 apply_edit consults the governor (check_can_edit before, record after).
- #203 GraceCodebook + apply_grace_edit + install_grace_hook + Registry kinds.
- #193 utils/unlearn_kernels.py (NPO/SimNPO/RMU) + live UnlearnTrainerWrapper
+ soup train --task unlearn.
Validated on SmolLM2-135M: ROME 0.0016->0.96, NPO/SimNPO forget loss down.
+81 tests (tests/test_v0719.py). 2 review waves, all findings fixed.
2026-06-03 17:04:03 +05:00
Alpamys
1f63393421
feat(v0.71.5): ingest/data/prompt/drift polish
...
Closes #157 , #205 , #207 , #149 , #164 , #163 . Defers #204 (live SaaS pull —
paid accounts, infra-blocked, kept open).
- #164 : get_metric_series falls back to eval_results when metrics is empty
- #163 : build_verdict confidence biased by advise_history (same project+choice,
>=3 precedents); decision never changes
- #207 : shared utils/webhooks.py (SSRF-hardened) + --slack-url/--discord-url on
ingest/prune-prompt/ab/active-sample; ab fires only on a decision
- #205 : soup prune-prompt --tokenizer (token-prefix detect + decode remainder,
boundary-safe)
- #149 : DynamicCurriculumCallback buckets by loss/perplexity percentile;
length keeps round-robin
- #157 : soup data push/forge --hub modelscope|modelers (data score N/A)
107 new tests in tests/test_v0715.py (12474 -> 12581). ruff clean.
2026-06-02 14:34:36 +05:00
Alpamys
514761c89a
feat(env,lock,serve,eval): v0.71.1 — quick wins + wiring (7 closures)
...
Closes #195 #210 #214 #224 #230 #233 #209 .
- soup env fix: print-only install-plan renderer from soup-env.lock
(uv-pip / requirements; non-pip entries surfaced as comments). (#209 )
- soup lock write --env-lock: auto-derive --env-hash from soup-env.lock
via new compute_env_hash (excludes created_at). (#224 )
- soup serve --record-thumbs <db>: capture thumbs-up/down into the
local-RL SQLite + POST /v1/thumbs (transformers backend). (#230 )
- Judge-calibration persistence: JudgeCalibrationReport.to_dict +
write/load_judge_calibration + judge_calibration registry kind; load
re-validates the frozen dataclass with cwd/symlink containment. (#214 )
- soup completions: introspect a base model's real LoRA target modules
(config-only AutoConfig, local_files_only, never networks/raises). (#210 )
- Bundled MUSE + WMDP unlearning eval fixtures; WMDP forget rows ship
REDACTED (Soup never bundles verbatim hazardous content). (#195 )
- build_dag.validate_build_source: cwd-containment + symlink rejection. (#233 )
Review-fix hardening (consolidated python+code+security+tdd, 0 CRIT/0 HIGH):
load_judge_calibration containment + friendly missing-field ValueError;
serve thumbs success-print escape; env_fix --output Optional[str];
empty --env-hash auto-derives; render_install_plan PEP 440 docstring note.
Tests: 12071 -> 12134 (12044 passed, 90 skipped, 2 deselected).
2026-06-01 14:12:30 +05:00
Alpamys
7724353ac6
docs: trim README to a 238-line front door; move feature reference to docs/
...
The README had grown to 5046 lines (195 sections) — roughly one deep-dive per
feature accreted over 70 releases. Split it into a concise front door plus a
public docs/ tree:
- README (5046 -> 238 lines): hero, why, quickstart, config, a Documentation
map, data formats, common commands, models, Docker, requirements, dev.
- docs/*.md: all 185 feature sections preserved verbatim, grouped into 10 themed
guides + an index. Every original line is accounted for (content-conservation
checked); all 235 internal links + anchors verified to resolve.
- un-gitignore docs/ (it was empty); fix a pre-existing dangling
docs/QUANTIZATION.md link; correct the stale `ruff check soup_cli/` ->
`src/soup_cli/` reference in the Development section.
No version bump: docs-only — rides into the 0.71.0 deps-split release.
2026-05-31 20:10:59 +05:00