Commit Graph

260 Commits

Author SHA1 Message Date
Alpamys 0db7404a6a test(train): strip ANSI + widen terminal in reward-hack help assertions (v0.71.26)
The two TestRewardHackMitigationCli help-text assertions grepped the raw
`soup train --help` output for the new flags. CI runners emit color codes
that split long option names into non-contiguous characters (same failure
mode documented in test_eval_gate.py for --gate), so the raw substring check
failed on every OS/Python combo while passing locally on a no-color terminal.

Fix follows the established repo pattern: render with COLUMNS=200 (no option
wrapping) and strip ANSI escapes before the membership check. Verified under
FORCE_COLOR=1: both assertions pass. Test-only change, no version bump.
2026-07-01 18:16:53 +05:00
Alpamys 2ffc3743ae feat(train): --reward-hack-detector / --reward-hack-halt CLI flags (v0.71.26)
Folds in the doc-vs-reality cleanup the plan flagged: the docs referenced
soup train --reward-hack-detector / --reward-hack-halt but they were config-only.
Add both as CLI passthroughs mirroring --reward-hack-mitigation (validate value,
set cfg.training field, accelerate re-exec passthrough) so the docs are true and
the reward-hack CLI is consistent. +4 tests (test_v07126: 180 -> 184; full suite
-> 14788).
2026-07-01 17:02:37 +05:00
Alpamys fa992bf381 fix(train): reward-hack mitigation review fixes (v0.71.26)
Fixes from 5 sequential ECC reviews (python/code/security/tdd/verification).

python-review (2 CRITICAL + HIGH/MED/LOW):
- signal/vote coherence: schema now requires the active detector in
  reward_hack_signals + rejects the inactive detector name (was silently
  dropping the primary signal from the vote).
- integral_clamp is its own field (was wrongly hard-wired to beta_ceil).
- task/backend gate runs before controller-config checks; EMA convention
  corrected; type hints; mutable-list default -> tuple + normalised compare.

code-review (4 HIGH + MED/LOW):
- _prune now trims _saved in sync with disk (rollback can't target a deleted
  checkpoint); bang-bang release_count resets after each relaxation (hysteretic
  descent); EMA formula uses standard convention; _escalate no longer burns a
  recovery attempt on a None target; max_recovery_attempts>=1 required with
  rollback; _action_history capped; on_step_end logs errors once; loud warning
  when the mitigation callback can't attach (was a silent safety-off).

security-review (HIGH + MED):
- restore_checkpoint / save_checkpoint refuse a SYMLINKED optimizer.pt
  (torch.load weights_only=False was an RCE via attacker-placed symlink);
  bool-before-int/float guards on all new numeric fields; reward_hack_signals
  max_length=4; empty-signals guard in the callback.

tdd-review: +13 coverage tests (dead-band hold, shim verbatim-on-error,
conservative boundary, read-only-beta dual-write, escalation postconditions,
both-restore, PID D exact, log cap + concurrency, no-top-level-import, fuzz
field-validity).

Test count 152 -> 180 (+2 POSIX-only symlink skips).
2026-07-01 16:38:00 +05:00
Alpamys eb2edb1a51 feat(train): anti-gaming hardening + reward shaping + give-up explainer (v0.71.26 Part D)
Stage 3 of closed-loop reward-hacking mitigation: the controller itself must
not be gameable.

- schema: 6 Stage-3 tunables (signal_smoothing, smoothing_window,
  conservative_on_disagreement, reward_shaping, shaping_kind, shaping_strength)
  + bool guards; reward_shaping requires a control mode + strength>0.
- reward_hack_control: combine_conservative (disagreement -> MAX, stay cautious);
  detect_reward_distribution_drift (bimodal-collapse heuristic); shape_reward_fn
  + apply_reward_shaping (bounded length/repetition/sentinel penalty over the
  wrap_reward_funcs seam, inner called once, verbatim on shim error);
  explain_giveup (plain-English, mirrors why.py). Callback wires per-signal
  smoothing + conservative vote + opt-in drift guard (keep KL high) + logs the
  give-up explanation on early-stop.
- peft_wiring: thread smoothing/conservative params from tcfg.
- grpo.py / ppo.py: apply_reward_shaping BEFORE the buffer capture.

Includes an adversarial-fuzz suite (sawtooth/step/noisy/flip/out-of-range
traces): bounded output, no unbounded jump, no flap, anti-windup holds.

+35 tests (test_v07126: 117 -> 152).
2026-07-01 15:33:46 +05:00
Alpamys 21557b43c1 feat(train): PID-Lagrangian controller + rollback escalation ladder (v0.71.26 Part C)
Stage 2 of closed-loop reward-hacking mitigation: principled control + safety net.

- schema: 7 Stage-2 tunables (pid_kp/ki/kd, signal_target, rollback,
  rollback_patience, max_recovery_attempts) + bool guard; cross-validators —
  PID/rollback tunables require pid_lagrangian mode; reward_hack_rollback
  requires rl_checkpoint_save_every_steps (a cadence to roll back to).
- reward_hack_control: PIDLagrangianPolicy + pid_step (Stooke et al. — constraint
  'signal <= target' whose multiplier β is a PID law; integral anti-windup clamp,
  output clamp [floor,ceil], never crosses 0). Callback pid_lagrangian path drives
  β via PID + escalation ladder: raise -> rollback last-good checkpoint (after
  rollback_patience HACK steps) -> early-stop (after max_recovery_attempts),
  subsuming halt_on_hack.
- rl_checkpoint: RLCheckpointCallback.restore_checkpoint (reload PEFT adapter via
  set_peft_model_state_dict + optimizer.load_state_dict; best-effort, never raises).
- peft_wiring: build PIDLagrangianPolicy from tcfg; build the RL-checkpoint cb
  FIRST and hand its reference to the mitigation callback for rollback.

+26 tests (test_v07126: 91 -> 117; real-peft adapter round-trip included).
2026-07-01 15:18:07 +05:00
Alpamys 925e600702 feat(train): reversible bang-bang KL controller + hysteresis (v0.71.26 Part B)
Stage 1 of closed-loop reward-hacking mitigation: the first closed loop.

- schema: 8 Stage-1 tunables (beta_floor/ceil, trip/release band, dwell_steps,
  release_patience, kl_gain, signals allowlist) + bounds; extended
  _validate_reward_hack_compat (floor<ceil, release<trip, signal allowlist,
  control-mode XOR ref_model_ema_alpha, footgun-reject tunables while off).
- reward_hack_control: MitigationAction + BangBangPolicy + bang_bang_step (pure
  hysteresis: dwell before trip, release_patience before relax, beta geometric
  x/div kl_gain clamped [floor,ceil], never crosses 0; multi-signal vote).
  Callback kl_control path mutates BOTH trainer.beta and trainer.args.beta
  (GRPO stock + variant) / trainer.args.kl_coef (PPO); seeds beta from trainer;
  length-trend signal.
- peft_wiring: build BangBangPolicy from tcfg for kl_control.
- ppo.py: RLSignalBuffer parity + attach_rl_callbacks(task=ppo) (BETA — GRPO
  gets the on-GPU proof; PPO kl_coef mutation is unit-tested).
- train.py: --reward-hack-mitigation off|log_only|kl_control|pid_lagrangian
  flag + validation + accelerate re-exec passthrough.

+33 tests (test_v07126: 58 -> 91).
2026-07-01 15:05:32 +05:00
Alpamys bcb08ae0ab feat(train): reward-hack mitigation instrumentation + log_only telemetry (v0.71.26 Part A)
Stage 0 of closed-loop reward-hacking auto-mitigation: observe only, no
control action.

- schema: reward_hack_mitigation Literal[off/log_only/kl_control/pid_lagrangian]
  gated to grpo/ppo + non-mlx + requires reward_hack_detector; YAML-1.1 bool
  DWIM (off->"off", on/yes rejected with a quote hint).
- utils/reward_hack_control.py (no top-level torch): MitigationLogWriter
  (mirrors TraceLogWriter: thread-lock, rotation, redaction, cwd containment,
  symlink-reject), ControllerState (frozen), combine_signals, smooth_signal,
  telemetry helpers, RewardHackMitigationCallback (log_only path provably never
  mutates beta).
- reward_hacking.RewardHackCallback: last_drop_pct() accessor for the controller.
- peft_wiring: mitigation callback subsumes the plain detector when a mode is set.
- examples/reward_hacking/rewards.py: synthetic length-hack + sentinel proxies
  decoupled from a held-out true_score (the GPU-experiment fixture).
- trace_logger: public redact_value alias (DRY reuse).

+58 tests (tests/test_v07126.py).
2026-07-01 14:47:40 +05:00
Salil M 14de9ec9d5
feat(recipes): add ready-made SFT recipe for Qwen2.5-Coder-7B-Instruct (#285)
* feat(recipes): add qwen2.5-coder-7b-sft recipe to catalog

* test(recipes): update catalog size assertion to 134 in recipe tests

* test(recipes): update version catalog count to 134 in v0.71.24 tests

* docs(contrib): bump recipe count in project structure overview

* docs(commands): update recipe list count reference to 134

* docs(serve): update Web UI recipe count references in serving docs
2026-06-28 19:22:29 +05:00
Alpamys 6cb1abab8f feat(eval): soup ship — SHIP / DON'T-SHIP verdict (v0.71.25)
Add `soup ship`, a binary SHIP / DON'T-SHIP verdict after fine-tuning: it
SHIPs only when (leg 1) the task metric strictly improved AND (leg 2) no
general benchmark regressed past a forgetting threshold (default 0.05
absolute points) — otherwise DON'T SHIP, even if the task metric looks
great. The moat is leg 2 (catastrophic-forgetting gate) fused with the
task win into one decision. Exit: 0=SHIP, 2=DON'T SHIP, 1=runtime error.

- utils/ship_verdict.py: pure engine (no top-level torch) — frozen
  TaskWin/BenchmarkDelta/ShipVerdict + decide_ship (single source of
  truth for the threshold) + compute_benchmark_deltas + render/serialize.
- commands/ship.py: Typer command; --evidence offline path + live
  metric/judge leg-1 + mini(default)/lm-eval leg-2; --baseline/--output.
- Reuses run_eval / JudgeEvaluator / ForgettingDetector / resolve_baseline
  / _run_lm_eval / live_eval.make_generator.
- Hardening: --evidence O_NOFOLLOW + size cap; --task-eval cwd-contained;
  --judge-model urlparse SSRF guard; lm-eval model_args injection guard;
  --general-suite bounded.

Schema (ShipConfig) deferred — v1 is CLI-only. Pairwise judge win-rate is
a planned fast-follow. +79 tests (14514 -> 14593).
2026-06-27 23:37:35 +05:00
Kondamwar Akshaya Shrikant e4065d8741
Fix friendly error mappings 272 (#282)
* Improve friendly error mappings and tests

* Improve friendly error mappings and tests
2026-06-22 15:54:55 +05:00
Alpamys 418f86390a feat(recipes): 2026 model-family expansion — 17 SFT recipes, catalog 116→133 (v0.71.24)
Add ready-made SFT recipes for the open-weight models released Feb–Jun 2026,
each base repo-ID verified to resolve on Hugging Face:
- Qwen3.5 0.8B/2B/4B/9B/27B + MoE 35B-A3B/122B-A10B/397B-A17B (Apache-2.0)
- Qwen3.6 27B + 35B-A3B (Apache-2.0)
- DeepSeek-V4 Flash/Pro (MIT), GLM-5.1 (MIT)
- Kimi-K2.5/K2.6 (Modified MIT), MiniMax-M3 (MiniMax Community License)
- Mistral-Large-3 (Apache-2.0, 675B/41B-active multimodal MoE)

Fix stale glm-5-sft repo-ID THUDM/glm-5 -> zai-org/GLM-5 (org migration).
+220 tests (tests/test_v07124.py). Catalog count 116 -> 133.
2026-06-21 13:00:47 +05:00
Shatakshi Prasad 1d1d12c3a7
test: add unit tests for warmup.py helpers (#274)
* test: add unit tests for warmup.py helpers

* fix: ruff lint fixes (import order + trailing newline)

* fix: add trailing newline

* style: ruff --fix
2026-06-20 21:38:05 +05:00
Alpamys fcf4b33394 feat(train): native Spectrum targeted training — soup spectrum scan + training.unfrozen_parameters (v0.71.23)
Closes #266. `soup spectrum scan` streams safetensors per-tensor (no model
load, CPU-friendly) and computes a singular-value SNR per weight matrix
(Marchenko-Pastur, arXiv:2406.06623), emitting a ready-to-paste
training.unfrozen_parameters patch. The SFT trainer freezes all params then
unfreezes the matched set (full FT, LoRA off).

- utils/spectrum_scan.py: pure-numpy transpose-invariant SNR kernel +
  per-tensor safetensors streaming (2^31 SVD cap, symlink skip) + cache
  (~/.soup/spectrum, SOUP_SPECTRUM_CACHE_DIR containment) + hardened
  hubs.snapshot_download.
- commands/spectrum.py: soup spectrum scan (SNR table + YAML patch).
- schema: training.unfrozen_parameters (caps/NUL/invalid-regex/ReDoS reject)
  + gates (sft/transformers/text/quantization=none; mutually exclusive with
  LoRA features / freeze_layers / freeze_ratio / train_router_only /
  expand_layers).
- trainer/sft.py: full-FT branch via apply_unfrozen_parameters +
  enable_input_require_grads (fixes grad-checkpointing through frozen
  embeddings).

Existing spectrum trainer-plugin untouched (back-compat); LISA -> #267.
Live-validated on Windows + RTX 3050: CPU scan of SmolLM2-135M + top-25%
unfrozen full-FT train (loss 3.455 -> 0.719).

Note: after the version bump the editable install metadata was stale
(0.71.17); pip install -e . --force-reinstall --no-deps re-synced it so
test_cli_subprocess::test_version passes.

+94 tests in tests/test_v07123.py (14184 -> 14278).
2026-06-12 17:40:46 +05:00
Alpamys ccd5c80e4d feat(perf): MiniLLM/MoLE KV-cache + deploy-measure live factories + live-codec TTS (v0.71.22)
#263 MiniLLM on-policy KV-cache (PEFT-unwrap probe activates the cache for LoRA
students; per-step single-token forward), #262 serve --mole per-adapter KV cache
(fresh per generate, no cross-request leak, byte-identical to no-cache), #143
deploy-autopilot live generator factories (baseline scored once + up-front
candidate validation; injected seams retained), #265-partial live-codec TTS
(soundfile.info pre-probe + O_NOFOLLOW; SNAC Orpheus encode validated).

Review: 1 HIGH + 5 MEDIUM + ~10 LOW fixed across 2 review waves + verification +
step-6 live smoke (Windows + RTX 3050). Tests 14084 -> 14184 (+100 in
tests/test_v07122.py; 293 files). Full suite 14067 passed / 117 skipped, exit 0.
2026-06-10 21:30:01 +05:00
Alpamys ed5fc3a8b3 feat(precision,rollout): live fp8/nvfp4 + vLLM sleep + openenv rollout + apple-adapter + delinearize-llama4 (v0.71.21)
Closes #141, #124, #125, #228, #97.

- #141: apply_fp8_attention (torchao float8 on attention projections, Hopper
  gate) + apply_nvfp4 (NVFP4Config, Blackwell gate); partial-conversion honesty;
  wired into the v0.28 speed/memory pipeline with yellow-advisory degrade.
- #124: vllm_sleep_mode live - create_vllm_engine(sleep_mode=True) +
  vllm_sleep_cycle ctx (wake in finally) + TRL GRPOConfig hook probe.
- #125: openenv rollout fully live via training.rollout_func module:fn
  resolver; rows replace the prompt dataset; art/ruler/nemo_gym honest dep
  gates + _EXTERNAL_ROLLOUT_RUNNERS seam. Real GRPO train on SmolLM2-135M.
- #228: convert_apple_adapter live - PEFT LoRA <-> mlx-lm (both matrices
  transpose, bf16 upcast, adapters.safetensors + num_layers, npz legacy read,
  np.ascontiguousarray fix for safetensors non-contiguous mangling);
  *-to-apple upstream-gated exit 3.
- #97: delinearize-llama4 live - [E*din,dout] -> [E,din,dout] per shard,
  config.json expert-count probe + --num-experts, sidecar copy, atomic writes.

Review waves: 3 HIGH + ~8 MEDIUM + ~12 LOW fixed.
Tests: 13874 -> 14084 (+210 in tests/test_v07121.py).
Full suite: 13967 passed, 117 skipped. ruff clean.
2026-06-10 16:29:52 +05:00
Alpamys a4dfbb308c feat(trainer): live TTS / BitNet / MoE-expert-quant trainers (v0.71.20)
Lift three v0.52.0 schema-only NotImplementedError stubs to real code.

- #131 TTS: TTSTrainerWrapper(SFTTrainerWrapper) — TTS fine-tune = next-token
  CE over [text][audio-codec-token] chat; per-family emotion templating
  (Orpheus/Oute) + codec special-token registration. Pre-encoded chat path
  live-validated on SmolLM2-135M-Instruct; live-codec (data.format=audio)
  hardware-gated per family.
- #134 BitNet: BitNetTrainerWrapper gated on onebitllms; export --format
  bitnet|tq1_0 runs real llama.cpp TQ1_0 ternary GGUF export.
- #136 MoE: apply_moe_expert_quant swaps fused-MoE experts to bnb Linear4bit/
  Linear8bitLt (pre-LoRA); train_router_only freezes experts (post-LoRA).
  Live-validated on RTX 3050 (dequant err 0.0155).

Review fixes: H1 explicit Params4bit/Int8Params weight-carry; H2 quant
pre-LoRA / freeze post-LoRA + skip PEFT-wrapped modules; M4 device-aware
placement.

Tests 13807 -> 13874 (+69 in test_v07120.py, -2 lifted stubs in test_v0520.py).
2026-06-10 12:37:14 +05:00
Alpamys f51331d637 feat(quant,multipack): vision/audio Quant Menu + multipack FSDP sharding (v0.71.19)
#81 Quant Menu for vision/audio modality
- config/schema: drop the `modality != "text"` rejection in
  _validate_quant_menu_supported_tasks (mlx-backend gate retained) so the full
  Quant Menu (gptq/awq/hqq:Nbit/aqlm/eetq/mxfp4/fp8) applies to vision+audio.
- trainer/sft: _setup_vision_transformers + _setup_audio_transformers call
  build_quantization_config_for_loader (strict superset of the inline 4bit/8bit
  BNB blocks they replaced); drop BitsAndBytesConfig import; retain the
  prepare_model_for_kbit_training gate on (4bit,8bit,mxfp4).

#80 multipack DataLoader sharding under FSDP/DeepSpeed/DDP
- utils/multipack_trainer: get_train_dataloader routes the multipack DataLoader
  through accelerator.prepare when num_processes > 1 so accelerate's
  BatchSamplerShard shards whole FFD bins across ranks (even_batches=True avoids
  the epoch-boundary collective hang; seed identical across ranks, no `+ rank`).
  Single-process path unchanged. Defence-in-depth guard against an unconfigured
  MagicMock num_processes.

Tests: +37 in tests/test_v07119.py (13770 -> 13807). ruff clean.
Supersedes the v0.40.5 vision/audio Quant-Menu and v0.40.4 multipack-FSDP
known-limitations. Full multi-GPU validation remains an INFRA-BLOCKED QA item.
2026-06-09 12:37:50 +05:00
Alpamys ae6a18e7d7 test(v0.71.18): ANSI-strip the --minillm-on-policy --help assertion for CI FORCE_COLOR
TestTrainCliMinillmOnPolicy::test_flag_in_help removed newlines + spaces but
not ANSI codes, so under CI FORCE_COLOR the Rich-rendered long flag (ANSI codes
between the dashes) failed the substring check. Strip ANSI + remove all
whitespace before the check, matching the cloud/sandbox help tests + the
v0.71.17 precedent. Verified under FORCE_COLOR=1 (114 passed).
2026-06-09 00:05:36 +05:00
Alpamys 70fd5ee9f3 feat(distill,agent,cloud): on-policy MiniLLM + aligned ULD + agent sandbox eval + Modal cloud (v0.71.18)
Closes #257, #258, #110, #16.

- #257 MiniLLM true on-policy rollout: minillm_on_policy_rollout (Gu et al. §3.1
  autoregressive teacher-mixed rollout, reverse-KL on full distributions,
  grad-to-student-only) + on_policy_term + training.minillm_on_policy /
  minillm_rollout_length, wired into DistillTrainer.compute_loss.
- #258 cross-tokenizer ULD wasserstein_aligned: align_token_sequences (difflib
  char-span) + aggregate_aligned_logits + uld_aligned_loss for fully-disjoint
  tokenizers, wired into DistillTrainer.
- #110 soup agent eval --sandbox: build_eval_stub (base64-embed-as-data) +
  run_eval_in_sandbox (v0.25 RLVR isolation + SANDBOX_NETWORK_GUARD) +
  classify_sandbox_outcome (ok/tool_error/timeout/arg_error).
- #16 soup train --cloud modal: render a Modal app from soup.yaml (config
  base64-embedded), plan-only default, --cloud-submit token-gated; [modal] extra.

+114 tests (tests/test_v07118.py); 13656 -> 13770. Step-6 smoke on real input
(Windows + RTX 3050): Modal stub render, real subprocess sandbox scorecard,
on-policy distill (tiny-gpt2), cross-tokenizer aligned ULD (GPT-2 + Llama).
2026-06-08 23:49:14 +05:00
Alpamys 73fdbbb1b4 test(v0.71.17): ANSI-strip CLI --help assertions for CI FORCE_COLOR
Typer/Rich injects ANSI codes under CI FORCE_COLOR that split flag names
(--mole, --citation-style) and wrap panel text, breaking raw substring
asserts. Normalise output (strip ANSI + collapse whitespace) before matching.
Same pattern as the v0.71.1 --record-thumbs fix.
2026-06-08 20:21:05 +05:00
Alpamys 0d55ba9e46 feat(serve): serve-time MoLE + per-request vector banks + epoch RAFT shuffle (v0.71.17)
Closes #259 (soup serve --mole: load base + N frozen task LoRAs + mole_gate.pt,
blend per-token at decode; train writes mole_manifest.json).
Closes #260 (soup serve --bank active user per-request via contextvars.ContextVar,
no cross-request leak; streaming path re-selects in-context).
Closes #253 (data.raft_epoch_shuffle: re-permute golden/distractor docs each epoch;
epoch=0 == legacy order).
Closes #254 (soup diagnose --citation-style / --shuffle-seed into the live probe).

Fix: MoLE train() returns initial_loss/final_loss/total_steps/duration_secs so
task=moe_lora_routing completes cleanly (surfaced by the #259 smoke).

Validated live on SmolLM2-135M (RTX 3050). 13595 -> 13656 tests.
2026-06-08 20:06:43 +05:00
Alpamys cb1fa2e69f test(local-rl): de-flake stamp-before-train concurrent-thumbs test on Windows
Pre-existing v0.71.13 test flaked on windows-latest CI: record_thumb stamps
time.time() and count_new_thumbs_since uses strict `>`, so on Windows' ~15.6ms
clock resolution the two mid-train thumbs could land in the same tick as
run_started (ts == run_started, dropped). Sleep one clock tick at the start of
slow_train so the thumbs are strictly later. Prod semantics unchanged; tests-only.
2026-06-07 16:41:31 +05:00
Alpamys 28a2d8cf82 feat(edit): GPT-2 Conv1D edits, covariance ROME, atomic governor, Mixtral LongLoRA (v0.71.16)
Knowledge edit depth — 4-issue patch (validated on real gpt2 + SmolLM2-135M, RTX 3050).

- #251 GPT-2 (transformer.h / mlp.c_proj Conv1D) support in ROME/MEMIT/AlphaEdit:
  transpose-aware _rank1_update + _alphaedit_project + MEMIT dim-check + PEFT unwrap.
- #250 covariance-preconditioned ROME via --cov-corpus (estimate_key_covariance +
  C^-1 k* solve; fail-loud reject for non-rome; cwd-contained O_NOFOLLOW loader).
- #252 atomic EditGovernor increment: save_state merges this run's delta under the
  cross-process lock (baseline-delta, mirrors namespace_pin).
- #147 Mixtral in the LongLoRA allowlist (is_mixtral_model + MixtralAttention
  forward override + _SEPARATE_QKV_FAMILIES).

Tests: 13511 -> 13595 (+81 in tests/test_v07116.py). Full suite green; ruff clean.
2026-06-07 16:11:56 +05:00
Alpamys d5f98c67d5 fix(loop): config-render, cmaes base-reuse, cost gate, energy hand-off, rank guard (v0.71.15)
#261 iterative_dpo._default_train_fn rendered output as a {dir: ...} mapping
that SoupConfig rejected; render it flat so the spawned soup train succeeds.
#246 CMA-ES merge now loads the base model once and reuses it across the
candidate population (_CachedBaseScorer) instead of reloading per candidate.
#245 soup loop estimate_cost wires run_cost.estimate_run_cost_usd off the last
completed run instead of a 0.0 placeholder; never crashes the daemon.
#244 soup train --track-energy --energy-out persists the measurement JSON so
soup bom emit --energy can attach it to an ML-BOM.
#170 --diagnose-gate is RANK-aware: gate once per cluster (RANK==0), not per node.

Tests: 13476 -> 13511 (+35 in tests/test_v07115.py). Validated end-to-end on
SmolLM2-135M / RTX 3050.
2026-06-07 13:58:18 +05:00
Alpamys 79f01a52a8 fix(bom): emit energy in SPDX output too + strengthen energy tests
PR #256 attached energy only to CycloneDX; #244's contract is "both
outputs". Add _energy_annotations() so the SPDX model package carries
energy as OTHER annotations (same soup:<field>=<value> naming). Strengthen
the happy-path test to assert energy actually lands in the BOM (not just
that a file is written) and add a --format both case asserting energy in
BOTH cdx + spdx.
2026-06-06 19:11:38 +05:00
gittihub-jpg 8254091e44
feat(bom): thread --track-energy measurement into soup bom emit (#256)
Closes #244

Adds --energy flag to soup bom emit: validates the measurement JSON (cwd containment + symlink rejection + JSON + EnergyMeasurement shape) and attaches energy properties to both CycloneDX and SPDX outputs.

Co-authored-by: gittihub-jpg <gittihub-jpg@users.noreply.github.com>
2026-06-06 19:08:43 +05:00
Alpamys 6ec36b5ff6 feat: live FSDP-shard consolidation + serve KV-cache type + ONNX QA (v0.71.14)
Close the doable tail of the export-QA + deferred-stub family.

- #96 consolidate_shards: lazy-torch safetensors merge, TOCTOU-hardened
  (enforce_under_cwd_and_no_symlink, weights_only=True, atomic_write_bytes),
  16 GiB/shard cap, dup-key shape-conflict reject. --plan-only flag.
- #140 apply_kv_cache_type -> KvCacheRuntime; soup serve --kv-cache-type
  (bf16/f16 dtype, q8_0 quantized cache + hqq probe, fp8 Hopper-gated).
- #71 ONNX export QA: tiny-GPT2 PASS, TinyLlama host-RAM-bound (qa doc).
- #70/#72/#144/#74/#79 deferred to INFRA-BLOCKED tail (kept open).

Tests 13430 -> 13476 (+46). ruff clean. Suite green, 78.52% cov.
2026-06-05 13:39:36 +05:00
Alpamys 3ac9e305ce test(ci): harden flaky MiniLLM anchor test + restore coverage network-free
CI failed on the HF-rate-limited runners: test_anchor_term_with_file did a
live from_pretrained that 429'd, so it failed AND its unique MiniLLM-anchor
lines went uncovered, tipping the 77% gate to 76.77% on exactly those jobs
(macos + 3.11 stayed green where the cache warmed).

- Skip test_anchor_term_with_file on OSError (offline / rate-limited) instead
  of failing.
- Add test_anchor_term_with_fake_model: a fake tokenizer + tiny nn.Module
  exercise the identical _load_anchor + anchor_term lines with no network, so
  coverage no longer depends on HF availability.
- Add TestReachableInternals cushion (prompt_compile._resolve_metric,
  prompt_distill._build_provider_fn + default-provider wiring) so the gate
  sits comfortably above 77% (the DSPy/TextGrad/GEPA optimiser bodies are
  uncoverable without the [compile] extra).

Tests 13424 -> 13430.
2026-06-04 22:41:22 +05:00
Alpamys f528da5328 feat(prompt-compile): live soup compile / distill-prompt / compile-tools / local-rl train (v0.71.13)
Lift the v0.68.0 deferred-stub family to live (closes #225, #226, #227, #229):

- #229 local-rl train --once: harvest thumbs -> DPO/KTO/ORPO train via a
  soup train subprocess (argv list, no shell); state table tracks last_train_at
  (skip-on-no-new-thumbs + skip-on-insufficient-pairs); no --once renders a
  systemd/launchd nightly scheduler scaffold. New local_rl_scheduler.py.
- #226 distill-prompt: call the teacher once per trace (Ollama/Anthropic/vLLM)
  and write a real dataset (sft/kl -> messages; preference -> chosen/rejected).
- #225 compile / #227 compile-tools: live DSPy/GEPA/TextGrad dispatch behind the
  new [compile] extra with a friendly ImportError when absent; injectable seams.

Security: reject \n/\r in the model id + shell-quote ExecStart args (systemd
injection defence). Fix: render train output as a plain string (schema-valid),
with a regression test against SoupConfig.

Tests 13329 -> 13424. Smoked end-to-end: real DPO train on SmolLM2-135M (RTX 3050)
+ real Ollama teacher distillation.
2026-06-04 22:14:43 +05:00
Alpamys a66e4b9ebe feat: architecture + distill + adapter-train live wiring (v0.71.12)
Lift seven schema-only stubs to live, validated on tiny models:

- #145 distill_mode token|sequence — sequence-level teacher-continuation KD
- #146 classifier LoRA — frozen encoder + adapter (classifier/reranker/cross_encoder)
- #148 LLaMA Pro block expansion — per-arch (Llama/Qwen/Mistral) zero-init blocks
- #158 LongLoRA S2 — shifted-sparse attention on Q/K projections (Llama/Mistral/Qwen/Phi)
- #84 Mixture-of-Depths — per-layer top-k token router (use_mod; Llama/Qwen/Mistral)
- #221 VeRA/VB-LoRA serving — soup serve --bank, per-user delta via X-User-Id header
- #222 MoLE — task=moe_lora_routing, per-token gate over N frozen task LoRAs (gate-only)

Tests 13203 -> 13329 (+126; tests/test_v07112.py). ruff clean, coverage 78.49%.
2026-06-04 19:36:04 +05:00
gittihub-jpg 8d3a7640c8
feat(build): manifest-level dotted-path custom transforms (#255)
Closes #249

Extends resolve_transform to import module:func dotted paths (lazy, lru-cached, arity-validated). Trusted-input posture documented in docs/data.md.

Co-authored-by: gittihub-jpg <gittihub-jpg@users.noreply.github.com>
2026-06-04 17:37:12 +05:00
Alpamys f316d334bc feat(rl): live GRPO/RL callbacks — reward-hack, echo-trap, RL ckpt, ULD, MiniLLM, iterative-DPO (v0.71.11)
Lifts the v0.70.0 schema-only build_*_callback / build_uld_projection /
run_iterative_dpo stubs. Validated end-to-end on SmolLM2-135M.

Closes #235, #236, #237, #238, #239, #240, #159, #160

- #235 RewardHackCallback: info_rm cluster-sep / rm_ensemble divergence,
  OK/WARN/HACK, halt on HACK. Shared thread-safe RLSignalBuffer captures
  per-step rewards by wrapping the reward fns (no TRL monkeypatching).
- #236 ULD: Wasserstein-1 / top-k aligned distill loss in DistillTrainer.
- #237 MiniLLM: teacher-mixed length-normalised reverse-KL + pretrain anchor.
- #238 RLCheckpointCallback: adapter + optimizer.pt + manifest + keep_last prune.
- #239 run_iterative_dpo: sample -> RM-score -> build-pairs -> DPO-train per round.
- #240 EchoTrapCallback: n-gram repetition OK/WARN/TRAP, halt on TRAP.
- #159 one-shot WARNING when a GRPO variant compute_loss falls back to super().
- #160 in-place ref-model EMA (no state_dict round-trip) + 0-overlap warning.

Tests 13142 -> 13203 (+62 in tests/test_v07111.py).
2026-06-04 17:14:08 +05:00
Alpamys 6df553cb4e test(ci): ANSI-strip --steer/--output/--top-k help asserts (FORCE_COLOR-robust)
CI renders Typer help with ANSI colour codes under FORCE_COLOR that split the
leading `--` from the flag name, so raw-substring assertions on `--steer`
(serve) and `--output`/`--top-k` (steer train) passed locally but failed in CI.
Strip ANSI before the membership check (same fix as the v0.71.1 --record-thumbs
help assert). Local + FORCE_COLOR=1: 142/142 pass. tests-only, no version bump.
2026-06-03 23:00:13 +05:00
Alpamys a2287dd6a5 feat(rag): RAFT span-mask trainer + RA-DIT auto-link + live steering + eval citation (v0.71.10)
Lifts the v0.62.0 RAG-family schema-only stubs to live, validated on SmolLM2-135M:

- #199 RAFT: data.format=raft trains answer-only (prompt span masked to -100,
  [doc-N] citation ids, deterministic doc shuffle by raft_shuffle_seed); rows
  whose prompt fills max_length are dropped with a warning. New utils/raft.py +
  trainer/raft.py (RaftDataCollator + weighted-CE _RaftTrainer).
- #200 soup ra-dit: one-shot two-stage orchestrator (train retriever -> record
  it as the generator's paired retriever -> train generator); a generator-stage
  `soup train` with no retriever set auto-links the latest RA-DIT retriever from
  the Registry. New utils/ra_dit_run.py + commands/ra_dit.py.
- #201 soup steer train/apply + soup serve --steer: live CAA/ITI/RepE fit from
  {positive, negative} pairs + decode-time forward hook (transformers backend).
  Lifts the steering.py apply_steering/build_steering_vector stubs.
- #202 soup eval citation + citation-span per-token loss boost + 7th `citation`
  failure mode in soup diagnose. New commands/_eval_v07110.py +
  diagnose/citation.py.

Review fixes (3 agents, all CRITICAL->LOW): markup-escaped autolink advisory;
shared enforce_under_cwd_and_no_symlink + O_NOFOLLOW on every new file read;
steering-artifact containment; honest RA-DIT docs (records pairing, no weight
fusion); public validate_ra_dit_config_path + render_raft_prompt; repe/iti
require >=2 pairs; eval citation --shuffle-seed.

Full suite: 13034 passed, 106 skipped (13142 collected). ruff clean.
2026-06-03 19:28:45 +05:00
Alpamys e27bf81eb7 test(ci): assert governor-db null-byte rejection on validator directly (POSIX-safe)
POSIX os.putenv forbids null bytes in env values, so
monkeypatch.setenv(SOUP_EDIT_GOVERNOR_DB, 'x\x00.db') raised
ValueError at setenv time on ubuntu/macos before the code under
test ran (windows tolerated it). Assert _validate_governor_db_override
rejects the null byte directly; the validated-None fallback branch is
already covered cross-platform by test_env_override_out_of_bounds_falls_back.
2026-06-03 17:18:33 +05:00
Alpamys 96c339f184 feat(edit): live ROME/MEMIT/AlphaEdit + GRACE + NPO/SimNPO/RMU unlearn (v0.71.9)
Closes #193, #194, #196, #197, #203.

- #194 utils/edit_kernels.py: covariance-free rank-1 ROME/MEMIT/AlphaEdit;
  apply_edit live (load -> optimise residual -> rank-1 update -> save);
  edit diff live before/after generation.
- #196 EditGovernorStore SQLite persistence + cross-process lock.
- #197 apply_edit consults the governor (check_can_edit before, record after).
- #203 GraceCodebook + apply_grace_edit + install_grace_hook + Registry kinds.
- #193 utils/unlearn_kernels.py (NPO/SimNPO/RMU) + live UnlearnTrainerWrapper
  + soup train --task unlearn.

Validated on SmolLM2-135M: ROME 0.0016->0.96, NPO/SimNPO forget loss down.
+81 tests (tests/test_v0719.py). 2 review waves, all findings fixed.
2026-06-03 17:04:03 +05:00
Alpamys eb7655f81e test(ci): ANSI-strip help/error asserts in test_v0718 (FORCE_COLOR-robust)
CI (FORCE_COLOR) makes Rich/Typer split flag tokens at colorized hyphens
(--auto-download -> -auto -download) and auto-highlight `=` in error text
(name=path), so contiguous-substring asserts fail. Add the _clean_help helper
(strip ANSI + all whitespace, matching the v0.71.1 / test_v0717 pattern) and
apply it to the sae-diff / train / sleeper / interference --help asserts plus
the bad-adapter-spec name=path error assert. Reproduced + verified with
FORCE_COLOR=1 locally. No source change; test count unchanged.
2026-06-03 15:19:53 +05:00
Alpamys 823456c1a5 feat(probe): real probe weights, SAE auto-download, truth/harm, interference --measure, capture-activations (v0.71.8)
Closes #216, #217, #218, #219. Partial #215 (calibrated vectors upstream-gated).

- #215 probe_kernel.py: compute_contrast_probe + load_probe_weights
  (.npz/.npy/.safetensors, O_NOFOLLOW, allow_pickle=False, cwd-contained);
  soup probe sleeper --weights. Synthetic seed fallback retained.
- #216 hubs.snapshot_download (SSRF-hardened, home/cwd/tmp cache, TOFU gate)
  + sae_diff.download_sae (allowlist-before-network + symlink-escape guard);
  soup probe sae-diff --auto-download.
- #217 truth_probe.py + harm_probe.py over probe_kernel; soup probe truth/harm;
  probe pack ships truth+harm per base.
- #218 interference_live.measure_interference_losses (live PEFT multi-adapter,
  add_weighted_adapter cat off-diagonal); soup probe interference --measure.
- #219 live_eval.extract_layer_activations + resolve_layer_module PEFT-fallback;
  soup train --capture-activations writes <output>/activations/activations.json.

Test count 12771 -> 12917 (+146 in tests/test_v0718.py). Step-6 smoke on
SmolLM2-135M (RTX 3050) green; caught + fixed a PEFT-wrapper layer-resolution bug.
2026-06-03 15:00:27 +05:00
Alpamys f0118ccffe test(ci): ANSI-strip help-text asserts in test_v0717 (FORCE_COLOR-robust) 2026-06-03 11:41:53 +05:00
Alpamys f097528ac0 feat(eval): live eval runners — advise/tunability/capability/behavior/diagnose (v0.71.7)
Closes #161, #162, #208, #211, #212, #165.

New utils/live_eval.py shared model-loading layer (lazy torch/transformers/peft):
load_model_and_tokenizer, make_generator/make_multi_generator, compute_eval_loss,
lora_probe, measure_logit_agreement, token_f1.

- #161 soup advise --probe-model: live zero/few-shot token-F1 + LoRA probe
- #162 base_model_proximity via held-out logit agreement
- #208 soup tunability --live: per-candidate LoRA probe
- #211 soup eval capability --live --model: lm-eval-harness per task (per-task isolation)
- #212 soup eval behavior --base-model: live pre/post battery diff
- #165 soup diagnose --base-model: utils/diagnose/live.py runs all 6 probes live

Heuristic/neutral paths preserved when no model is supplied. Both new JSONL
readers open with O_NOFOLLOW after cwd-containment (TOCTOU close). +68 tests
(12703 -> 12771). Smoked end-to-end on SmolLM2-135M (RTX 3050).
2026-06-03 00:05:36 +05:00
Alpamys a1463bf716 feat(v0.71.6): live build runner + Magpie generator + 2PL/3PL IRT + augment fix
Lift the v0.69.0 deferred stubs to live + extend IRT + fix a real bug:

- #231 soup build materialises (5 built-in transforms, table/view/incremental
  with SQLite-tracked config-fingerprint cache key, atomic JSONL, --output-dir)
- #232 soup data gen-magpie live (ollama/vllm raw-completion harvest; anthropic
  rejected; optional --quality-filter; dedup-before-response)
- #167 tokenizer-aware memorization probe (sub-word/BPE overlap, library-only)
- #213 soup eval irt-subset --model 2pl|3pl (joint coordinate-ascent MLE)
- #75 fix soup data augment --provider ollama|vllm ImportError + QA log

Security: validate_ollama_url/validate_vllm_url reject 0.0.0.0; augment output
containment+symlink reject; magpie response-body cap.

Tests 12581 -> 12703 (+122 in tests/test_v0716.py). Full suite green, ruff clean.
2026-06-02 22:06:12 +05:00
Alpamys 1f63393421 feat(v0.71.5): ingest/data/prompt/drift polish
Closes #157, #205, #207, #149, #164, #163. Defers #204 (live SaaS pull —
paid accounts, infra-blocked, kept open).

- #164: get_metric_series falls back to eval_results when metrics is empty
- #163: build_verdict confidence biased by advise_history (same project+choice,
  >=3 precedents); decision never changes
- #207: shared utils/webhooks.py (SSRF-hardened) + --slack-url/--discord-url on
  ingest/prune-prompt/ab/active-sample; ab fires only on a decision
- #205: soup prune-prompt --tokenizer (token-prefix detect + decode remainder,
  boundary-safe)
- #149: DynamicCurriculumCallback buckets by loss/perplexity percentile;
  length keeps round-robin
- #157: soup data push/forge --hub modelscope|modelers (data score N/A)

107 new tests in tests/test_v0715.py (12474 -> 12581). ruff clean.
2026-06-02 14:34:36 +05:00
Alpamys 76fdd848cb test(ci): ANSI-strip help-text asserts in test_v0714 (FORCE_COLOR-robust)
Rich splits `--pre-wired` / `--pack-cans` / `--push` with ANSI escapes under
CI FORCE_COLOR; _clean_help() strips them before the substring check (same
fix family as v0.71.1/v0.71.3). No src change.
2026-06-02 12:37:37 +05:00
Alpamys 5652215d4a feat(adapters,loop): v0.71.4 — live canary verdict + cmaes merge + PR push + pre-wired loop + can lineage + branch↔registry
Closes #172, #173, #176, #177, #220, #223.

- #172 soup adapters merge --canary/--strict-verdict: live OK/MINOR/MAJOR verdict (was UNKNOWN stub)
- #220 soup adapters merge --strategy cmaes: live merge→score→write-best loop (was plan-only)
- #223 soup adapters pr --push owner/repo#N: post PR comment via gh api
- #176 soup loop --pre-wired: real traces→DPO→eval-gate→canary stages
- #177 soup loop --pack-cans / replay --extract: iterations as Soup Cans + Registry lineage DAG
- #173 soup adapters branch --from-registry / --attach-to-registry

Security: backdoor-scan + license gates now run for ALL merge strategies (incl cmaes);
loop canary deploy restricted to loopback/RFC1918; gh child env from allowlist;
canary read uses O_NOFOLLOW+fstat (TOCTOU); pack-entry failure rolls back registry entry.

Tests: 12342 → 12474 (+130 in tests/test_v0714.py).
2026-06-02 12:25:01 +05:00
Alpamys 22d5c4f226 test(ci): ANSI-robust help asserts + POSIX-safe audit test (v0.71.3)
CI (FORCE_COLOR) renders --track-energy / --no-audit-log as split ANSI colour
segments, and monkeypatch.setenv with a null byte raises at setup on POSIX
(Windows tolerated both). Strip ANSI via a shared `_plain()` helper for every
--help substring assert, and rewrite the never-raises audit test to monkeypatch
append_audit_event to throw instead of injecting a null-byte env path.
2026-06-01 19:25:33 +05:00
Alpamys 21a2bf8e8c feat(governance,energy): v0.71.3 — annex PDF, audit auto-log, energy hook, can v3, airgap receipt
Closes #180 #181 #182 #183 #184 #188.

- #180 EnergyTracker (codecarbon offline) + `soup train --track-energy`; [carbon] extra
- #181 PDF Annex XI/XII (reportlab) + paths.atomic_write_bytes; [pdf] extra
- #182 Soup Can manifest v3 + attestations field + `can pack --attest`
- #183 per-command audit-log auto-instrumentation (--no-audit-log / SOUP_NO_AUDIT_LOG)
- #184 auto-populate Annex top_domains from the training JSONL
- #188 embed repro-receipt into `soup airgap-bundle`

New [pdf]+[carbon] extras (reportlab also in [dev]). +83 tests (12259 -> 12342),
79.10% coverage. Reviewed (security/code/python/tdd): security M1 raw-size gate on
--attest before parse, python H1/H2 type hints, +22 negative tests.
2026-06-01 19:12:24 +05:00
Alpamys 5d7828d40b test(ci): ANSI-robust help-text asserts in test_v0712 (v0.71.2)
CI installs [dev] with FORCE_COLOR, so Rich colorizes Typer --help and
splits an option name like --key into ANSI-wrapped segments
(\x1b[1;36m-\x1b[0m\x1b[1;36m-key\x1b[0m). The 4 raw-substring help
asserts passed locally (no color) but failed on all 9 CI test jobs.

Add a module-level _strip_ansi() helper and route the sign/verify/merge/
attest-emit --help substring checks through it (mirrors the v0.71.1
test_serve --record-thumbs fix). Confirmed locally under FORCE_COLOR=1:
all 4 pass; ANSI-strip alone is sufficient (no flag line-wraps).

Test-only change on the unreleased v0.71.2 — no version bump.
2026-06-01 17:36:17 +05:00
Alpamys 9300ee3412 feat(governance,sign): v0.71.2 — ed25519 signing + supply-chain gates
ed25519 signing (#179/#185 ed25519 half): new utils/signing.py + [sign]
extra. `soup adapters sign/verify --backend ed25519` (--key/--generate-key/
--public-key) and `soup attest emit --sign ed25519` + new `soup attest
verify`. Sigstore keyless stays infra-blocked (OIDC/Fulcio/Rekor network).

#186 namespace-pin gate wired into download_repo (anti-AI-Jacking; fail-open)
#187 adapters merge auto-detects each adapter's license
#190 license-override reason recorded to the audit log
#191 NamespacePinStore WAL + busy_timeout + cross-process lock
#192 merge refuses scan-FAIL inputs unless --allow-unscanned

+106 tests (12153 -> 12259); full suite 79.07% cov. cryptography added to
[dev] for CI. 5 review waves, all findings fixed.
2026-06-01 17:17:44 +05:00
Alpamys 0ee5f78986 fix(ci): ANSI-robust serve help test + restore 77% coverage gate (v0.71.1)
The v0.71.1 release commit (514761c) went red on CI for two reasons:

- test_flag_in_help asserted a raw "--record-thumbs" substring, but Rich
  splits an option name's dashes with ANSI codes under CI's FORCE_COLOR
  (it passes locally without color). Strip ANSI before the substring check.
- Coverage fell to 76.96% (< 77% gate): CI installs [dev], which has no
  FastAPI, so the new /v1/thumbs endpoint + record-thumbs startup block in
  serve.py are uncovered there. Restore the gate honestly (no lowering, no
  pragma) by adding 19 genuine no-FastAPI tests for previously-uncovered
  pure-CLI paths: lock show / lock check (no-drift / drift exit 3 / missing),
  env check (no-drift / missing / drift exit 3), env fix error branches,
  env lock null-byte output, and load_evidence_file (the
  `eval unlearning --evidence` loader).

CI-equivalent (no-fastapi) coverage: 76.96% -> 77.24%. Tests: 12134 -> 12153.
2026-06-01 15:03:37 +05:00
Alpamys 514761c89a feat(env,lock,serve,eval): v0.71.1 — quick wins + wiring (7 closures)
Closes #195 #210 #214 #224 #230 #233 #209.

- soup env fix: print-only install-plan renderer from soup-env.lock
  (uv-pip / requirements; non-pip entries surfaced as comments). (#209)
- soup lock write --env-lock: auto-derive --env-hash from soup-env.lock
  via new compute_env_hash (excludes created_at). (#224)
- soup serve --record-thumbs <db>: capture thumbs-up/down into the
  local-RL SQLite + POST /v1/thumbs (transformers backend). (#230)
- Judge-calibration persistence: JudgeCalibrationReport.to_dict +
  write/load_judge_calibration + judge_calibration registry kind; load
  re-validates the frozen dataclass with cwd/symlink containment. (#214)
- soup completions: introspect a base model's real LoRA target modules
  (config-only AutoConfig, local_files_only, never networks/raises). (#210)
- Bundled MUSE + WMDP unlearning eval fixtures; WMDP forget rows ship
  REDACTED (Soup never bundles verbatim hazardous content). (#195)
- build_dag.validate_build_source: cwd-containment + symlink rejection. (#233)

Review-fix hardening (consolidated python+code+security+tdd, 0 CRIT/0 HIGH):
load_judge_calibration containment + friendly missing-field ValueError;
serve thumbs success-print escape; env_fix --output Optional[str];
empty --env-hash auto-derives; render_install_plan PEP 440 docstring note.

Tests: 12071 -> 12134 (12044 passed, 90 skipped, 2 deselected).
2026-06-01 14:12:30 +05:00