Commit Graph

444 Commits

Author SHA1 Message Date
Alpamys a108c8183b fix(cli): ASCII-safe mcp help string (em-dash -> hyphen)
The v0.71.28 mcp command help used an em-dash, tripping the long-standing
test_cli_subprocess.py::TestEncoding::test_help_output_is_ascii_safe guard
(Windows encoding safety). Replace with an ASCII hyphen; main CI was red.
2026-07-04 19:08:27 +05:00
Alpamys 808c93587c docs: credit @CODING-DARSH for ORPO/SimPO/GRPO vocab expansion (#295) 2026-07-04 18:44:19 +05:00
Alpamys 9df0ee573e Merge branch 'main' of https://github.com/MakazhanAlpamys/Soup 2026-07-04 18:43:27 +05:00
Darsh cd6cbddbbf
fix: honor vocabulary expansion in ORPO, SIMPO, and GRPO trainers (#295) 2026-07-04 18:43:23 +05:00
Alpamys 96c04f9716 feat(mcp): stdio server + soup mcp serve command + [mcp] extra (v0.71.28 Part D) 2026-07-04 16:55:00 +05:00
Alpamys d4fe661f05 feat(mcp): plan-only mutating tools + --allow-mutating gate (v0.71.28 Part C) 2026-07-04 16:45:47 +05:00
Alpamys 370f562d22 feat(mcp): profile + diagnose/ship evidence tools (v0.71.28 Part B) 2026-07-04 16:40:47 +05:00
Alpamys 3d797afcae feat(mcp): tool registry + read-only handlers (v0.71.28 Part A) 2026-07-04 16:35:58 +05:00
Alpamys d034b23d19 docs: credit @CODING-DARSH for DPO/IPO/KTO/BCO vocab expansion (#293) 2026-07-04 15:46:59 +05:00
Darsh 1cc4bf48ad
fix: honor vocabulary expansion in DPO/IPO/KTO/BCO trainers (#293) 2026-07-04 15:46:08 +05:00
Alpamys 5c3a95312a docs(changelog): fold #291 vision/audio vocab fix into [0.71.27]
The #291 fix (vision/audio vocab expansion) landed on main below the
v0.71.27 version bump, so it ships in the v0.71.27 wheel — move its note
from [Unreleased] into [0.71.27] ### Fixed, next to its text-path sibling
#287, so the release notes accurately reflect what ships.
2026-07-04 14:09:06 +05:00
Alpamys 411ac504eb docs: credit @CODING-DARSH for vision/audio vocab expansion (#291) 2026-07-04 14:03:46 +05:00
Alpamys 937abb9e0d feat(data): soup data doctor + soup data lint — Fine-tune Doctor (v0.71.27)
Add `soup data doctor` and `soup data lint`, killing the top *silent*
fine-tune failures before a single training step: EOS-missing-from-labels
(the #1 "model never stops generating" bug), BOS duplication, no-system-role
templates, and preference-data length bias (the #1 silent DPO degradation) —
none of which any competitor (Unsloth/Axolotl/LlamaFactory) checks for.

- utils/data_doctor.py: 8-check chat-template compat report over a
  tokenizer + sampled rows, OK/MINOR/MAJOR taxonomy mirroring diagnose;
  --show-mask N renders per-token trained/masked colouring through the
  SAME masking dispatch (_build_row_labels) the report itself uses, so
  the two can never disagree about what's actually trained.
- utils/data_lint.py: preference-data linter (dpo/orpo/simpo/ipo/bco/kto)
  — length bias (Cohen's d), label imbalance, near-duplicates (MinHash),
  identical chosen==rejected pairs, prompt leakage.
- commands/data_doctor.py: Typer layer for both commands; strips C0
  control bytes before untrusted dataset content reaches the terminal.
- commands/diagnose.py: hardens the --evidence loader against a TOCTOU
  symlink swap (O_NOFOLLOW + fstat-on-open-fd), backporting the pattern
  soup ship shipped in v0.71.25 (closes v0.71.25 known-limitation (4)).

Live smoke against the real HuggingFaceTB/SmolLM2-135M-Instruct tokenizer
(Windows + RTX 3050) found and fixed two genuine bugs beyond the synthetic
fixtures: the EOS check needed to span-search the whole trained region
(not just the last token), and two apply_chat_template call sites needed
a broad except Exception for jinja2.exceptions.TemplateError.

+173 tests (14788 -> 15042). 5 sequential ECC reviews, every finding fixed.
2026-07-04 14:03:30 +05:00
Darsh 23da524554
fix: reuse shared vocabulary expansion in vision/audio SFT (#291) 2026-07-04 14:02:52 +05:00
Alpamys 1946e3fdde docs: credit @CODING-DARSH for judge-URL SSRF fix (#288) + SFT vocab expansion (#287) 2026-07-03 18:51:05 +05:00
Darsh d8519c5f80
fix(eval): prevent hostname prefix bypass in judge URL validation (#288)
* fix(eval): prevent hostname prefix bypass in judge URL validation

* style: ruff --fix import order + trailing newline (unblock CI lint)

---------

Co-authored-by: Alpamys <vpn.alpamys@gmail.com>
2026-07-03 18:49:48 +05:00
Darsh abf8fef1e0
fix(sft): apply configured vocabulary expansion (#287) 2026-07-03 18:38:09 +05:00
Alpamys 7476b75a73 docs(registry): recommend is_under_cwd over resolve()+relative_to() in resolve_dataset docstring
Last holdout of the recurring Path.resolve()+relative_to() containment class —
an advisory docstring (no behaviour change) that still recommended the
Windows-broken idiom.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:41:25 +05:00
Alpamys cba51e153b fix: close out the 3 deferred code-review items (EMA, hardware-fit, MoD)
These were documented as known limitations in the MEDIUM/LOW pass; now fixed.

1. reward_hack EMA smoothing window — smooth_signal("ema") folded only
   window[-1], so reward_hack_smoothing_window had no effect. Now a windowed
   EMA folds alpha over the whole retained window (oldest→newest) then the new
   sample, so a larger window incorporates more history; a 1-element window
   reduces to the old 2-tap form. Updated test_v07126 (0.3 → 0.275).

2. hardware_fit OOM gate wired into `soup train` — the analytical VRAM
   predictor was never called despite its docstring. Added
   _build_hardware_fit_input (SoupConfig → HardwareFitInput, best-effort;
   None when not statically predictable, e.g. batch_size="auto") and
   _hardware_fit_preflight, run after device detection. Refuses on predicted
   OOM (peak × 1.1 > available) unless the documented --allow-oom-attempt
   opt-out is passed; skips silently on CPU / unknown VRAM, and the flag is
   threaded through the --gpus re-exec.

3. MoD real token-dropping — mod_forward ran the full block on ALL tokens then
   masked (zero compute savings). Now the top-k tokens are gathered into a
   shorter sub-sequence, the block runs on ONLY those tokens (real saving),
   the gated result is scattered back, and unselected tokens pass through
   unchanged. Positional inputs (RoPE cos/sin, 4D-causal attention_mask,
   position_ids, cache_position) are gathered to the sub-sequence; any
   unsafe-to-gather case (positional forward args, KV cache, non-4D mask)
   falls back to the prior correct blend so attention can never be silently
   mis-computed. Validated on CPU (gather/scatter/passthrough/savings +
   fallback); the sub-sequence-attention numerics still warrant GPU validation
   at scale.

Adds tests/test_code_review_deferred.py (7 tests). ruff clean; full suite
14867 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:36:59 +05:00
Alpamys defd3151cf fix: remediate the MEDIUM/LOW code-review findings
- license_matrix: permissive ↔ weak-copyleft is now symmetric (MIT + LGPL no
  longer flagged incompatible regardless of order).
- formats.detect_format: check tool-calling before audio so an audio+tools row
  keeps its tool_calls instead of being classified audio.
- formats: reject null message content in the alpaca / sharegpt / vision text
  converters (routes the row to the drop path instead of literal None content).
  DPO only rejects an explicit null (chosen/rejected may be message lists).
- eval/custom.tool_call_args_subset: hallucinated args on a no-arg expected
  call now score 0.0 (was a dead `0.5 if ... else 0.5` ternary).
- monitoring/callback: SSE metric push uses `is not None` so a real 0.0 loss/lr
  is not reported as None.
- cans/schema.DeployTarget: reject Windows drive-absolute paths (C:\..., C:/...).
- commands/diagnose: reject a non-numeric evidence score with a clear
  BadParameter (was ValueError -> exit 1 with zero output).
- commands/generate: partial-save accumulated examples on a mid-run failure so
  paid API spend is not discarded.
- __init__.py: fix the byte-corrupted em dash in the package docstring.

Two MEDIUM/LOW items reverted to documented known limitations after they broke
existing behaviour locked by tests: (1) the reward_hack EMA smoother is a
recursive 2-tap by design — smoothing_window only affects `median`; (2)
`soup train`'s MoD compute-savings and the hardware_fit OOM preflight wiring
are architectural, GPU-validation work left as follow-ups.

Adds tests/test_code_review_medium_low.py (10 tests). ruff clean; full suite
14861 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 17:20:05 +05:00
Alpamys e150cba1e7 fix: remediate the recurring-pattern classes from the code review
Path.resolve()+relative_to() cwd-containment (breaks on Windows 8.3 short
names) -> utils.paths.is_under / is_under_cwd across all holdouts:
migrate/common.py, utils/ollama.py, commands/serve.py, commands/export.py,
commands/generate.py, data/loader.py, commands/data.py (×3), and
eval/checkpoint_intelligence.py (the pre-rmtree prune guard).

Unescaped Rich markup from external data: commands/_eval_v0550.py now
escapes the dataset-read exception message (×2). (adapters.py:315 already
escapes layer.name — that flagged instance was a false positive.)

Symlink-following writes -> project-safe helpers:
- utils/active_sampler.py: atomic_write_text (cwd-contained, symlink-reject).
- ui/app.py: tempfile.mkstemp (O_EXCL, unpredictable name) instead of a
  fixed shared-temp path a local attacker could pre-symlink.

typer.Exit-vs-SystemExit audit: the two named instances (bench.py, eval.py)
were fixed in the HIGH tier; cli.py's top-level run() already handles both
SystemExit and typer.Exit correctly, so no further sites needed changes.

Adds tests/test_code_review_recurring.py (7 regression tests). ruff clean;
full suite 14851 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 16:57:08 +05:00
Alpamys 58d280cb6f fix: remediate the HIGH-tier code-review findings
Training correctness (silent wrong results):
- ppo: refuse a randomly-initialised reward head when only reward_fn is set
  (trl 0.19.1 PPO can't use a reward_fn) instead of training against noise.
- edit_kernels (AlphaEdit): reject a non-finite key-norm (NaN <= 0.0 is False).
- preference_combine (ORPO): length-normalise log-probs so exp() doesn't
  underflow and kill the odds-ratio correction (+_read_lens caller wiring).
- ipo: anneal the beta schedule from ipo_tau, not the DPO default dpo_beta.
- distill: mask padding + prompt tokens in the default KL term (labels!=-100).
- block_expansion: freeze all-but-the-ACTUAL-added blocks (clamp over-request).
- formats (KTO): map a -1 label to False (bool(-1) was silently True).

Features that silently did nothing:
- sft: actually install the LongLoRA S² attention override (defensive).
- train --gpus re-exec: pass through --gate/--push-as/--trust-remote-code/
  --tracker/--diagnose-gate/--annex-xi/--repro-receipt/--profile/energy flags.
- eval gate-install hook: pass $GATE_SUITE to `soup eval against`, which now
  validates the locked suite as a precondition (block on missing/tampered).
- deploy_measure: fold the candidate list into the cache key.

Security:
- sglang: loopback-only CORS (was wildcard).
- fetch: lstat the ORIGINAL path (realpath resolved the symlink -> S_ISLNK
  never fired -> write followed the link).
- ui /api/data/inspect: is_under_cwd (commonpath) instead of str.startswith.
- registry lineage: unbounded cycle check (the depth-10 cap accepted a
  far-away cycle-closing edge).
- gguf calib: read from the O_NOFOLLOW fd (no close+reopen TOCTOU window).
- namespace_pin: flag ANY created_at drift (repo-recreation moves it forward).

Robustness / cross-platform:
- bench: re-raise typer.Exit (RuntimeError subclass) instead of masking it.
- eval auto: catch typer.Exit so a benchmark failure falls through.
- data split: reject negative --val/--test (negative slice inverted the split).
- trace parser: read utf-8-sig so a BOM'd first record isn't dropped.
- terraform plan: tolerate batch_size="auto" in the runtime estimate.
- rl_checkpoint: only rank-0 writes; atomic optimizer save.

Adds tests/test_code_review_high.py (29 regression tests). ruff clean;
full suite 14844 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 16:43:29 +05:00
Alpamys b587deea94 fix: remediate 6 CRITICAL code-review findings
1. RLVR verifiable rewards (grpo.py, ppo.py): pass verifiable_domain=
   to load_reward_fn so `reward_fn: verifiable` + `verifiable_domain: math`
   no longer crashes at setup() — the headline RLVR feature was 100% broken.
2. serve: default --host to 127.0.0.1 (was 0.0.0.0) and add an opt-in
   --tool-auth-token wired to _create_app so the code-exec tool endpoints
   are not exposed unauthenticated on all interfaces; warn on non-loopback
   bind without a token.
3. eval benchmark: reject ','/'=' in the adapter's base_model_name_or_path
   (and --model path) before lm-eval model_args interpolation, closing the
   trust_remote_code=True injection (mirrors the ship.py guard).
4. webhooks: run the private/link-local SSRF check for BOTH http and https
   (was nested in the http-only branch, so https://169.254.169.254 and
   10.x/192.168.x sailed through). Adds allow_private_hosts= so the trusted
   loop_stages LAN-deploy caller (which does its own loopback/LAN tightening)
   keeps working.
5. cloud/modal: emit the output_dir via the already-repr'd _LOCAL_OUTPUT
   variable in a generated f-string instead of raw interpolation, closing
   the stub code-injection hole (also quote-safe, unlike bare {output_dir!r}).
6. eval gate: wire cfg.training.eval_gate through all 14 trainers to
   SoupTrainerCallback, and load the suite/baseline + build a live generator
   in on_train_begin so --gate actually halts training. Warn honestly when
   forgetting/checkpoint/early-stop knobs are set (they are not yet enforced).

Adds tests/test_code_review_critical.py (27 end-to-end regression tests).
ruff clean; full suite 14815 passed / 120 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 15:56:48 +05:00
Alpamys 0db7404a6a test(train): strip ANSI + widen terminal in reward-hack help assertions (v0.71.26)
The two TestRewardHackMitigationCli help-text assertions grepped the raw
`soup train --help` output for the new flags. CI runners emit color codes
that split long option names into non-contiguous characters (same failure
mode documented in test_eval_gate.py for --gate), so the raw substring check
failed on every OS/Python combo while passing locally on a no-color terminal.

Fix follows the established repo pattern: render with COLUMNS=200 (no option
wrapping) and strip ANSI escapes before the membership check. Verified under
FORCE_COLOR=1: both assertions pass. Test-only change, no version bump.
2026-07-01 18:16:53 +05:00
Alpamys 2ffc3743ae feat(train): --reward-hack-detector / --reward-hack-halt CLI flags (v0.71.26)
Folds in the doc-vs-reality cleanup the plan flagged: the docs referenced
soup train --reward-hack-detector / --reward-hack-halt but they were config-only.
Add both as CLI passthroughs mirroring --reward-hack-mitigation (validate value,
set cfg.training field, accelerate re-exec passthrough) so the docs are true and
the reward-hack CLI is consistent. +4 tests (test_v07126: 180 -> 184; full suite
-> 14788).
2026-07-01 17:02:37 +05:00
Alpamys dca58c4107 docs(train): v0.71.26 release — closed-loop reward-hacking mitigation
Version 0.71.25 -> 0.71.26 (pyproject + __init__). CHANGELOG [0.71.26] entry
(feature + security). README What's New slot. docs/training.md mitigation
section + docs/commands.md flag. CONTRIBUTING + examples/README. Also folds in
the already-merged qwen2.5-coder-7b-sft recipe (#285) that rides this release.
2026-07-01 16:55:59 +05:00
Alpamys fa992bf381 fix(train): reward-hack mitigation review fixes (v0.71.26)
Fixes from 5 sequential ECC reviews (python/code/security/tdd/verification).

python-review (2 CRITICAL + HIGH/MED/LOW):
- signal/vote coherence: schema now requires the active detector in
  reward_hack_signals + rejects the inactive detector name (was silently
  dropping the primary signal from the vote).
- integral_clamp is its own field (was wrongly hard-wired to beta_ceil).
- task/backend gate runs before controller-config checks; EMA convention
  corrected; type hints; mutable-list default -> tuple + normalised compare.

code-review (4 HIGH + MED/LOW):
- _prune now trims _saved in sync with disk (rollback can't target a deleted
  checkpoint); bang-bang release_count resets after each relaxation (hysteretic
  descent); EMA formula uses standard convention; _escalate no longer burns a
  recovery attempt on a None target; max_recovery_attempts>=1 required with
  rollback; _action_history capped; on_step_end logs errors once; loud warning
  when the mitigation callback can't attach (was a silent safety-off).

security-review (HIGH + MED):
- restore_checkpoint / save_checkpoint refuse a SYMLINKED optimizer.pt
  (torch.load weights_only=False was an RCE via attacker-placed symlink);
  bool-before-int/float guards on all new numeric fields; reward_hack_signals
  max_length=4; empty-signals guard in the callback.

tdd-review: +13 coverage tests (dead-band hold, shim verbatim-on-error,
conservative boundary, read-only-beta dual-write, escalation postconditions,
both-restore, PID D exact, log cap + concurrency, no-top-level-import, fuzz
field-validity).

Test count 152 -> 180 (+2 POSIX-only symlink skips).
2026-07-01 16:38:00 +05:00
Alpamys eb2edb1a51 feat(train): anti-gaming hardening + reward shaping + give-up explainer (v0.71.26 Part D)
Stage 3 of closed-loop reward-hacking mitigation: the controller itself must
not be gameable.

- schema: 6 Stage-3 tunables (signal_smoothing, smoothing_window,
  conservative_on_disagreement, reward_shaping, shaping_kind, shaping_strength)
  + bool guards; reward_shaping requires a control mode + strength>0.
- reward_hack_control: combine_conservative (disagreement -> MAX, stay cautious);
  detect_reward_distribution_drift (bimodal-collapse heuristic); shape_reward_fn
  + apply_reward_shaping (bounded length/repetition/sentinel penalty over the
  wrap_reward_funcs seam, inner called once, verbatim on shim error);
  explain_giveup (plain-English, mirrors why.py). Callback wires per-signal
  smoothing + conservative vote + opt-in drift guard (keep KL high) + logs the
  give-up explanation on early-stop.
- peft_wiring: thread smoothing/conservative params from tcfg.
- grpo.py / ppo.py: apply_reward_shaping BEFORE the buffer capture.

Includes an adversarial-fuzz suite (sawtooth/step/noisy/flip/out-of-range
traces): bounded output, no unbounded jump, no flap, anti-windup holds.

+35 tests (test_v07126: 117 -> 152).
2026-07-01 15:33:46 +05:00
Alpamys 21557b43c1 feat(train): PID-Lagrangian controller + rollback escalation ladder (v0.71.26 Part C)
Stage 2 of closed-loop reward-hacking mitigation: principled control + safety net.

- schema: 7 Stage-2 tunables (pid_kp/ki/kd, signal_target, rollback,
  rollback_patience, max_recovery_attempts) + bool guard; cross-validators —
  PID/rollback tunables require pid_lagrangian mode; reward_hack_rollback
  requires rl_checkpoint_save_every_steps (a cadence to roll back to).
- reward_hack_control: PIDLagrangianPolicy + pid_step (Stooke et al. — constraint
  'signal <= target' whose multiplier β is a PID law; integral anti-windup clamp,
  output clamp [floor,ceil], never crosses 0). Callback pid_lagrangian path drives
  β via PID + escalation ladder: raise -> rollback last-good checkpoint (after
  rollback_patience HACK steps) -> early-stop (after max_recovery_attempts),
  subsuming halt_on_hack.
- rl_checkpoint: RLCheckpointCallback.restore_checkpoint (reload PEFT adapter via
  set_peft_model_state_dict + optimizer.load_state_dict; best-effort, never raises).
- peft_wiring: build PIDLagrangianPolicy from tcfg; build the RL-checkpoint cb
  FIRST and hand its reference to the mitigation callback for rollback.

+26 tests (test_v07126: 91 -> 117; real-peft adapter round-trip included).
2026-07-01 15:18:07 +05:00
Alpamys 925e600702 feat(train): reversible bang-bang KL controller + hysteresis (v0.71.26 Part B)
Stage 1 of closed-loop reward-hacking mitigation: the first closed loop.

- schema: 8 Stage-1 tunables (beta_floor/ceil, trip/release band, dwell_steps,
  release_patience, kl_gain, signals allowlist) + bounds; extended
  _validate_reward_hack_compat (floor<ceil, release<trip, signal allowlist,
  control-mode XOR ref_model_ema_alpha, footgun-reject tunables while off).
- reward_hack_control: MitigationAction + BangBangPolicy + bang_bang_step (pure
  hysteresis: dwell before trip, release_patience before relax, beta geometric
  x/div kl_gain clamped [floor,ceil], never crosses 0; multi-signal vote).
  Callback kl_control path mutates BOTH trainer.beta and trainer.args.beta
  (GRPO stock + variant) / trainer.args.kl_coef (PPO); seeds beta from trainer;
  length-trend signal.
- peft_wiring: build BangBangPolicy from tcfg for kl_control.
- ppo.py: RLSignalBuffer parity + attach_rl_callbacks(task=ppo) (BETA — GRPO
  gets the on-GPU proof; PPO kl_coef mutation is unit-tested).
- train.py: --reward-hack-mitigation off|log_only|kl_control|pid_lagrangian
  flag + validation + accelerate re-exec passthrough.

+33 tests (test_v07126: 58 -> 91).
2026-07-01 15:05:32 +05:00
Alpamys bcb08ae0ab feat(train): reward-hack mitigation instrumentation + log_only telemetry (v0.71.26 Part A)
Stage 0 of closed-loop reward-hacking auto-mitigation: observe only, no
control action.

- schema: reward_hack_mitigation Literal[off/log_only/kl_control/pid_lagrangian]
  gated to grpo/ppo + non-mlx + requires reward_hack_detector; YAML-1.1 bool
  DWIM (off->"off", on/yes rejected with a quote hint).
- utils/reward_hack_control.py (no top-level torch): MitigationLogWriter
  (mirrors TraceLogWriter: thread-lock, rotation, redaction, cwd containment,
  symlink-reject), ControllerState (frozen), combine_signals, smooth_signal,
  telemetry helpers, RewardHackMitigationCallback (log_only path provably never
  mutates beta).
- reward_hacking.RewardHackCallback: last_drop_pct() accessor for the controller.
- peft_wiring: mitigation callback subsumes the plain detector when a mode is set.
- examples/reward_hacking/rewards.py: synthetic length-hack + sentinel proxies
  decoupled from a held-out true_score (the GPU-experiment fixture).
- trace_logger: public redact_value alias (DRY reuse).

+58 tests (tests/test_v07126.py).
2026-07-01 14:47:40 +05:00
Alpamys 4f0dff79cb docs: credit @Deadpool2000 for qwen2.5-coder-7b-sft recipe (#285) 2026-06-28 19:23:25 +05:00
Salil M 14de9ec9d5
feat(recipes): add ready-made SFT recipe for Qwen2.5-Coder-7B-Instruct (#285)
* feat(recipes): add qwen2.5-coder-7b-sft recipe to catalog

* test(recipes): update catalog size assertion to 134 in recipe tests

* test(recipes): update version catalog count to 134 in v0.71.24 tests

* docs(contrib): bump recipe count in project structure overview

* docs(commands): update recipe list count reference to 134

* docs(serve): update Web UI recipe count references in serving docs
2026-06-28 19:22:29 +05:00
Alpamys 6cb2e0201e docs: index soup ship in docs/README + CONTRIBUTING utils list 2026-06-28 00:02:23 +05:00
Alpamys 6cb1abab8f feat(eval): soup ship — SHIP / DON'T-SHIP verdict (v0.71.25)
Add `soup ship`, a binary SHIP / DON'T-SHIP verdict after fine-tuning: it
SHIPs only when (leg 1) the task metric strictly improved AND (leg 2) no
general benchmark regressed past a forgetting threshold (default 0.05
absolute points) — otherwise DON'T SHIP, even if the task metric looks
great. The moat is leg 2 (catastrophic-forgetting gate) fused with the
task win into one decision. Exit: 0=SHIP, 2=DON'T SHIP, 1=runtime error.

- utils/ship_verdict.py: pure engine (no top-level torch) — frozen
  TaskWin/BenchmarkDelta/ShipVerdict + decide_ship (single source of
  truth for the threshold) + compute_benchmark_deltas + render/serialize.
- commands/ship.py: Typer command; --evidence offline path + live
  metric/judge leg-1 + mini(default)/lm-eval leg-2; --baseline/--output.
- Reuses run_eval / JudgeEvaluator / ForgettingDetector / resolve_baseline
  / _run_lm_eval / live_eval.make_generator.
- Hardening: --evidence O_NOFOLLOW + size cap; --task-eval cwd-contained;
  --judge-model urlparse SSRF guard; lm-eval model_args injection guard;
  --general-suite bounded.

Schema (ShipConfig) deferred — v1 is CLI-only. Pairwise judge win-rate is
a planned fast-follow. +79 tests (14514 -> 14593).
2026-06-27 23:37:35 +05:00
Alpamys e0579383e5 docs: surface contributors on README front page 2026-06-22 16:12:27 +05:00
Alpamys bd9ef0e578 docs: credit @Akshaya-reddy18 for friendly error mappings (#282) 2026-06-22 15:56:44 +05:00
Kondamwar Akshaya Shrikant e4065d8741
Fix friendly error mappings 272 (#282)
* Improve friendly error mappings and tests

* Improve friendly error mappings and tests
2026-06-22 15:54:55 +05:00
Alpamys 418f86390a feat(recipes): 2026 model-family expansion — 17 SFT recipes, catalog 116→133 (v0.71.24)
Add ready-made SFT recipes for the open-weight models released Feb–Jun 2026,
each base repo-ID verified to resolve on Hugging Face:
- Qwen3.5 0.8B/2B/4B/9B/27B + MoE 35B-A3B/122B-A10B/397B-A17B (Apache-2.0)
- Qwen3.6 27B + 35B-A3B (Apache-2.0)
- DeepSeek-V4 Flash/Pro (MIT), GLM-5.1 (MIT)
- Kimi-K2.5/K2.6 (Modified MIT), MiniMax-M3 (MiniMax Community License)
- Mistral-Large-3 (Apache-2.0, 675B/41B-active multimodal MoE)

Fix stale glm-5-sft repo-ID THUDM/glm-5 -> zai-org/GLM-5 (org migration).
+220 tests (tests/test_v07124.py). Catalog count 116 -> 133.
2026-06-21 13:00:47 +05:00
Alpamys da81b61798 docs: credit @shatakshi-1404 for warmup.py tests (#274)
Add CONTRIBUTORS.md entry + CHANGELOG [Unreleased] note for the
first merged PR from @shatakshi-1404 (unit tests for the warmup
auto-steps helper).
2026-06-20 21:38:56 +05:00
Shatakshi Prasad 1d1d12c3a7
test: add unit tests for warmup.py helpers (#274)
* test: add unit tests for warmup.py helpers

* fix: ruff lint fixes (import order + trailing newline)

* fix: add trailing newline

* style: ruff --fix
2026-06-20 21:38:05 +05:00
Alpamys e91da922c3 docs: correct recipe count to 116 and refresh stale TTS/BitNet blurbs
Recipe-count drift: the CLI help (commands.md) and Web UI pages
(serving-and-export.md) claimed 43 ready-made recipes and the catalog
docstring said ~30, while RECIPES actually holds 116 (test_recipes already
asserts len == 116). Aligned every user-facing count to 116.

Also refreshed 6 recipe descriptions still tagged "schema-only stub
(live in v0.52.1)": the TTS task (#131) and BitNet 1.58 SFT (#134) went
live in v0.71.20, so orpheus/sesame/llasa/spark/oute TTS and the Falcon-E
BitNet recipe now read "live (v0.71.20)".

Doc-drift only — description strings + one comment; no functional change,
no version bump.
2026-06-20 12:30:00 +05:00
Alpamys fcf4b33394 feat(train): native Spectrum targeted training — soup spectrum scan + training.unfrozen_parameters (v0.71.23)
Closes #266. `soup spectrum scan` streams safetensors per-tensor (no model
load, CPU-friendly) and computes a singular-value SNR per weight matrix
(Marchenko-Pastur, arXiv:2406.06623), emitting a ready-to-paste
training.unfrozen_parameters patch. The SFT trainer freezes all params then
unfreezes the matched set (full FT, LoRA off).

- utils/spectrum_scan.py: pure-numpy transpose-invariant SNR kernel +
  per-tensor safetensors streaming (2^31 SVD cap, symlink skip) + cache
  (~/.soup/spectrum, SOUP_SPECTRUM_CACHE_DIR containment) + hardened
  hubs.snapshot_download.
- commands/spectrum.py: soup spectrum scan (SNR table + YAML patch).
- schema: training.unfrozen_parameters (caps/NUL/invalid-regex/ReDoS reject)
  + gates (sft/transformers/text/quantization=none; mutually exclusive with
  LoRA features / freeze_layers / freeze_ratio / train_router_only /
  expand_layers).
- trainer/sft.py: full-FT branch via apply_unfrozen_parameters +
  enable_input_require_grads (fixes grad-checkpointing through frozen
  embeddings).

Existing spectrum trainer-plugin untouched (back-compat); LISA -> #267.
Live-validated on Windows + RTX 3050: CPU scan of SmolLM2-135M + top-25%
unfrozen full-FT train (loss 3.455 -> 0.719).

Note: after the version bump the editable install metadata was stale
(0.71.17); pip install -e . --force-reinstall --no-deps re-synced it so
test_cli_subprocess::test_version passes.

+94 tests in tests/test_v07123.py (14184 -> 14278).
2026-06-12 17:40:46 +05:00
Alpamys 0c82c443d6 chore: simplify .claude gitignore to a single entry 2026-06-10 23:26:45 +05:00
Alpamys 8b7d63f944 docs: clarify Orpheus live-codec TTS is live in training.md (v0.71.22)
The live-codec block claimed the entire data.format=audio path was "not
validated on the maintainer's box" and only surfaced a RuntimeError. v0.71.22
made the Orpheus SNAC encode live + validated; note that while the other four
families stay dependency-gated. Docs-only, no version bump.
2026-06-10 22:02:50 +05:00
Alpamys ccd5c80e4d feat(perf): MiniLLM/MoLE KV-cache + deploy-measure live factories + live-codec TTS (v0.71.22)
#263 MiniLLM on-policy KV-cache (PEFT-unwrap probe activates the cache for LoRA
students; per-step single-token forward), #262 serve --mole per-adapter KV cache
(fresh per generate, no cross-request leak, byte-identical to no-cache), #143
deploy-autopilot live generator factories (baseline scored once + up-front
candidate validation; injected seams retained), #265-partial live-codec TTS
(soundfile.info pre-probe + O_NOFOLLOW; SNAC Orpheus encode validated).

Review: 1 HIGH + 5 MEDIUM + ~10 LOW fixed across 2 review waves + verification +
step-6 live smoke (Windows + RTX 3050). Tests 14084 -> 14184 (+100 in
tests/test_v07122.py; 293 files). Full suite 14067 passed / 117 skipped, exit 0.
2026-06-10 21:30:01 +05:00
Alpamys 9fd356b8ae docs: bump CONTRIBUTING test counts to 292 files / 14084 tests (v0.71.21) 2026-06-10 16:50:31 +05:00
Alpamys ed5fc3a8b3 feat(precision,rollout): live fp8/nvfp4 + vLLM sleep + openenv rollout + apple-adapter + delinearize-llama4 (v0.71.21)
Closes #141, #124, #125, #228, #97.

- #141: apply_fp8_attention (torchao float8 on attention projections, Hopper
  gate) + apply_nvfp4 (NVFP4Config, Blackwell gate); partial-conversion honesty;
  wired into the v0.28 speed/memory pipeline with yellow-advisory degrade.
- #124: vllm_sleep_mode live - create_vllm_engine(sleep_mode=True) +
  vllm_sleep_cycle ctx (wake in finally) + TRL GRPOConfig hook probe.
- #125: openenv rollout fully live via training.rollout_func module:fn
  resolver; rows replace the prompt dataset; art/ruler/nemo_gym honest dep
  gates + _EXTERNAL_ROLLOUT_RUNNERS seam. Real GRPO train on SmolLM2-135M.
- #228: convert_apple_adapter live - PEFT LoRA <-> mlx-lm (both matrices
  transpose, bf16 upcast, adapters.safetensors + num_layers, npz legacy read,
  np.ascontiguousarray fix for safetensors non-contiguous mangling);
  *-to-apple upstream-gated exit 3.
- #97: delinearize-llama4 live - [E*din,dout] -> [E,din,dout] per shard,
  config.json expert-count probe + --num-experts, sidecar copy, atomic writes.

Review waves: 3 HIGH + ~8 MEDIUM + ~12 LOW fixed.
Tests: 13874 -> 14084 (+210 in tests/test_v07121.py).
Full suite: 13967 passed, 117 skipped. ruff clean.
2026-06-10 16:29:52 +05:00
Alpamys a4dfbb308c feat(trainer): live TTS / BitNet / MoE-expert-quant trainers (v0.71.20)
Lift three v0.52.0 schema-only NotImplementedError stubs to real code.

- #131 TTS: TTSTrainerWrapper(SFTTrainerWrapper) — TTS fine-tune = next-token
  CE over [text][audio-codec-token] chat; per-family emotion templating
  (Orpheus/Oute) + codec special-token registration. Pre-encoded chat path
  live-validated on SmolLM2-135M-Instruct; live-codec (data.format=audio)
  hardware-gated per family.
- #134 BitNet: BitNetTrainerWrapper gated on onebitllms; export --format
  bitnet|tq1_0 runs real llama.cpp TQ1_0 ternary GGUF export.
- #136 MoE: apply_moe_expert_quant swaps fused-MoE experts to bnb Linear4bit/
  Linear8bitLt (pre-LoRA); train_router_only freezes experts (post-LoRA).
  Live-validated on RTX 3050 (dequant err 0.0155).

Review fixes: H1 explicit Params4bit/Int8Params weight-carry; H2 quant
pre-LoRA / freeze post-LoRA + skip PEFT-wrapped modules; M4 device-aware
placement.

Tests 13807 -> 13874 (+69 in test_v07120.py, -2 lifted stubs in test_v0520.py).
2026-06-10 12:37:14 +05:00
Alpamys 853b348898 docs: refresh quant-menu modality + multipack sharding notes (v0.71.19)
- performance-and-quantization.md: the Quant Menu multi-trainer note said
  "vision / audio modality is still SFT-only inline-BNB (wiring tracked as a
  follow-up)" — stale after v0.71.19 #81 dropped the modality gate. Now states
  vision/audio thread the unified loader (full gptq/awq/hqq/aqlm/eetq/mxfp4/fp8
  menu), with the upstream class+kernel caveat.
- peft-and-efficiency.md: added the v0.71.19 #80 multi-GPU sharding paragraph to
  the Multipack section (accelerator.prepare + BatchSamplerShard under
  num_processes>1; identical bin seed across ranks; single-GPU unchanged).

Docs-only — no version bump / tag (the v0.71.19 code already shipped at f51331d).
2026-06-09 13:02:30 +05:00