mirror of https://github.com/razor-ai/soup.git
3 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
f316d334bc |
feat(rl): live GRPO/RL callbacks — reward-hack, echo-trap, RL ckpt, ULD, MiniLLM, iterative-DPO (v0.71.11)
Lifts the v0.70.0 schema-only build_*_callback / build_uld_projection / run_iterative_dpo stubs. Validated end-to-end on SmolLM2-135M. Closes #235, #236, #237, #238, #239, #240, #159, #160 - #235 RewardHackCallback: info_rm cluster-sep / rm_ensemble divergence, OK/WARN/HACK, halt on HACK. Shared thread-safe RLSignalBuffer captures per-step rewards by wrapping the reward fns (no TRL monkeypatching). - #236 ULD: Wasserstein-1 / top-k aligned distill loss in DistillTrainer. - #237 MiniLLM: teacher-mixed length-normalised reverse-KL + pretrain anchor. - #238 RLCheckpointCallback: adapter + optimizer.pt + manifest + keep_last prune. - #239 run_iterative_dpo: sample -> RM-score -> build-pairs -> DPO-train per round. - #240 EchoTrapCallback: n-gram repetition OK/WARN/TRAP, halt on TRAP. - #159 one-shot WARNING when a GRPO variant compute_loss falls back to super(). - #160 in-place ref-model EMA (no state_dict round-trip) + 0-overlap warning. Tests 13142 -> 13203 (+62 in tests/test_v07111.py). |
|
|
|
06d8ea7cc5 |
chore: migrate to src-layout
Move soup_cli/ -> src/soup_cli/ (history preserved via git mv). src-layout
forces the test suite to import the installed package instead of the
repo-root source tree, surfacing packaging bugs that flat-layout masks —
e.g. the v0.53.8 double-shipped-fixtures regression, invisible because
`pytest tests/` imports ./soup_cli directly and never from the wheel.
- pyproject: packages = ["src/soup_cli"]; artifacts globs -> src/soup_cli/...
The import name is unchanged, so the `soup` entry point, --cov=soup_cli,
and report_to/module-path strings stay `soup_cli`.
- CI / ownership: ruff lint path (ci.yml), recipe-validation `paths:` filters,
CODEOWNERS patterns, and the PR-template checklist all repointed to
src/soup_cli/.
- tests: source-grep regression tests that read package files by repo-relative
path repointed to src/soup_cli/ (64 files; 170 path literals). Lines pushed
over 100 chars by the prefix were wrapped to keep ruff E501 clean. Module
references (`import soup_cli`, `-m soup_cli`, mock.patch("soup_cli.x")) and
the `--cov=soup_cli` coverage target are deliberately unchanged.
- docs: AGENTS.md + CONTRIBUTING.md structure tree and lint commands.
Verified locally: ruff clean (src/soup_cli + tests); `import soup_cli`
resolves to src/soup_cli/__init__.py; built wheel ships
soup_cli/data/_fixtures/*.jsonl (10 files, no duplicates, no src/ prefix);
3174 tests across every touched test file pass. Packaging-only — no version bump.
|
|
|
|
74edac95d1 |
feat(v0.70.0): Loop Hardening — reward-hacking + ULD + MiniLLM + RL ckpt + iterative DPO + echo-trap
Six-part schema-only release shipping the axis-3 + axis-13 training-loop hardening. Every live trainer-callback / math kernel is deferred to v0.70.1 per the project's established stub-then-live cadence (matches v0.50.0 / v0.62.0 / v0.69.0). Part A — Reward-hacking detector (soup_cli/utils/reward_hacking.py): InfoRM Cluster-Separation Index (Wang et al. 2024 arXiv:2402.09345) + RM-ensemble pairwise variance + OK/WARN/HACK taxonomy at 0.10/0.30 thresholds. TrainingConfig.reward_hack_detector / reward_hack_halt; SoupConfig task-gate (grpo/ppo only, mlx rejected, halt requires detector). Part B — Cross-tokenizer ULD (soup_cli/utils/uld.py): Universal Logit Distillation (Boizard et al. 2024 arXiv:2402.12030). wasserstein + topk_align allowlist; ULDConfig frozen with topk/strategy cross-validators; vocab-size cap 262144. Schema-gated to task='distill'. Part C — MiniLLM reverse-KL on-policy distillation (soup_cli/utils/minillm.py): Gu et al. 2024 (arXiv:2306.08543) — bundles teacher-mixed sampling + length-norm + pretrain-loss anchor stability tricks. Anchor weight↔path mutual-requirement cross-validators reject silent no-op combos. Part D — Mid-epoch RL checkpoint (soup_cli/utils/rl_checkpoint.py): Optimizer-state serialization TorchTune explicitly punts. RLCheckpointConfig + RLCheckpointState frozen + JSON-serialisable manifest. RL-task gate. Part E — Iterative DPO loop driver (soup_cli/utils/iterative_dpo.py + commands/iterative_dpo.py): sample → RM-score → re-pair → retrain over N rounds. IterativeDPOPlan with consecutive-round_index invariant. New `soup iterative-dpo` CLI; --plan-only live, runner deferred. Part F — RAGEN echo-trap detector (soup_cli/utils/echo_trap.py): Zhu et al. 2025 (arXiv:2504.14437) — n-gram trajectory-repetition kernels with DoS caps (max 32 ngram_n, 1M tokens, 100k trajectories). OK/WARN/TRAP at 0.30/0.60. Composes with v0.53.11 #127 GRPOStabilityCallback. Cross-cutting hardening: - 6 new util modules + 1 new top-level CLI + 11 new TrainingConfig fields + 6 new SoupConfig cross-validators + 3 new field validators - Closed allowlists (frozenset) + MappingProxyType registries everywhere - Frozen dataclasses with post-init validation on every public record - Bool-as-int rejection on every numeric (matches v0.30.0 / v0.41.0 policy) - math.isfinite NaN/Inf rejection on every float - Null-byte rejection + per-field length caps on every string - No top-level torch imports (4 source-grep regression tests) - Deferred-live stubs validate inputs FIRST then raise NotImplementedError with explicit v0.70.1 marker - CLI exit codes split: 2 = validation rejection, 3 = deferred-live Test count: 11487 → 11824 (+337 net). 12-invariant self-review against the full project checklist (closed allowlists, frozen dataclasses, MappingProxyType, bool-as-int rejection, finite check, null-byte, length caps, no top-level torch, TypeError/ValueError split, deferred-live, tuples-not-lists, CLI exit codes) all green across all 6 Parts. Manual CPU smokes (Step 6): every CLI happy + failure path exercised — `soup iterative-dpo --plan-only` 3-round plan rendered end-to-end with per-round artifacts; 5 happy-path YAML loads + 5 failure-mode rejections across reward_hack / uld / minillm / rl_checkpoint / echo_trap. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |