mirror of https://github.com/razor-ai/soup.git
3 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
cfdabf2b3b |
feat(v0.53.11): GRPO Plus finish + preference live
Closes v0.50.1 (#123, #126, #127), v0.49.1 (#119), v0.40.1 (#68). #123 — live math kernels for 6 GRPO variants (gspo/dapo/dr_grpo/bnpo/ two_sided/rft) + `_GRPOTrainerVariant` HF Trainer subclass via `make_grpo_trainer_variant` factory. Variant compute_loss reads kernel inputs FIRST (no double-forward); falls back to super() only on missing attrs. Case-insensitive variant normalisation before lru_cache. #126 — PRMTrainerWrapper + `_PRMTrainer` HF Trainer subclass with real compute_loss (gather hidden states at step_positions -> reward_head -> MSE via compute_prm_loss). Dataset wrapped in datasets.Dataset.from_list for HF Trainer compatibility. Bool-before-isinstance guard on batch_size. #127 — GRPOStabilityCallback inherits transformers.TrainerCallback (lazy), live EMA ref-model update in on_step_end with strict=True + fallback-to-strict=False-with-WARNING on key mismatch (silent corruption defence). math.isfinite guard on alpha. #119 — LongLoRA forward override via LongLoRAForwardOverride context manager with idempotent install (_soup_longlora_patched marker prevents re-entry double-wrap), 256-char class name cap on regex match, restore on __exit__ AND on exception. #68 — true per-batch weighted-sum preference combine reading policy/ref logps from TRL inputs + each compute_*_term kernel + combine_losses. Explicit None checks on trainer attrs (no `or` on possibly-tensor), DEBUG log on per-term skip. Review fixes from 4 agents (python/code/security/tdd): 10 HIGH + 8 MEDIUM + 7 LOW — see CLAUDE.md v0.53.11 entry for the full list. Test count: 8330 -> 8400 (+75 in test_v05311.py: 54 initial + 21 review-fix coverage gaps). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
89b06efb8c |
feat(longctx): v0.53.4 — Long Context + Architecture
Six closes lifting the v0.49.0 LongLoRA hardening + v0.41.0 LLaMA Pro deferred stubs, plus a UX upgrade to the CUDA-OOM friendly message: - #11 utils/errors.py: OOM hint now names --batch-size / --grad-accum - #122 flash_attn.is_flash_attn_v3_available() + LongLoRA+FA3 schema reject - #120 LongLoRA arch allowlist expansion (Mistral / Qwen / Phi); Mixtral intentionally excluded (regex matches the bare 'mistral' token only) - #121 apply_long_context_config auto-detects 'llama3' when caller passes rope_scaling_type=None and the model config carries a Llama 3.1 rope_scaling block - #83 block_expansion.expand_model_blocks LIVE (deepcopy last-N blocks, zero-init residual projections, append, bump num_hidden_layers) + apply_llama_pro_freeze + shared apply_block_expansion_if_configured helper wired into SFT + Pretrain (mirrors v0.40.6 peft_wiring centralisation policy so SFT and Pretrain stay in lock-step) - #74 HF push surface QA — test plan recorded in tests/qa/v053_qa.md; live execution against a private HF repo deferred to a credentialed contributor Review pipeline (python / code / security / tdd agents) ran; every CRITICAL -> LOW finding addressed: - bool-first guards in _check_model_name (defends against int subclass) - is_supported_longlora_arch defensive non-string surface (returns False, never raises) matching v0.53.3 is_known_vlm_base policy - _truncate_for_message(value, limit=64) bounds the base echo in LongLoRA error messages (security MEDIUM, mirrors v0.34.0 crash.py) - null-byte + non-string TypeError guards on validate_longlora_compat task / backend params (matches v0.50.0 validate_long_context_grpo_compat) - _get_layers_module uses explicit `is None` not falsy shortcut (defends against nn.Module.__bool__ overrides on subclasses) - _zero_init_block_residual returns bool + warnings.warn when neither standard projection matches the cloned block (non-Llama-shaped arches still train but lose the LLaMA Pro identity-init guarantee) Test count: 7879 -> 7935 (+56 net; +49 in new tests/test_v0534.py). Lint clean. CPU smoke verified on a real transformers.LlamaForCausalLM: 4 -> 6 layers, down_proj + o_proj actually zeroed on PyTorch tensors, old blocks frozen + new blocks trainable, forward pass finite. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
161c81ee7f |
feat(long-context): v0.49.0 — YaRN, Dynamic NTK, LongLoRA S², Llama 3.1 NTK
- Part A: YaRN RoPE scaling — math kernels (yarn_find_correction_dim /
yarn_find_correction_range / yarn_linear_ramp_mask / yarn_get_mscale) +
4 yarn_* schema fields + cross-validator rejecting yarn fields outside
rope_scaling_type=yarn. Pure-Python implementations of the upstream YaRN
paper §3.4/§3.5 with bool/NaN/Inf rejection on every numeric input.
- Part B: Dynamic NTK — existing path verified, explicit test coverage.
- Part C: LongLoRA S² shifted-sparse attention (schema-only) — new
soup_cli/utils/longlora.py with is_llama_model (word-boundary regex
mirroring v0.39.0 is_gemma4_model policy) + validate_longlora_compat.
TrainingConfig.use_longlora + SoupConfig._validate_longlora_compat.
Live LlamaAttention.forward override deferred to v0.49.1 (stub-then-live
pattern, mirrors v0.27.0 MII / v0.37.0 multipack).
- Part D: Llama 3.1 NTK-aware (full impl) — scale_inv_freq_llama3
smooth-transition kernel + detect_llama3_rope_in_config HF-config probe
+ "llama3" added to rope_scaling_type Literal. LLAMA3_DEFAULT_*
constants per Unsloth models/llama.py:1853.
Public-boundary input validation on get_rope_scaling_config (bool/NaN/Inf
rejection on target_length / original_length / yarn_factor) per security
review — prevents direct callers from emitting {factor: NaN} into HF
model configs when bypassing the Pydantic schema.
Reviews: code-reviewer (2 HIGH + 1 MEDIUM + 1 LOW), security-reviewer
(1 MEDIUM + 2 LOW), python-reviewer (4 findings), tdd-guide (8 coverage
gaps) — all findings fixed. verification-loop done as manual equivalent
(version + --help + happy/failure YAML smoke).
+80 net new tests (6410 → 6490). Full suite green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|