mirror of https://github.com/razor-ai/soup.git
2 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
ed5fc3a8b3 |
feat(precision,rollout): live fp8/nvfp4 + vLLM sleep + openenv rollout + apple-adapter + delinearize-llama4 (v0.71.21)
Closes #141, #124, #125, #228, #97. - #141: apply_fp8_attention (torchao float8 on attention projections, Hopper gate) + apply_nvfp4 (NVFP4Config, Blackwell gate); partial-conversion honesty; wired into the v0.28 speed/memory pipeline with yellow-advisory degrade. - #124: vllm_sleep_mode live - create_vllm_engine(sleep_mode=True) + vllm_sleep_cycle ctx (wake in finally) + TRL GRPOConfig hook probe. - #125: openenv rollout fully live via training.rollout_func module:fn resolver; rows replace the prompt dataset; art/ruler/nemo_gym honest dep gates + _EXTERNAL_ROLLOUT_RUNNERS seam. Real GRPO train on SmolLM2-135M. - #228: convert_apple_adapter live - PEFT LoRA <-> mlx-lm (both matrices transpose, bf16 upcast, adapters.safetensors + num_layers, npz legacy read, np.ascontiguousarray fix for safetensors non-contiguous mangling); *-to-apple upstream-gated exit 3. - #97: delinearize-llama4 live - [E*din,dout] -> [E,din,dout] per shard, config.json expert-count probe + --num-experts, sidecar copy, atomic writes. Review waves: 3 HIGH + ~8 MEDIUM + ~12 LOW fixed. Tests: 13874 -> 14084 (+210 in tests/test_v07121.py). Full suite: 13967 passed, 117 skipped. ruff clean. |
|
|
|
33c60b4c1f |
feat(grpo): v0.50.0 — GRPO Plus (unsloth + axolotl RL parity)
22 features across 5 internal Parts shipped as schema-only — closed allowlists, Pydantic validators, NotImplementedError stubs for live wiring deferred to v0.50.1 (mirrors v0.27.0 MII / v0.37.0 multipack / v0.41.0 LLaMA Pro / v0.45.0 plugins / v0.48.0 curriculum / v0.49.0 LongLoRA stub-then-live pattern). Part A — 7 GRPO objective variants (gspo / dapo / dr_grpo / bnpo / two_sided / rft / standard) with `validate_grpo_variant` + frozen `GRPOVariantSpec` metadata + `MappingProxyType`-wrapped registry. `validate_grpo_delta` is bool-first / math.isfinite / (0, 1] bounded. `apply_variant_loss` raises NotImplementedError with v0.50.1 marker for the 6 new variants and is a no-op for standard. Part B — long_context_grpo + vllm_sleep_mode schema gates with compat validators (null-byte rejection on task + backend, bool guard on use_ring_attention). vllm_sleep_mode requires task='grpo' AND a transformers/unsloth backend (code-review HIGH fix — sleep is a between-rollouts feature). Part C — 4 multi-turn rollout backends (art / ruler / nemo_gym / openenv) with closed allowlist + per-entry required_package mapping. Part D — 7 stability/efficiency knobs (ref_model_ema_alpha, replay_buffer_size, async_grpo_prefetch, tis_threshold, mask_truncated_completions, defer_rerolling, skip_zero_advantage, off_policy_mask_threshold). Every numeric field rejects bool via a shared `_reject_bool_on_grpo_numerics` field_validator (tdd-guide HIGH fix — Pydantic v2 coerces True→1 by default). The `mask_truncated_completions` + `tis_threshold` pairing is enforced by a cross-validator (matches v0.32.0 spike-recovery+watchdog policy). Part E — top-level task='prm' Literal addition (Process Reward Model / stepwise-supervised, paired with data.format='prm' from v0.42.0) + `vision_grpo: bool` flag for VLM-RL on Qwen2-VL / Pixtral / InternVL. Compat helpers gate task / modality / backend. Review-round fixes applied (5 sequential reviews per CLAUDE.md): - python-review: list_variants annotation, frozenset[str] params, Optional[str] → str | None, module-level math import, D401 imperative docstrings, dropped *args/**kwargs on stubs. - code-review: grpo_fp16 added to GRPO-only task-gate; vllm_sleep_mode requires task='grpo'. - security-review: explicit field_validator for grpo_delta NaN/Inf rejection (Pydantic le=1.0 incidentally rejects NaN, made explicit); null-byte rejection on backend/task in grpo_long_context helpers; use_ring_attention bool guard. - tdd-guide: bool-rejecting field_validator on all Part D numeric fields + grpo_delta; missing bool-rejection tests added on validate_grpo_variant / validate_rollout_backend; null-byte test on validate_vllm_sleep_mode_compat; required_rollout_package rejection path; RolloutBackendSpec.live_wired; PPO+vision_grpo round-trip; _DEFERRED_LIVE invariant. Test count: 6490 → 6729 (+239 across 5 new test files). Notes for future maintainers: - v0.50.0 has zero new CLI commands and zero new trainer wirings; every step 6d/6e is intentionally n/a. Step 6 smoke runs schema happy + every documented cross-validator rejection. - All `task='grpo'` gates use `if self.task != 'grpo'` literal comparisons; do NOT switch to a set membership check until Part D knobs are wired into PPO/preference trainers in v0.50.x. - Multi-modal Vision RL does not yet verify the base model is actually a VLM — upstream trainer surfaces the error loudly when it fails to load the vision tower. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |