mirror of https://github.com/razor-ai/soup.git
3 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
892fd33f9e |
feat(trainers): v0.35.0 — Trainer Coverage (closes #60, #61, #45)
Wires v0.28.0 speed/memory features into every transformer-backend trainer (grpo / kto / orpo / simpo / ipo / ppo / reward_model / embedding) plus closes the v0.33.0 #43 oversight where dpo / pretrain accepted activation_offloading without installing offload hooks. Auto-quant --auto-quant now forwards the picked candidate's quantization to vLLM via an explicit named parameter (kwarg-splat hazard removed). Kernel auto-compose runs a forward-only benchmark loop on the trainer's actual model under torch.no_grad() so live training gradients aren't polluted (this was a critical-class bug caught by code-review pre-tag and fixed before merge). Schema gate lifted with distinct MLX-backend vs unknown-task error messages so users get the right fix. fp8 / int8 QAT guard fixed in 6 trainers (the legacy unguarded `if tcfg.quantization_aware:` would have crashed the int8 path with the string "fp8"). Net +187 tests (3928 -> 4115). New file tests/test_trainer_coverage_v035.py provides a parametrised matrix proof that every trainer x every feature is exercised on every CI matrix job. All four review-agent waves (python / code / security / tdd) clean with every CRITICAL -> LOW finding fixed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
55d1b9312c |
feat(speed,memory): v0.28.0 features go multi-trainer (v0.33.0 Part C)
Closes #43, #44, #47. #43 Multi-trainer wiring (sft/dpo/pretrain): - New utils/v028_features.apply_v028_speed_memory(model, tcfg, base_model, console) — single shared helper for use_cut_ce, quantization_aware="fp8", kernel_auto_compose. Each feature degrades silently to a yellow advisory if the underlying lib is missing; never crashes training kick-off. - Helpers supports_v028_features(task) and warn_unsupported_features(tcfg, task) drive both the schema validator and runtime advisories. - soup_cli/trainer/dpo.py and trainer/pretrain.py now call the helper after model load (post-LoRA, post-QAT) — same hook point as SFT. - soup_cli/config/schema.py validator _validate_v028_speed_memory_sft_only renamed _validate_v028_speed_memory_supported_tasks; allowlist now {sft, dpo, pretrain}. GRPO/KTO/ORPO/SimPO/IPO/PPO/RewardModel/Embedding still error out at config-load with a precise multi-trainer message. #44 Selective gradient-checkpoint hooks: - New utils/gradient_ckpt.install_selective_hooks(model, granularity) iterates ``model.named_modules()`` looking for transformer-block-shaped names (numeric suffix on layer path), wraps each module's ``forward`` with torch.utils.checkpoint.checkpoint based on tier: - selective: only attention sub-modules - medium: every second transformer block - full: every transformer block - Returns hook count so callers can fall back to HF native checkpointing when zero blocks were found. #47 CrossDocCollator: - New soup_cli/data/collators.CrossDocCollator wraps any base data collator and injects a block-diagonal causal ``cross_doc_attn_mask`` built from per-example ``doc_lengths``. Preferred over TRL's ``packing_strategy="attention_free"`` flag (best-effort across TRL versions). Degrades gracefully when doc_lengths is missing or shapes don't match — base attention_mask preserved, no crash. Tests: +16 in tests/test_part_c.py covering apply_v028_speed_memory (no-features, cut_ce graceful failure), supports/warn helpers extension, schema gate (dpo + pretrain accept, kto still rejects), selective hook installation across full/medium/selective with fake transformer-shaped models, CrossDocCollator passthrough + strip + injection. One existing test in test_training_speed.py updated: dpo+use_cut_ce now accepted. Known limitations: - 7 trainers (GRPO/KTO/ORPO/SimPO/IPO/PPO/RewardModel/Embedding) still reject v0.28.0 flags at config-load. Each is a 5-line addition once schema validation is satisfied; tracked as a v0.33.x follow-up. - install_selective_hooks doesn't undo earlier hooks — caller must be re-init aware. Not an issue for the typical "construct wrapper, train, exit" flow but worth noting. - CrossDocCollator emits ``cross_doc_attn_mask`` (not ``attention_mask``) to avoid clobbering the base collator's contract; downstream consumers must read the new key explicitly. The plan calls for "preferred over TRL's packing_strategy" which we satisfy via opt-in collation, not silent override. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
e09a742167 |
feat(training): Training Speed & Memory — CCE, FP8, grad-ckpt tiers, kernel picker, cross-doc attn, activation offload (v0.28.0)
Six new training speed/memory features, SFT-only in v0.28.0: - use_cut_ce: Cut Cross-Entropy for 128k-vocab models (8-24GB save) - quantization_aware: "fp8" — Hopper+ float8 training via torchao.float8 - gradient_checkpointing: bool | selective|medium|full|auto (VRAM-based auto) - kernel_auto_compose: benchmark + pick fastest kernel combo - packing_cross_doc_attn_mask: block-diagonal mask for sample packing - activation_offloading: cpu|disk saved-tensor offload Config-load validator rejects non-SFT tasks when speed/memory flags are set — prevents int8-QAT-wrapper crash on the string "fp8" and silent no-ops on DPO/GRPO/KTO/ORPO/SimPO/IPO/PPO/Pretrain/Reward/Embedding. Multi-trainer wiring tracked for v0.28.1. Security: - FP8 path: CUDA + Hopper+ SM capability + transformers backend - Activation-offload disk: is_under_cwd containment, TOCTOU-safe mkstemp (fd held through torch.save), weights_only=True reload, crash-safe cleanup - Kernel picker raises when all candidates lack finite time_ms - Cut CE detector matches last path component only (deepseek-ai/...-phi-... org-prefix does not trigger Phi patch on DeepSeek) - Cross-doc mask numpy-vectorised (np.tril) at max_length=1M - @model_validator gates: packing_cross_doc_attn_mask requires packing=true; v0.28.0 features require task=sft New optional extra: pip install 'soup-cli[cce]' Tests: 2585 -> 2685 (+100 in tests/test_training_speed.py, +1 file). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |