mirror of https://github.com/razor-ai/soup.git
2 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
89b06efb8c |
feat(longctx): v0.53.4 — Long Context + Architecture
Six closes lifting the v0.49.0 LongLoRA hardening + v0.41.0 LLaMA Pro deferred stubs, plus a UX upgrade to the CUDA-OOM friendly message: - #11 utils/errors.py: OOM hint now names --batch-size / --grad-accum - #122 flash_attn.is_flash_attn_v3_available() + LongLoRA+FA3 schema reject - #120 LongLoRA arch allowlist expansion (Mistral / Qwen / Phi); Mixtral intentionally excluded (regex matches the bare 'mistral' token only) - #121 apply_long_context_config auto-detects 'llama3' when caller passes rope_scaling_type=None and the model config carries a Llama 3.1 rope_scaling block - #83 block_expansion.expand_model_blocks LIVE (deepcopy last-N blocks, zero-init residual projections, append, bump num_hidden_layers) + apply_llama_pro_freeze + shared apply_block_expansion_if_configured helper wired into SFT + Pretrain (mirrors v0.40.6 peft_wiring centralisation policy so SFT and Pretrain stay in lock-step) - #74 HF push surface QA — test plan recorded in tests/qa/v053_qa.md; live execution against a private HF repo deferred to a credentialed contributor Review pipeline (python / code / security / tdd agents) ran; every CRITICAL -> LOW finding addressed: - bool-first guards in _check_model_name (defends against int subclass) - is_supported_longlora_arch defensive non-string surface (returns False, never raises) matching v0.53.3 is_known_vlm_base policy - _truncate_for_message(value, limit=64) bounds the base echo in LongLoRA error messages (security MEDIUM, mirrors v0.34.0 crash.py) - null-byte + non-string TypeError guards on validate_longlora_compat task / backend params (matches v0.50.0 validate_long_context_grpo_compat) - _get_layers_module uses explicit `is None` not falsy shortcut (defends against nn.Module.__bool__ overrides on subclasses) - _zero_init_block_residual returns bool + warnings.warn when neither standard projection matches the cloned block (non-Llama-shaped arches still train but lose the LLaMA Pro identity-init guarantee) Test count: 7879 -> 7935 (+56 net; +49 in new tests/test_v0534.py). Lint clean. CPU smoke verified on a real transformers.LlamaForCausalLM: 4 -> 6 layers, down_proj + o_proj actually zeroed on PyTorch tensors, old blocks frozen + new blocks trainable, forward pass finite. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
1491025a36 |
feat(trainer): Optimizer & PEFT Zoo (v0.41.0)
Optimizer Zoo (Part A): closed-allowlist SUPPORTED_OPTIMIZERS adds 14 new
entries — BAdam, APOLLO (apollo_adamw), Adam-mini, lomo / adalomo,
grokadamw, schedule_free_adamw / schedule_free_sgd, muon, dion,
came_pytorch, ao_adamw_{fp8,4bit,8bit}. Validates name + lower-cases
deterministically; rejects non-string / empty / null-byte / >64-char.
_OPTIMIZER_PACKAGES wrapped in MappingProxyType (matches v0.36.0 _REGISTRY).
Per-module LR (Part B): training.lr_groups accepts list-of-pairs /
list-of-dicts / {pattern: lr} mapping; canonical [{pattern, lr}, ...]
storage. Capped at MAX_LR_GROUPS=32; per-pattern non-empty string
≤256 chars + null-byte rejection + re.compile + best-effort ReDoS probe;
per-LR (0.0, 1.0] + math.isfinite (rejects NaN AND ±inf) + bool rejection
(matches v0.30.0 Candidate policy); duplicates rejected. lr_groups_from_schema
bridges canonical schema shape into runtime List[LrGroup] for
build_optimizer_param_groups (first-match-wins routing). LrGroup is
@dataclass(frozen=True). PyYAML scientific-notation (1e-4) parses as
string in YAML 1.1; _validate_lr coerces str → float so soup.yaml
round-trips work. base_lr rejects bool / non-positive (defence-in-depth).
PEFT methods (Part C): LoraConfig.init_strategy="loftq" + loftq_iter
∈ [1, 10] + loftq_bits ∈ {2, 4, 8}; cross-validator rejects loftq +
use_dora / use_vera. utils/loftq_init.py exposes validators +
build_loftq_config (lazy peft.LoftQConfig with actionable ImportError
hint). LLaMA Pro: TrainingConfig.expand_layers ∈ [1, 64] +
freeze_trainable_layers (signed, |x| ≤ 1000); cross-validator requires
the pair (LLaMA Pro freezes original layers and trains only new blocks).
field_validator(mode="before") on both rejects bool BEFORE Pydantic ge/le
silently coerces True → 1. expand_model_blocks raises NotImplementedError
with v0.41.1 marker — schema-only release (mirrors v0.27.0 / v0.37.0
stub-then-live pattern). utils/block_expansion._count_layers uses
hasattr(__len__) instead of try/except TypeError so legitimate __len__
bugs surface loudly. use_mod boolean for Mixture-of-Depths (schema only —
live patch deferred to v0.41.1).
Friendly aliases: load_in_8bit / load_in_16bit (Optional[bool]) for
LlamaFactory / Axolotl users. is True policy on both — explicit False is
"no preference", not "off"; mutually-exclusive both-True rejected; alias
combined with explicit Quant Menu format raises rather than silently
overriding. Alias-driven quantization rewrite via direct
self.quantization = ... (Pydantic v2 BaseModel non-frozen path), NOT
object.__setattr__ — code review caught that the latter would silently
bypass any future field_validator on quantization.
Five review agents (4 in parallel + manual smoke for verification-loop):
all CRITICAL/HIGH/MEDIUM/LOW findings fixed. Local smoke caught a real
bug — PyYAML parsing 1e-4 as string — fixed pre-commit with 2 added tests.
5242 tests pass (+120 net new); ruff clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|