Alpamys
|
ed5fc3a8b3
|
feat(precision,rollout): live fp8/nvfp4 + vLLM sleep + openenv rollout + apple-adapter + delinearize-llama4 (v0.71.21)
Closes #141, #124, #125, #228, #97.
- #141: apply_fp8_attention (torchao float8 on attention projections, Hopper
gate) + apply_nvfp4 (NVFP4Config, Blackwell gate); partial-conversion honesty;
wired into the v0.28 speed/memory pipeline with yellow-advisory degrade.
- #124: vllm_sleep_mode live - create_vllm_engine(sleep_mode=True) +
vllm_sleep_cycle ctx (wake in finally) + TRL GRPOConfig hook probe.
- #125: openenv rollout fully live via training.rollout_func module:fn
resolver; rows replace the prompt dataset; art/ruler/nemo_gym honest dep
gates + _EXTERNAL_ROLLOUT_RUNNERS seam. Real GRPO train on SmolLM2-135M.
- #228: convert_apple_adapter live - PEFT LoRA <-> mlx-lm (both matrices
transpose, bf16 upcast, adapters.safetensors + num_layers, npz legacy read,
np.ascontiguousarray fix for safetensors non-contiguous mangling);
*-to-apple upstream-gated exit 3.
- #97: delinearize-llama4 live - [E*din,dout] -> [E,din,dout] per shard,
config.json expert-count probe + --num-experts, sidecar copy, atomic writes.
Review waves: 3 HIGH + ~8 MEDIUM + ~12 LOW fixed.
Tests: 13874 -> 14084 (+210 in tests/test_v07121.py).
Full suite: 13967 passed, 117 skipped. ruff clean.
|
2026-06-10 16:29:52 +05:00 |