mirror of https://github.com/razor-ai/soup.git
1 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
725696b1da |
feat(v0.53.1): Quant Menu II + Export pipeline live
Lift six v0.53.0 deferred stubs from NotImplementedError to live wiring: - #82 autopilot pre-quantized base detection utils name regex over gptq/awq/aqlm/eetq/fp8/mxfp4 with word-boundary anchoring + HQQ Nbit extraction + config.json quantization_config probe (cwd-contained + symlink-rejected). decide_quantization() short-circuits the VRAM heuristic when prequantized is set. autopilot pipeline auto- applies so TheBloke/Llama-2-7B-Chat-GPTQ is recommended gptq instead of 4bit-on-top-of-quantized. - #142 merge_4bit + export_torchao live writers soup merge --save-format {fp16|4bit|4bit_forced}: single BNB-4bit merged checkpoint without the dequant->merge->requant cycle (fixes wrong-name llm_int8_skip_modules to bnb_4bit_skip_modules per code- review). soup export --format torchao --quant-config <yaml>: torchao .quantize_ + save_pretrained with per-scheme closed kwarg allowlist (Int4WeightOnly accepts {group_size, inner_k_tiles}, NVFP4 accepts nothing extra; dunder + unknown keys rejected per security-review H1). load_quant_config enforces yaml.safe_load + 256 KB cap + extension allowlist + cwd containment + S_ISLNK rejection. - #139 export_advanced_gguf via llama.cpp imatrix 3-stage pipeline: convert_hf_to_gguf.py -> optional imatrix -> quantize. argv-list subprocess (no shell), 30-min timeout, realpath- verified convert script stays inside llama_cpp_dir (security-review M5). _prepare_calibration_text accepts JSONL with text/prompt/content field aliases + raw text fallback; strips null bytes, collapses newlines, 8 KB per-line + 50 MB total cap (security-review M1); POSIX O_NOFOLLOW closes the TOCTOU window between dispatch-time check and open() (security-review M3). UD- prefix stripped before passing to llama-quantize. _safe_stderr Rich-escapes subprocess stderr before embedding in RuntimeError (security-review L4). - #109 soup deploy autopilot --measure Live Quant-Lobotomy scorecard: classifies each candidate quant OK / MINOR / MAJOR (thresholds 2% / 5% mirror v0.26.0 Part D). Results cached at ~/.soup/deploy_autopilot_cache.json (atomic write, 0o600 perms on POSIX, S_ISLNK rejection on BOTH load and save). pick_best soft-fallback now picks max-by-delta (was max-by-after) matching the v0.33.0 #54 design intent. _DEPLOY_MEASURE_BEFORE_GEN / _AFTER_FACTORY module-level hooks act as the stop-gap escape hatch until v0.46.1 ships first-party transformers / vLLM generator factories. - #70/#72 manual QA log scripted at tests/qa/v053_qa.md with exact reproduction recipes + acceptance criteria for the CUDA + llama.cpp smokes that can't run on the CI runners. Shared cleanup: - soup_cli/utils/paths.enforce_under_cwd_and_no_symlink consolidates the v0.33.0 #22 TOCTOU pattern previously copy-pasted in save_formats.py and gguf_quant.py (code-review HIGH fix). Reviews ran: python / code / security / tdd. Every CRITICAL / HIGH / MEDIUM / LOW finding fixed or documented. Test count: 7610 -> 7722 (+112 across 4 new files). Known limitations: live GPU + bitsandbytes / torchao smokes for the new merge / export paths remain pending (recipes in QA log); injected- generator escape hatch is non-public until v0.46.1; cache key truncates base_sha to 16 hex (1-in-2^32 collision floor). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |