mirror of https://github.com/razor-ai/soup.git
2 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
f6bc8e7bdd |
feat(ship): make soup ship's leg-2 regression gate real (v0.71.38)
soup ship's leg 2 — the catastrophic-forgetting / regression gate that carries the whole SHIP / DON'T-SHIP claim — was 15 trivia prompts scored by raw substring containment (it credited "B" for "Berlin", "3" for "13") with zero coverage for tool-calling, safety, or JSON. This makes the gate real. - forgetting.py: score_answer/extract_mcq_letter replace the substring scorer with answer-extraction (cue -> paren -> clause-terminating bare letter) + boundary-aware token match. MINI_BENCHMARKS expanded (mmlu 26 / common_sense 24 / instruction 24) + new mini_arithmetic (36) so a 1-item flip trips 0.05. BREAKING: an existing run's verdict can change (the old gate under-reported). - eval/gate_suites.py (new): bundled offline general-suite registry, no torch. DEFAULT_GENERAL_SUITE = the 4 MCQ suites + 3 behavioural JSONL suites (mini_tool_call / mini_format_json / mini_safety) scored per-model-absolute by the pure custom/diagnose scorers. _fraction_passing isolates a per-item scorer exception (deep-JSON RecursionError scores as a failed item). - ship.py: leg-2 scores bundled suites offline (base+tuned) before routing any non-bundled name to lm-eval; default general suite = the full bundled set. Exit-code taxonomy: usage errors move 2 -> 3 so exit 2 means only DON'T-SHIP (a typo'd flag was previously indistinguishable from a caught regression). - diagnose/__init__: "Six" -> "Seven" probes + re-export all 7 score_* fns; removed the dead SUPPORTED_TASK_MODES "pairwise reserved" gate. - Bundled gate fixtures ship in the wheel via the pyproject artifacts glob. Every bundled item is original, hand-authored (no MMLU/GSM8K rows copied). Test count 16288 -> 16330 (+42 in tests/test_v07138.py). |
|
|
|
e4c3042a56 |
feat(v0.25.0): Beyond the Wrapper — 8 major features
Ships v0.25.0 with eight new capabilities (Parts A–H) that close every competitive gap vs LLaMA-Factory/Axolotl/Unsloth and add unique differentiators: Part A — 9 new model recipes: Llama 4 Scout (sft/dpo/grpo), Qwen 3 14B/32B/8B-grpo, Gemma 3 12B/27B-dpo, DeepSeek V3 (MoE LoRA). Part B — Tool-calling / agentic fine-tuning: new "tool-calling" data format with detection + normalization, synth data template, init template, eval scoring (tool_call_match / tool_call_name_match / tool_call_args_subset), plus qwen3-8b-tools and llama4-scout-tools recipes. Part C — RLVR (RL from Verifiable Rewards): reward_fn=verifiable routing to math_verify_reward (regex-only, no eval), code_exec_reward (subprocess sandbox with RLIMIT_AS/RLIMIT_CPU on POSIX, ephemeral tempdir cwd, concurrency cap, one-time warning panel), and json_schema_reward. verifiable_domain Literal validated via model_validator. Part D — VeRA + OLoRA PEFT methods: LoraConfig.use_vera / use_olora with mutual-exclusion validator and a unified peft_builder helper that returns either LoraConfig or VeraConfig with the right init kwargs. Part E — Apple Silicon MLX backend: detection + hardware profiling in utils/mlx, MLXSFTTrainerWrapper via mlx-lm, scaffolding DPO/GRPO wrappers rejected at config load time by SoupConfig._validate_mlx_task_support, lazy trainer registry, doctor integration, 3 MLX SFT recipes, [mlx] extra in pyproject. Part F — Data augmentation: soup data augment with rephrase / translate / style strategies, path-traversal-protected input/output, count capped 1-10, lang/styles lists bounded (10 entries × 32 chars), rate limiting, and optional --dedup. Part G — Training intelligence: forgetting detection (ForgettingDetector with 3 built-in mini benchmarks and warning levels) and checkpoint intelligence (CheckpointTracker with composite metric, early-stop on regression, safe top-N pruning refusing symlinks and non-checkpoint dirs). SQLite schema extended with checkpoint_quality + forgetting_eval tables. Part H — Autopilot: soup autopilot command with dataset/model/hardware profilers, decision engine (task/quant/peft/batch/lr/epochs/max_length/perf flags), YAML generator, and full CLI with dry-run + --yes + path-traversal protection + goal whitelist + gpu_budget bounds [1GB, 1TB]. Bakes forgetting detection + checkpoint intelligence + early-stop into the generated config. Totals: - 2313 tests passing (183 new, up from 2130) - 86 test files (8 new) - 43 ready-made recipes (14 new) - 16 built-in templates (tool-calling added) - Review findings: all CRITICAL/HIGH/MEDIUM/LOW addressed (3 documented design limitations: code_exec best-effort sandbox, prune_checkpoints TOCTOU, MLX training integration test requires real hardware) Docs: CLAUDE.md, README.md, SECURITY.md, CONTRIBUTING.md updated. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |