mirror of https://github.com/razor-ai/soup.git
4 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
ddab34115c |
feat(v0.26.0): Parts B-E — Eval Gate, Trace-to-Pref, Quant-Check, Soup Cans
Closes the v0.26.0 "Red and Blue Ocean" flywheel after Part A (Registry): Train (eval-gated) -> Registry -> Deploy (quant-check) -> Trace-to-Pref -> Train. Part B — Eval-Gated Training: - soup_cli/config/schema.py: EvalGateConfig (enabled/suite/every_n_epochs/ regression_threshold/baseline/on_regression) + TrainingConfig.eval_gate field - soup_cli/eval/gate.py: EvalSuite, GateTask, run_gate, resolve_baseline, load_suite; baselines from registry:// or file - soup_cli/monitoring/callback.py: on_epoch_end + _run_eval_gate with fail-safe error handling (structured errors treated as regressions under on_regression=stop) - soup_cli/commands/train.py: --gate <suite.yaml> shortcut flag - soup_cli/commands/eval.py: gate subcommand (stub generator; live scoring v0.26.1) Part C — Trace-to-Preference: - soup_cli/data/traces/: parse_langchain, parse_openai, parse_soup_serve; build_pairs from thumbs_up / regenerations / user_edit - soup_cli/commands/data.py: from-traces + review subcommands - PII warning panel, 100,000-line cap, path containment, Literal validation Part D — Quant-Lobotomy Checker: - soup_cli/eval/quant_check.py: classify_delta (OK/MINOR/MAJOR), run_quant_check, resolve_model_ref with artifact kinds filter, table/json/markdown renderers - soup_cli/commands/eval.py: quant-check subcommand Part E — Soup Cans: - soup_cli/cans/: Manifest + DataRef (Pydantic v2); pack_entry + fork_can (100MB cap, dunder-key guard); safe tar extraction (filter='data' on py3.12+, narrow fallback, manual symlink rejection + commonpath check) - soup_cli/commands/can.py: pack/inspect/verify/fork subcommands Shared utility: - soup_cli/utils/paths.py: single is_under_cwd helper replacing 5 duplicates (os.path.realpath + commonpath — Windows 8.3 short-name safe) Tests: 103 new (29 eval_gate + 24 trace_to_pref + 23 quant_check + 27 cans) Full suite: 2511 passed on Windows Python 3.10. Security hardening (review-driven, all severities fixed): - EvalGateConfig bounds; GateTask null-byte + judge URL scheme allowlist - Narrow except in _safe_extract so TarError from filter='data' is not swallowed - resolve_model_ref artifact kinds filter (avoid wrong artifact) - Manifest.author cap + null/newline rejection; created_at ISO-8601 validation - fork_can dunder-key + null-byte rejection (prototype pollution prevention) - fork_can size cap (100MB matches pack_entry) - inspect_can/read_config refuse paths outside cwd Docs: - README.md: v0.26.0 "New in" block (flywheel); 43 recipes; all new commands in All Commands list; version examples bumped to 0.26.0; Windows-safe arrows - CLAUDE.md: architecture + test table + schema + CLI + security section extended with B/C/D/E; phase vs Part terminology clarified; release checklist step 18 adds Known Limitations section; step 20 adds comment template; step 21 adds completeness check via gh issue list --milestone - SECURITY.md: per-Part security notes (B/C/D/E) under v0.26.0 - CONTRIBUTING.md: test count + directory tree updates Local smoke: version, eval gate, eval quant-check (table + json), data from-traces, data review, can pack/inspect/verify/fork — all happy-path end-to-end. Fixed Unicode arrows (U+2192) in can.py + gate.py that crashed on Windows CP1252 consoles. Deferred to v0.26.1 (known limitations, filed as issues post-release): - eval gate/quant-check live model scoring (stub generator currently) - data from-traces quality.py judge validation; serve --trace-log collector - can run + can publish + orchestrator - eval --attach-to-registry flag; export auto-artifact registration Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
|
|
|
e4c3042a56 |
feat(v0.25.0): Beyond the Wrapper — 8 major features
Ships v0.25.0 with eight new capabilities (Parts A–H) that close every competitive gap vs LLaMA-Factory/Axolotl/Unsloth and add unique differentiators: Part A — 9 new model recipes: Llama 4 Scout (sft/dpo/grpo), Qwen 3 14B/32B/8B-grpo, Gemma 3 12B/27B-dpo, DeepSeek V3 (MoE LoRA). Part B — Tool-calling / agentic fine-tuning: new "tool-calling" data format with detection + normalization, synth data template, init template, eval scoring (tool_call_match / tool_call_name_match / tool_call_args_subset), plus qwen3-8b-tools and llama4-scout-tools recipes. Part C — RLVR (RL from Verifiable Rewards): reward_fn=verifiable routing to math_verify_reward (regex-only, no eval), code_exec_reward (subprocess sandbox with RLIMIT_AS/RLIMIT_CPU on POSIX, ephemeral tempdir cwd, concurrency cap, one-time warning panel), and json_schema_reward. verifiable_domain Literal validated via model_validator. Part D — VeRA + OLoRA PEFT methods: LoraConfig.use_vera / use_olora with mutual-exclusion validator and a unified peft_builder helper that returns either LoraConfig or VeraConfig with the right init kwargs. Part E — Apple Silicon MLX backend: detection + hardware profiling in utils/mlx, MLXSFTTrainerWrapper via mlx-lm, scaffolding DPO/GRPO wrappers rejected at config load time by SoupConfig._validate_mlx_task_support, lazy trainer registry, doctor integration, 3 MLX SFT recipes, [mlx] extra in pyproject. Part F — Data augmentation: soup data augment with rephrase / translate / style strategies, path-traversal-protected input/output, count capped 1-10, lang/styles lists bounded (10 entries × 32 chars), rate limiting, and optional --dedup. Part G — Training intelligence: forgetting detection (ForgettingDetector with 3 built-in mini benchmarks and warning levels) and checkpoint intelligence (CheckpointTracker with composite metric, early-stop on regression, safe top-N pruning refusing symlinks and non-checkpoint dirs). SQLite schema extended with checkpoint_quality + forgetting_eval tables. Part H — Autopilot: soup autopilot command with dataset/model/hardware profilers, decision engine (task/quant/peft/batch/lr/epochs/max_length/perf flags), YAML generator, and full CLI with dry-run + --yes + path-traversal protection + goal whitelist + gpu_budget bounds [1GB, 1TB]. Bakes forgetting detection + checkpoint intelligence + early-stop into the generated config. Totals: - 2313 tests passing (183 new, up from 2130) - 86 test files (8 new) - 43 ready-made recipes (14 new) - 16 built-in templates (tool-calling added) - Review findings: all CRITICAL/HIGH/MEDIUM/LOW addressed (3 documented design limitations: code_exec best-effort sandbox, prune_checkpoints TOCTOU, MLX training integration test requires real hardware) Docs: CLAUDE.md, README.md, SECURITY.md, CONTRIBUTING.md updated. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
|
|
|
1b1d679141 |
feat: v0.21.0 — migrate, recipes, NEFTune, rsLoRA
- `soup migrate` — import configs from LLaMA-Factory, Axolotl, Unsloth notebooks (AST-only .ipynb parsing, path traversal protection) - `soup recipes` — 30 ready-made configs for popular models (list/show/use/search with path traversal protection) - NEFTune (`neftune_alpha`) — noisy embeddings for SFT/DPO/KTO/ORPO/SimPO/IPO - rsLoRA (`use_rslora`) — rank-stabilized LoRA scaling in all 11 trainers - Fix: `soup doctor` torchvision circular import crash - Fix: `load_eval_tasks()` now accepts str in addition to Path - Security: Rich markup injection prevention in migration warnings - Security: 10 MB file size limit on migration input files - 1789 tests, 62 test files, 64% coverage |
|
|
|
c46265fd18 |
feat: add eval platform with custom evals, LLM judge, human eval, leaderboard (v0.19.0)
Full-featured evaluation system with 7 subcommands: - soup eval benchmark: standard benchmarks via lm-evaluation-harness - soup eval custom: custom JSONL eval tasks with 4 scoring modes - soup eval judge: LLM-as-a-judge (OpenAI/Ollama/server backends) - soup eval auto: automatic post-training evaluation from config - soup eval compare: side-by-side eval comparison with regression detection - soup eval leaderboard: local model leaderboard with JSON/CSV export - soup eval human: terminal A/B comparison with Elo ratings New modules: soup_cli/eval/ (custom.py, judge.py, human.py, leaderboard.py) Config: EvalConfig added to schema.py (auto_eval, benchmarks, custom_tasks, judge) Callback: SoupTrainerCallback.on_train_end triggers auto-eval when configured Security: SSRF protection on judge API, ReDoS guard on regex scoring, API key isolation per provider, 10k task/prompt caps, read-only SQL queries 1585 tests, 58 test files, ruff clean |