mirror of https://github.com/razor-ai/soup.git
4 Commits
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
06d8ea7cc5 |
chore: migrate to src-layout
Move soup_cli/ -> src/soup_cli/ (history preserved via git mv). src-layout
forces the test suite to import the installed package instead of the
repo-root source tree, surfacing packaging bugs that flat-layout masks —
e.g. the v0.53.8 double-shipped-fixtures regression, invisible because
`pytest tests/` imports ./soup_cli directly and never from the wheel.
- pyproject: packages = ["src/soup_cli"]; artifacts globs -> src/soup_cli/...
The import name is unchanged, so the `soup` entry point, --cov=soup_cli,
and report_to/module-path strings stay `soup_cli`.
- CI / ownership: ruff lint path (ci.yml), recipe-validation `paths:` filters,
CODEOWNERS patterns, and the PR-template checklist all repointed to
src/soup_cli/.
- tests: source-grep regression tests that read package files by repo-relative
path repointed to src/soup_cli/ (64 files; 170 path literals). Lines pushed
over 100 chars by the prefix were wrapped to keep ruff E501 clean. Module
references (`import soup_cli`, `-m soup_cli`, mock.patch("soup_cli.x")) and
the `--cov=soup_cli` coverage target are deliberately unchanged.
- docs: AGENTS.md + CONTRIBUTING.md structure tree and lint commands.
Verified locally: ruff clean (src/soup_cli + tests); `import soup_cli`
resolves to src/soup_cli/__init__.py; built wheel ships
soup_cli/data/_fixtures/*.jsonl (10 files, no duplicates, no src/ prefix);
3174 tests across every touched test file pass. Packaging-only — no version bump.
|
|
|
|
600686cd70 |
feat(advise): soup advise — pre-flight decision (v0.54.0)
`soup advise <data.jsonl> --goal "..."` returns one of PROMPT_ENG / RAG / SFT / DPO / GRPO with a confidence, reason, and reverse-when criterion BEFORE the user spends 8 hours on a GPU. Layer above autopilot — autopilot picks hyperparams AFTER the training decision; advise picks the training decision itself. Three Parts: - Part A: Verdict engine — TASK_CATEGORIES + CHOICES allowlists, frozen Verdict / DatasetProfile / ROIEstimate dataclasses, pure- Python classify_task + compute_dataset_profile + build_verdict rubric (DPO / GRPO floor 500 / PROMPT_ENG floor 50 / RAG / SFT). - Part B: Probe runner — synth_probe_baselines + synth_probe_lora_delta heuristic stubs with forward-compat model/device/lr/timeout_seconds kwargs (v0.54.1 lifts to live model loading per stub-then-live cadence used by v0.27.0 MII / v0.37.0 multipack / v0.50.0 GRPO Plus). - Part C: Cross-project learning — ~/.soup/advise_history.jsonl with cross-process file locking (fcntl on POSIX, sidecar <path>.lock + msvcrt on Windows). `soup advise compare` reads history; env override SOUP_ADVISE_HISTORY_PATH containment-checked to $HOME / $CWD / tempdir (mirrors v0.36.0 SOUP_BATCH_CACHE_PATH policy). CLI: Typer subcommand group `run` / `explain` / `compare` plus argv preprocessor in cli.py that maps `soup advise data.jsonl` → `soup advise run data.jsonl`. Scoped to argv[1] == "advise" only (code-review HIGH fix — defends against rewrites when an unrelated arg contains the literal string "advise"). Schema: AdviseConfig (goal / probe / record) field on SoupConfig honors the plan's cross-cutting bullet. Security: cwd-containment + os.lstat + S_ISLNK symlink reject on every path input; atomic writes via tempfile.mkstemp + os.replace on scratch + history; per-line 64 KB cap + 16 MiB file cap on history reads; bool / finite / NUL / oversize guards on every public input; Rich markup escape on user-controlled output. Reviewed by python / code / security / tdd / architect agents — every finding fixed before commit (0 CRITICAL + 5 HIGH + 7 MEDIUM + 4 LOW). Test count: 8400 → 8571 (+136 in tests/test_v0540.py, +35 net adjustments to v0.53.x version-pin assertions to forward-compat >=). Note: Windows CRLF / LF warnings during stage are .gitattributes- governed and benign. CI runs on ubuntu-latest / windows-latest / macos-latest × Python 3.9 / 3.11 / 3.12. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
cfdabf2b3b |
feat(v0.53.11): GRPO Plus finish + preference live
Closes v0.50.1 (#123, #126, #127), v0.49.1 (#119), v0.40.1 (#68). #123 — live math kernels for 6 GRPO variants (gspo/dapo/dr_grpo/bnpo/ two_sided/rft) + `_GRPOTrainerVariant` HF Trainer subclass via `make_grpo_trainer_variant` factory. Variant compute_loss reads kernel inputs FIRST (no double-forward); falls back to super() only on missing attrs. Case-insensitive variant normalisation before lru_cache. #126 — PRMTrainerWrapper + `_PRMTrainer` HF Trainer subclass with real compute_loss (gather hidden states at step_positions -> reward_head -> MSE via compute_prm_loss). Dataset wrapped in datasets.Dataset.from_list for HF Trainer compatibility. Bool-before-isinstance guard on batch_size. #127 — GRPOStabilityCallback inherits transformers.TrainerCallback (lazy), live EMA ref-model update in on_step_end with strict=True + fallback-to-strict=False-with-WARNING on key mismatch (silent corruption defence). math.isfinite guard on alpha. #119 — LongLoRA forward override via LongLoRAForwardOverride context manager with idempotent install (_soup_longlora_patched marker prevents re-entry double-wrap), 256-char class name cap on regex match, restore on __exit__ AND on exception. #68 — true per-batch weighted-sum preference combine reading policy/ref logps from TRL inputs + each compute_*_term kernel + combine_losses. Explicit None checks on trainer attrs (no `or` on possibly-tensor), DEBUG log on per-term skip. Review fixes from 4 agents (python/code/security/tdd): 10 HIGH + 8 MEDIUM + 7 LOW — see CLAUDE.md v0.53.11 entry for the full list. Test count: 8330 -> 8400 (+75 in test_v05311.py: 54 initial + 21 review-fix coverage gaps). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
|
|
|
76f033a6ff |
feat(v0.53.10): Quick wins + packaging + UX wiring
7 issues closed: - #150 [mix] pyproject extra bundles scikit-optimize so `soup data mix --optimize` runs the Bayesian loop instead of the v0.48.0 Dirichlet fallback; new describe_default_optimizer() helper labels the active backend without paying skopt's import cost. - #113 [data-pro] extras (langdetect + presidio-analyzer) with lazy fall-through helpers in utils/data_score (broader language coverage + Presidio entity recognition on top of the v0.47.0 regex baseline). Llama-Guard-3-1B documented as a manual recipe (license + size). - #154 SOUP_POSTHOG_KEY / SOUP_POSTHOG_ENDPOINT env override via sentinel-based explicit-vs-env precedence; HTTPS-only + RFC1918/link-local rejection on the endpoint; null-byte / control-char / >256-char rejection on the key. - #152 --hub flag plumbed on chat / serve / infer / merge / export / push via shared utils/hubs.apply_hub_to_cli_model + prefetch_model_from_hub helpers; push uses upload_repo (skips HF-specific Collections + model-card auto-render on non-HF hubs). - #153 `soup data download --hub modelscope|modelers` live SDK (lifts the v0.53.8 advisory-only path); friendly ImportError advisory when the SDK is missing. - #155 Web UI Tool Outputs panel — `loadToolOutputs` polls /api/tool-outputs every 3s; XSS-safe DOM-built table (textContent per cell, no innerHTML for user-controlled fields); Bearer token threaded via the v0.53.9 window._authToken bootstrap. - #156 SoupTrainerCallback.on_step_end records tool_calls counts from kwargs['inputs'] into the global tool buffer. Best-effort (# noqa: BLE001 per project policy — training must never crash). 13 review-fixes applied (4 HIGH / 5 MEDIUM / 4 LOW): - HIGH PostHog explicit-endpoint precedence sentinel - HIGH absolute path leak in local_path advisory reduced to relpath - HIGH Rich markup escape on base / local_path / cache_dir - HIGH callback # noqa: BLE001 per project policy - MED `import time` moved out of try block - MED oversize key + explicit-empty key rejection tests - MED source-grep regression guards (advisory-removal, helper imports across 5 non-push commands) - MED `prefetch_model_from_hub` outside-cwd cache_root rejection - LOW empty-list + bool-True tool_calls no-op tests - LOW push.py uses upload_repo + validate_hub_name regression guard Test count: 8285 -> 8330 (+45 in tests/test_v05310.py). Full suite green; ruff clean; on Win+Py3.10. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |