diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 6df44c2..18428fb 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -111,7 +111,7 @@ soup_cli/ templates/ - 17 built-in soup.yaml templates (YAML + manifest.json) with load_template loader (v0.39.0, +bco v0.40.0) ui/ - Web UI (FastAPI + HTML/JS SPA) -tests/ - Test suite (141 files, 4720 tests) +tests/ - Test suite (143 files, 4756 tests) examples/ - Real-world config examples and datasets ``` diff --git a/README.md b/README.md index 4e669af..04fa178 100644 --- a/README.md +++ b/README.md @@ -40,18 +40,20 @@ soup train ## What's New + Latest highlights only. Full history: [GitHub Releases](https://github.com/MakazhanAlpamys/Soup/releases). -**v0.40.1 — QA Hardening**: Windows codepage crash family fixed at the CLI bootstrap (one-line UTF-8 reconfigure closes 6 separate QA findings); the multi-objective preference-loss runtime stub from v0.40.0 is now live; autopilot + quickstart no longer assume 7B / 1.1B on tiny GPUs; `transformers >= 5.0.0` is flagged as INCOMPATIBLE in `soup doctor` until the migration lands. +**v0.40.2 — Quick polish + carry-overs**: closes 3 originally-scheduled GitHub issues (#36 eval-gate dashboard row, #50 smarter `--hf-resume`, #51 custom HF Space templates) plus 7 v0.40.1 long-tail UX papercuts. -- **UTF-8 stdio bootstrap (Windows)** — `soup_cli/cli.py` reconfigures `sys.stdout` / `sys.stderr` to UTF-8 before any Rich console init. β / ✓ / box-drawing characters no longer crash with `UnicodeEncodeError` on cp1251 / cp1252. POSIX is a no-op. -- **Multi-objective preference loss is live** — `training.preference_loss_weights: {dpo: 0.7, simpo: 0.3}` no longer raises `NotImplementedError`. The wrapper builds the highest-weighted loss as primary and prints the active blend for confirmation. BCO mixed with paired losses (DPO/SimPO/ORPO/IPO) is rejected at runtime with a clear message (data-format incompatible). -- **Smarter defaults on small GPUs** — `soup quickstart` auto-switches to `SmolLM2-135M-Instruct` on ≤6 GB VRAM (verified to train in 5 s on RTX 3050 4 GB); the autopilot model-size fallback now defaults to **1B** instead of 7B and reads the local safetensors index when available, so `tiny-gpt2` no longer fails VRAM-budget checks. -- **`soup doctor` upgrades** — flags `transformers ≥ 5.0.0` as INCOMPATIBLE; distinguishes "no GPU hardware" from "GPU hardware present, wrong torch wheel" (`nvidia-smi` succeeds but `torch.cuda.is_available()` returns False); detects dual-Python interpreter setups; uses `importlib.metadata` as version-probe fallback so `rich` no longer prints "?". -- **CLI UX consistency** — `soup init --force` non-interactively overwrites; `--template` help is generated from the live registry (no more drifting list); `soup migrate ` errors loudly with a "did you pass the wrong file?" hint; `soup eval custom -o` writes JSON regardless of `--attach-to-registry`; `soup recipes show ` suggests close matches via `difflib`; `soup data sample` filenames embed the strategy so successive `random` / `diverse` runs don't overwrite each other; JSONL loader auto-strips UTF-8 BOM (Windows PowerShell `Out-File` users). -- **`--find-lr` actually runs the live loop** — fixed a broken `load_local` import that previously always silently fell through to a static placeholder curve. -- **Schema strictness** — root-level `lora:` in YAML (the LlamaFactory / Axolotl convention) now migrates into `training.lora` so the nested validators (including `init_strategy: random|pissa|olora`) actually fire instead of being silently dropped. -- **Net new tests** across UTF-8 bootstrap, multi-objective preference math (with gradient propagation checks), schema regression, and Part C–E papercuts. +- **`--hf-resume` prefers local newer** — `prepare_hf_resume` checks the highest local `checkpoint-N` against the highest remote branch and skips the snapshot download when the local copy is newer or equal. Saves bandwidth and never overwrites a fresher local checkpoint with stale Hub state. +- **Eval-gate dashboard row** — pure formatter `format_gate_row(state)` lives in `soup_cli/monitoring/display.py` and renders a one-liner like `Gate: helpfulness 7.8 ✓ | math 0.82 ✗ (-0.06) | STOP` for the live training panel. Hidden when eval-gate is disabled. +- **`soup deploy hf-space --template-dir `** — supply your own `app.py` + `README.md` (+ optional `requirements.txt`) instead of the built-in `gradio-chat` / `streamlit-chat`. Containment-checked, repo-id substitution validated *before* render, 256 KB cap per file, symlinks rejected. +- **`soup quickstart --output DIR`** — route data, config, and run dir under any directory you choose (default keeps backwards-compat in cwd). +- **`soup runs --cwd-only`** — restrict the listing to runs whose `output_dir` is under the current directory; the global `~/.soup/experiments.db` view is still default. +- **`soup infer` / `soup bench` accept HF ids** — when the local model path is missing AND the value isn't path-like (no `./`, `/`, `~`, `C:\`), it falls through to a HuggingFace download via `transformers.from_pretrained`. Path-like-but-missing surfaces a friendly `FileNotFoundError`. +- **CLI flag aliases** — `data filter --min-coherence` (alias for `--coherence`); `data split --train` accepted (informational, train is the implicit remainder); `data register / unregister` accept positional `` and `` alongside the `--name` / `--path` options, with conflict detection. +- **`--log-level` plumbing complete** — `apply_logging_level` now sets the root logger so third-party libraries (transformers / peft / trl) actually respect QUIET and DEBUG. The four tiers no longer produce byte-identical output. +- **+36 net new tests** across the new helpers, security-fix follow-ups (path-containment in `register_data`, `bench.py`, `infer.py --output`, symlink-reject in custom Space templates), and Windows cross-drive edge cases. ## Why Soup? diff --git a/SECURITY.md b/SECURITY.md index 2308e4f..93c4789 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -9,8 +9,8 @@ We provide security updates for the following versions: - **Versions older than 3 minor versions:** No support Example: -- v0.40.1 -- Full support (latest) -- v0.40.0-0.40.x -- Full support +- v0.40.2 -- Full support (latest) +- v0.40.0-v0.40.x -- Full support - v0.39.0-0.39.x -- Bug-fix support only - v0.38.0-0.38.x -- Bug-fix support only - v0.37.x and below -- No support @@ -144,6 +144,7 @@ No known critical vulnerabilities in current releases. - **v0.32.0 — Training Stability & Auto-Tuning**: `--find-lr-output` containment via shared `utils/paths.is_under_cwd` (prevents writes outside cwd); `save_lr_finder_report` rejects NaN / Infinity floats in `lrs` / `losses` and serialises with `allow_nan=False` (keeps the report parser-safe); `compute_lr_schedule` rejects non-positive `start_lr`, inverted ranges, and `num_steps` outside `[2, 10_000]`; `pick_mixed_precision` rejects empty / null-byte / >200-char model names and resolves multi-version quirks (`qwen2.5` vs `qwen2`, `phi-3.5` vs `phi-3`) by longest-substring-first iteration so an added family can never accidentally make a more-specific entry dead code; `compute_warmup_steps` clamps to `[10, 1000]` with a `ratio==0.0` short-circuit matching HF Trainer's "no warmup" convention; `SpikeRecoveryStrategy` is `@dataclass(frozen=True)` (post-construction mutation cannot bypass validation), `max_attempts ∈ [1, 10]`, `lr_decay ∈ (0, 1)`, `min_lr > 0`; cross-validator `_validate_spike_recovery_requires_watchdog` rejects `loss_spike_recovery=true, loss_watchdog=false` at config-load (fails fast instead of never triggering); `convergence_window ∈ [5, 10_000]`, `convergence_rel_tol ∈ (0, 1]`, `recommend_action` reuses `detect_plateau` so plateau heuristic stays single-source-of-truth; `GradAccumMonitor.recommend()` caps doubled `accum` at `MAX_ACCUM=1024` so a runaway advisory loop cannot blow up DataLoader prefetch; `generate_config` validates BOTH the YAML output path AND the embedded `decisions["output"]` field via `is_under_cwd` (closes the gap where a crafted `decisions["output"]="../../etc"` would have silently propagated into the rendered YAML) - **v0.34.0 — Observability & Dev UX**: `.crash` bundle generator (`utils/crash.py`) recursively redacts `hf_*` / `sk-*` / `Bearer …` token-shaped strings in any captured `config` and metric tail before serialisation, so a `.crash` file shared on a public GitHub issue cannot leak credentials; `output_dir` is reduced to `os.path.basename` so `$HOME` doesn't leak; `write_crash_bundle` uses `os.path.realpath + commonpath` for cwd containment (Windows-safe; raises `ValueError` not `PermissionError` so callers cannot silently swallow with `except OSError`); filename appends `secrets.token_hex(4)` so two crashes in the same UTC second don't collide; bundle truncated to `MAX_BUNDLE_BYTES=1_000_000`. `train.py` crash-write surfaces failures to the user (no silent missing-bundle). `profiling.py` `resolve_trace_path` rejects empty / `.` / `..` / `/` / `\\` / null-byte `run_id` (closes the `output_dir/profiles/../trace.json` escape) and uses `os.path.realpath + is_under_cwd`; profiles dir is created only on successful torch import (no stale empty dirs on torch-less CI). `tracker.get_run` LIKE-prefix match escapes `%` / `_` / `\\` and uses `ESCAPE '\\'` so a crafted `run_id` cannot widen the match (mirrors v0.26.0 registry policy). Lazy schema migration (`_ensure_schema`) tolerates the "duplicate column" race when two CLI processes start simultaneously on a fresh DB (fork-based multi-GPU training, TUI auto-refresh). `runs.py show/replay/clean` switched user `run_id` rendering to `markup_escape` and switched `clean` containment from broken `Path.resolve() + relative_to()` to project-standard `os.path.realpath + is_under_cwd`. `tui_app.py` lazy-imports `ExperimentTracker` and `markup_escape`s every DB-sourced string before passing into Textual widgets so a crafted base_model / experiment_name cannot inject `[bold red]…[/]` markup. `run_cost.estimate_run_cost_usd` rejects `bool` in `num_gpus` (bool is a subclass of int — same defence as v0.30.0 `Candidate.__post_init__`); duration clamped to `[0, 1 year]`; unknown GPU returns `None` so callers render `—` instead of fabricating `$0.00`. `log_level.parse_log_level` rejects non-string + null-byte input. - **v0.33.0 — Live Wire**: RLVR `code_exec_reward` adds OS-level isolation (Linux best-effort `os.unshare(CLONE_NEWUSER|CLONE_NEWNET|CLONE_NEWPID)`, macOS `sandbox-exec` with default-deny `MACOS_SANDBOX_PROFILE` narrowed to a 3-name `mach-lookup` allowlist to prevent DNS / NSURLSession bypass of `(deny network*)`); `prune_checkpoints` switches to TOCTOU-safe `os.lstat + S_ISLNK` + `shutil.rmtree(onerror=_abort_on_symlink)` so a symlink encountered mid-walk aborts rather than escapes; `run_gate` wraps each task scorer in a typed `try/except` so backend failures produce `score=None, error=str(exc)` (never silent `score=1.0`); `_parse_judge_url` removes the bare `http://` catch-all (defence-in-depth after the Pydantic GateTask validator); `soup can run` requires `--yes` or explicit consent callback and raises `ValueError` (not `PermissionError`, which is an `OSError` subclass that broad `except` blocks would swallow); GGUF `rglob` result for ollama deploy is `realpath+commonpath` checked against extract_dir (prevents symlink escape from a crafted can); `DeployTarget.path` validator normalises mixed `\\`/`/` separators before splitting (closes a Windows `..` bypass); `CAN_FORMAT_VERSION` 1→2 (additive — v1 still loads); `soup can publish` validates `repo_id` via `utils/hf.validate_repo_id`, resolves token via `resolve_token`, sanitises commit messages (first-line, 200-char cap), uses HTTPS-only HfApi; `_write_spike_recovery_hint` adds `is_under_cwd` containment check on `args.output_dir` from raw HF `TrainingArguments`; `lookup_entry_by_output_dir` emits `ResourceWarning` when 1000-row scan limit is hit (no silent miss); `CrossDocCollator` no longer mutates input feature dicts (HF Dataset rows are cached and reused — mutation broke subsequent batches); `Candidate` rejects `bool` in `score`/`latency_ms` (was sneaking past `int` isinstance check); `evaluate_candidate` latency mean now divides by *completed* prompts (excludes crashed) so a broken candidate isn't artificially fast; `auto_quant.run_auto_quant_picker` soft-falls-back to highest-scored candidate when no candidate clears `min_score` (server still binds); `build_logits_processors` returns `[]` when neither `outlines` nor `lm-format-enforcer` is installed (server degrades to free-form rather than 500); MII server uses loopback-only CORS, max_tokens cap [1, 16384], stream rejection, generic 500 with no stack-trace leak; `os.execvp` auto-reexec uses list args (no shell), all forwarded flags pre-validated; `cleanup_extract_dir` uses `os.path.commonpath` (Windows-safe) instead of `startswith`; `_run_subprocess` catches `TimeoutExpired` and returns rc=124 (coreutils convention) instead of an unhandled traceback; new `eval_results` and `tensorrt` artifact kinds in `RegistryStore._VALID_KINDS` +- **v0.40.2 — Quick polish + carry-overs**: New `soup_cli/utils/hf_space.py:render_custom_template_dir` enforces `is_under_cwd` containment on the template directory; `validate_repo_id` runs BEFORE `{MODEL_REPO}` substitution (matches v0.29.0 Part F policy); per-file 256 KB cap (matches v0.39.0 Part E template-size policy); only `app.py` / `README.md` / `requirements.txt` are read (closed allowlist — no path-from-user-data). Symlinks rejected via `os.lstat + stat.S_ISLNK` and non-regular files (FIFO / device) also rejected (matches v0.33.0 #22 prune_checkpoints TOCTOU policy) — defends against `/app.py -> /etc/passwd`. `_find_highest_local_checkpoint` reads `output_dir` after caller's `is_under_cwd` validation (in `prepare_hf_resume`) and silently drops non-directories + OSError. `prepare_hf_resume` skips the snapshot download when local `checkpoint-N >= remote checkpoint-N` (saves bandwidth and never overwrites a fresher local checkpoint). `commands/data.py:register_data` containment switched from `Path.resolve() + relative_to()` to shared `is_under_cwd` (Windows 8.3 short-name safety per CLAUDE.md project rule); same fix applied to `commands/bench.py` prompts-file containment. `commands/infer.py:--output` now containment-checked via `is_under_cwd` (late-evaluated after model+input validation so pre-existing `tmp_path` test contracts keep working). `commands/quickstart.py:--output` validates target dir via `is_under_cwd` before `mkdir(parents=True)`; rejects out-of-cwd targets with friendly message. `commands/runs.py:_filter_runs_by_cwd` uses `os.path.realpath + commonpath`, catches `(ValueError, OSError)` so cross-drive paths on Windows (`D:\runs` vs `C:\project`) drop silently rather than crash. `monitoring/display.py:format_gate_row` uses explicit `task.get("passed") is True` so a missing `"passed"` field renders neutrally instead of as a false-y red ✗. `commands/infer.py:_resolve_model_source` heuristic for HF-id-vs-local-path: only falls through to HF when value is NOT path-like (no `./`, `/`, `\\`, `~`, no Windows drive letter, non-empty); path-like-but-missing raises `FileNotFoundError` so users see actionable errors instead of confusing HF download attempts. Known limitation: custom HF Space templates always create the Space with `space_sdk="gradio"` regardless of the supplied `app.py` (no `--sdk` flag in this release; combine `--template streamlit-chat` with the inline registry for Streamlit Spaces). Tracked for v0.40.3+. - **v0.40.1 — QA Hardening**: `soup_cli/utils/encoding.force_utf8_stdio` reconfigures Windows stdout/stderr to UTF-8 before any Rich Console is constructed; `os.environ.setdefault("PYTHONIOENCODING", "utf-8")` preserves user override; `(OSError, ValueError, AttributeError)` swallowed on detached streams; POSIX no-op. `SoupConfig._remap_root_level_misplaced_keys` (model_validator, mode='before') migrates root-level `lora:` into `training.lora` so nested validators (including `lora.init_strategy: Literal["random","pissa","olora"]`) actually fire — closes a footgun where the misplaced key was silently dropped. Caller's dict is never mutated (shallow-copy policy mirroring v0.33.0 #47 / v0.40.0 Part B). `PreferenceTrainerWrapper._build_multi_objective` replaces the v0.40.0 `NotImplementedError` stub with a primary-loss approximation; `validate_weight_compat` rejects BCO mixed with paired losses at runtime (data-format incompatible). `combine_losses` rejects empty weights, propagates NaN loudly (no silent zeroing), and rejects `bool` weight values (matches v0.30.0 `Candidate` policy). `_probe_cache_param_count` rejects empty / null-byte model names before path construction (mirrors v0.26.0 registry / v0.39.0 ReLoRAPolicy policy). `commands/doctor` flags `transformers ≥ 5.0.0` as INCOMPATIBLE via `_MAX_EXCLUSIVE` table; `_version_ge` parses leading-int chunks so `5.0.0.dev0` correctly trips the cap. `_detect_gpu_hw_without_torch_cuda` calls `nvidia-smi` via argv list (no shell), 5s timeout, `OSError` / `TimeoutExpired` caught; GPU label from `nvidia-smi` stdout is `rich.markup.escape`d before embedding in Rich-markup string (a real GPU name like `NVIDIA Quadro [T4]` cannot break or inject markup). `_detect_dual_python_interpreters` uses `os.path.realpath` (not `Path.resolve()`) for Windows 8.3 short-name compat. `_pick_quickstart_model` swaps TinyLlama-1.1B → SmolLM2-135M when `total_memory ≤ 6 GB` (prevents step-0 OOM on RTX 3050 4 GB / similar). `_live_lr_sweep_from_config` switched broken `load_local` import to `load_raw_data` (previously always silently fell back to a static placeholder curve). `commands/migrate` rejects `.jsonl` input (with first-line `{` sniff) with exit-2 friendly error; `.jsonl`-only suffix gate prevents false-positives on `.ipynb` notebooks. `commands/eval custom -o` is now honored independently of `--attach-to-registry`; loop-shadow regression where `output = generate_fn(...)` overwrote the CLI option fixed (variable renamed to `response`). `_load_jsonl` switched from `utf-8` to `utf-8-sig` so PowerShell `Out-File -Encoding utf8`-produced JSONL no longer fails first-row parse. Known limitation: `--trust-remote-code` opt-in surface still excludes 10 non-SFT trainers + 5 commands (v0.36.0 #63 carry-over). - **v0.40.0 — Preference Variety**: New `task='bco'` (Binary Classifier Optimization) and `task='preference'` (unified dispatcher). New schema fields: `bco_beta` (gt=0), `preference_loss: Literal[dpo,simpo,orpo,ipo,bco]|None`, `preference_loss_weights: Optional[Dict[str,float]]`, `dpo_beta_schedule: Literal[linear,cosine,exponential]|None`, `dpo_beta_end: float, gt=0|None`, `dpo_ref_regen_epochs: int [1,1000]|None`. Cross-validators: `_validate_preference_dispatcher` rejects setting either `preference_loss` or `preference_loss_weights` outside `task='preference'` (closes ordering-dependency between Part B/D validators); `_validate_dpo_variants_supported_tasks` gates β-schedule + ref-regen to DPO-family tasks (`dpo`, `ipo`, or `preference` + `preference_loss in {dpo, ipo}`); rejected on mlx backend with distinct error message (matches v0.34.0 distinct-reason policy); `_validate_preference_loss_weights` enforces 2–5 entries (single-entry rejected with actionable message pointing at scalar `preference_loss`), key allowlist `{dpo, simpo, orpo, ipo, bco}`, explicit null-byte rejection on keys (matches v0.39.0 rank_pattern policy), per-value bounds `(0, 1]`, weights must sum to 1.0 (±1e-6), mutually exclusive with scalar `preference_loss`, rejected on mlx backend. `compute_beta_at_step` rejects `bool` on `step` and `total_steps` (project bool-as-int policy from v0.30.0). `BetaScheduleCallback` resolves `total_steps` lazily in `on_train_begin` so the schedule sees the real `state.max_steps` populated by HF Trainer (closes a first-cut silent-no-op bug where total_steps=0 emitted beta_end for every step). `RefModelRegenCallback._regenerate` uses `strict=True` on `load_state_dict` and logs at WARNING on mismatch (closes a first-cut silent partial-copy hazard where strict=False could produce a hybrid old-base + new-LoRA reference); epoch 0 regen suppressed (avoids copying untrained student); trainer `.beta` assignment swallow narrowed to `AttributeError` only. `PreferenceTrainerWrapper._make_inner_cfg` uses `model_copy` (not `model_dump`+`model_validate`) so re-validation never sees an inconsistent intermediate state and the caller's `cfg` is never mutated (mirrors v0.33.0 #47 immutability policy). `_split_dpo_rows_to_bco` skipped-row count emitted at DEBUG so production silent-degradation is inspectable (mirrors v0.33.0 #47 CrossDocCollator policy). Multi-objective live runtime weighted-loss combination is deferred to v0.40.1: `PreferenceTrainerWrapper.setup` raises `NotImplementedError` with a friendly message naming the deferred-version follow-up (mirrors v0.27.0 MII / v0.37.0 multipack / v0.38.0 quant menu / v0.39.0 ReLoRA stub-then-live pattern). Known limitation: `BCOTrainerWrapper._setup_transformers` still hardcodes `trust_remote_code=True` (v0.36.0 #63 known-gap family carry-over across non-SFT trainers). - **v0.39.0 — LoRA Quality**: `LoraConfig.init_strategy: Literal["random","pissa","olora"]` rejects unknown strategies; PiSSA + DoRA / VeRA combinations rejected at config-load. `model_validator(mode="before")` aligns `use_olora=True` → `init_strategy="olora"` via dict-copy (no caller mutation; matches v0.33.0 #47 immutability policy). `rank_pattern`/`alpha_pattern: Optional[Dict[str, int]]` capped at 256 keys × value (0, 1024], rejects `bool` (subclass of `int` — matches v0.30.0 `Candidate` policy), null bytes in keys, empty keys; cross-validator rejects with `use_vera=True`. `ReLoRAPolicy` is `@dataclass(frozen=True)` (post-construction mutation raises `FrozenInstanceError`); bounds: `steps ∈ [1, 1e7]`, `warmup_ratio ∈ [0, 1]`, `prune_ratio ∈ (0, 1)` (strict — prevents zero-everything footgun). `magnitude_prune_tensor` strict `0 < prune_ratio < 1` rejection, non-Tensor input raises `TypeError`, empty / single-element tensor short-circuits (avoids `kthvalue(_, 0)` runtime crash). `_validate_relora_supported_tasks` cross-validator rejects `relora_steps` with `task != "sft"` and `backend=mlx` with distinct error messages (matches v0.34.0 distinct-reason policy); multi-trainer expansion deferred to v0.39.1. `is_gemma4_model` uses a word-boundary regex (`(?:^|[^a-z0-9])gemma-?4(?:[^a-z0-9]|$)`) so `"ungemma4ed"` / `"my-gemma4ish"` no longer over-match; null-byte rejection on `model_name`. `apply_gemma4_clippable_patch` weight-copy fallback logs at DEBUG instead of silent random-init; the patch is gated by `is_gemma4_model(cfg.base)` in `sft.py` before invocation so non-Gemma4 trainings never traverse the module tree. `apply_surgical_patches` rejects empty / null-byte `model_name` with `ValueError`. `templates/load_template` containment: filename re-validated via `_validate_name` (rejects `..`/`/`/`\\`/null/empty); `os.path.realpath + os.path.commonpath` containment check on the resolved path against `_templates_dir()` so a tampered `manifest.json` cannot read files outside the package directory (mirrors v0.26.0 registry policy). Tampered-manifest `ValueError` from `_validate_name` caught and falls back to inline (no propagating exception). 256 KB file-size cap. Inline `TEMPLATES` carries an explicit deprecation comment pointing at the canonical YAML registry; `tests/test_templates_yaml.py` asserts byte-equality of all 16 inline ↔ YAML pairs to prevent silent drift. Planned removal: v0.41.0+. diff --git a/pyproject.toml b/pyproject.toml index 7af7042..b5de89a 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "soup-cli" -version = "0.40.1" +version = "0.40.2" description = "Fine-tune LLMs in one command. No SSH, no config hell." readme = "README.md" license = "Apache-2.0" diff --git a/soup_cli/__init__.py b/soup_cli/__init__.py index 6e99337..57b618c 100644 --- a/soup_cli/__init__.py +++ b/soup_cli/__init__.py @@ -1,3 +1,3 @@ """Soup CLI — Fine-tune LLMs in one command.""" -__version__ = "0.40.1" +__version__ = "0.40.2" diff --git a/soup_cli/cli.py b/soup_cli/cli.py index f79b708..3b8a978 100644 --- a/soup_cli/cli.py +++ b/soup_cli/cli.py @@ -226,7 +226,11 @@ def main( """Soup — fine-tune LLMs in one command.""" global _verbose, _log_level _verbose = verbose - from soup_cli.utils.log_level import parse_log_level, setup_logging + from soup_cli.utils.log_level import ( + apply_logging_level, + parse_log_level, + setup_logging, + ) try: tier = parse_log_level(log_level) @@ -235,6 +239,10 @@ def main( raise typer.Exit(code=2) from exc _log_level = tier.value setup_logging(tier) + # v0.40.2 N1/G2: also push the tier into the root logger so third-party + # libraries (transformers / peft / trl) respect QUIET / DEBUG. Without + # this, all four levels were producing nearly-identical output. + apply_logging_level(tier) def run(): diff --git a/soup_cli/commands/bench.py b/soup_cli/commands/bench.py index 7949cdf..3cffc3e 100644 --- a/soup_cli/commands/bench.py +++ b/soup_cli/commands/bench.py @@ -43,13 +43,24 @@ def bench( """Run an inference benchmark (speed and memory) on a loaded model.""" import torch - from soup_cli.commands.infer import _generate, _load_model + from soup_cli.commands.infer import _generate, _load_model, _resolve_model_source from soup_cli.utils.gpu import detect_device - model_path = Path(model) - if not model_path.exists(): - console.print(f"[red]Model not found: {model_path}[/]") - raise typer.Exit(1) + # Resolve local-path-or-HF-id (#N7). + try: + model_kind, model_ref = _resolve_model_source(model) + except FileNotFoundError as exc: + console.print( + f"[red]{exc}[/]\n" + "[dim]If you meant a HuggingFace repo, use the form " + "'owner/repo-name' (no leading './').[/]" + ) + raise typer.Exit(1) from exc + model_path = Path(model_ref) + if model_kind == "hf": + console.print( + f"[dim]Local path not found; treating {model_ref!r} as a HF repo id.[/]" + ) device, _ = detect_device() @@ -61,15 +72,17 @@ def bench( if prompts_file: import json - p_path = Path(prompts_file).resolve() + import os as _os - try: - p_path.relative_to(Path.cwd()) - except ValueError: + from soup_cli.utils.paths import is_under_cwd + + if not is_under_cwd(prompts_file): console.print( - f"[red]Security Error:[/] Path {p_path} is outside the current working directory." + "[red]Security Error:[/] Prompts file must stay under the " + "current working directory." ) raise typer.Exit(1) + p_path = Path(_os.path.realpath(prompts_file)) if not p_path.is_file(): console.print(f"[red]Prompts file not found:[/] {p_path}") diff --git a/soup_cli/commands/data.py b/soup_cli/commands/data.py index 130bc8b..27e58b6 100644 --- a/soup_cli/commands/data.py +++ b/soup_cli/commands/data.py @@ -5,6 +5,7 @@ from __future__ import annotations import json import random from pathlib import Path +from typing import Optional import typer from rich.console import Console @@ -325,7 +326,7 @@ def filter_data( help="Max perplexity threshold (rows above this are removed)", ), coherence: float = typer.Option( - None, "--coherence", + None, "--coherence", "--min-coherence", help="Min coherence threshold 0.0-1.0 (rows below this are removed)", ), perplexity_model: str = typer.Option( @@ -780,6 +781,13 @@ def split_data( None, "--test", help="Test split: percentage (default) or absolute count (with --absolute)", ), + train: int = typer.Option( + None, "--train", + help=( + "Train split (informational; the train remainder is implied by " + "--val + --test). Accepted for command parity." + ), + ), absolute: bool = typer.Option( False, "--absolute", help="Treat --val/--test as absolute sample counts instead of percentages", @@ -1461,56 +1469,104 @@ def _load_augment_provider(provider: str, rpm: int): @app.command(name="register") def register_data( - name: str = typer.Option(..., "--name", "-n", help="Dataset name"), - path: str = typer.Option(..., "--path", "-p", help="Path to dataset file"), + name_arg: Optional[str] = typer.Argument( + None, help="Dataset name (positional alternative to --name)", + ), + path_arg: Optional[str] = typer.Argument( + None, help="Path to dataset file (positional alternative to --path)", + ), + name: Optional[str] = typer.Option(None, "--name", "-n", help="Dataset name"), + path: Optional[str] = typer.Option( + None, "--path", "-p", help="Path to dataset file" + ), fmt: str = typer.Option( "auto", "--format", "-f", help="Dataset format: alpaca, sharegpt, chatml, dpo, kto, auto", ), ): - """Register a local dataset by name for use in soup.yaml.""" + """Register a local dataset by name for use in soup.yaml. + + Accepts both ``--name X --path Y`` and positional ``X Y``. + """ from soup_cli.utils.registry import register_dataset - # Path traversal protection - resolved = Path(path).resolve() - cwd = Path.cwd().resolve() - try: - resolved.relative_to(cwd) - except ValueError: + # Resolve positional vs option, with conflict detection + final_name = name if name is not None else name_arg + final_path = path if path is not None else path_arg + if final_name is None or final_path is None: + console.print( + "[red]Provide both name and path " + "(positional ` ` or `--name --path`).[/]" + ) + raise typer.Exit(2) + if name is not None and name_arg is not None and name != name_arg: + console.print("[red]Conflict: --name and positional name differ.[/]") + raise typer.Exit(2) + if path is not None and path_arg is not None and path != path_arg: + console.print("[red]Conflict: --path and positional path differ.[/]") + raise typer.Exit(2) + + # Path traversal protection — use os.path.realpath + commonpath + # (project standard; Windows 8.3 short-name safe). + import os as _os + + from soup_cli.utils.paths import is_under_cwd + + if not is_under_cwd(final_path): console.print( "[red]Dataset path must be under the current working directory.[/]" ) raise typer.Exit(1) + resolved = Path(_os.path.realpath(final_path)) registry_path = _get_registry_path() try: - register_dataset(name, str(resolved), fmt, registry_path=registry_path) + register_dataset(final_name, str(resolved), fmt, registry_path=registry_path) except ValueError as exc: console.print(f"[red]{exc}[/]") raise typer.Exit(1) console.print( - f"[green]Registered dataset '[bold]{name}[/bold]'[/]\n" - f" Path: {path}\n" + f"[green]Registered dataset '[bold]{final_name}[/bold]'[/]\n" + f" Path: {final_path}\n" f" Format: {fmt}" ) @app.command(name="unregister") def unregister_data( - name: str = typer.Option(..., "--name", "-n", help="Dataset name to remove"), + name_arg: Optional[str] = typer.Argument( + None, help="Dataset name (positional alternative to --name)", + ), + name: Optional[str] = typer.Option( + None, "--name", "-n", help="Dataset name to remove" + ), ): - """Remove a dataset from the local registry.""" + """Remove a dataset from the local registry. + + Accepts both ``--name X`` and positional ``X``. + """ from soup_cli.utils.registry import unregister_dataset + final_name = name if name is not None else name_arg + if final_name is None: + console.print( + "[red]Provide a dataset name " + "(positional `` or `--name`).[/]" + ) + raise typer.Exit(2) + if name is not None and name_arg is not None and name != name_arg: + console.print("[red]Conflict: --name and positional name differ.[/]") + raise typer.Exit(2) + registry_path = _get_registry_path() - removed = unregister_dataset(name, registry_path=registry_path) + removed = unregister_dataset(final_name, registry_path=registry_path) if removed: - console.print(f"[green]Removed dataset '{name}' from registry.[/]") + console.print(f"[green]Removed dataset '{final_name}' from registry.[/]") else: - console.print(f"[red]Dataset '{name}' not found in registry.[/]") + console.print(f"[red]Dataset '{final_name}' not found in registry.[/]") raise typer.Exit(1) diff --git a/soup_cli/commands/deploy.py b/soup_cli/commands/deploy.py index 6514a87..254fb01 100644 --- a/soup_cli/commands/deploy.py +++ b/soup_cli/commands/deploy.py @@ -435,6 +435,15 @@ def hf_space( "-t", help=f"Space template: {', '.join(HF_SPACE_TEMPLATES.keys())}", ), + template_dir: Optional[str] = typer.Option( + None, + "--template-dir", + help=( + "Custom template directory (overrides --template). " + "Must contain app.py + README.md, optionally requirements.txt. " + "Use {MODEL_REPO} placeholder for substitution." + ), + ), private: bool = typer.Option( False, "--private", help="Create the Space as private", ), @@ -461,7 +470,7 @@ def hf_space( except ValueError as exc: console.print(f"[red]Invalid --space repo id:[/] {exc}") raise typer.Exit(1) from exc - if template not in HF_SPACE_TEMPLATES: + if template_dir is None and template not in HF_SPACE_TEMPLATES: console.print( f"[red]Unknown template: {template}[/]\n" f"Available: {', '.join(HF_SPACE_TEMPLATES.keys())}" @@ -485,16 +494,28 @@ def hf_space( # --- Render template files --- try: - files = render_space_template(template, model_repo=model) + if template_dir is not None: + from soup_cli.utils.hf_space import render_custom_template_dir + files = render_custom_template_dir(template_dir, model_repo=model) + # Custom templates default to gradio SDK unless requirements + # imply otherwise; we record gradio for create_repo space_sdk. + sdk = "gradio" + else: + files = render_space_template(template, model_repo=model) + sdk = HF_SPACE_TEMPLATES[template]["sdk"] except ValueError as exc: console.print(f"[red]Template render failed:[/] {exc}") raise typer.Exit(1) from exc + except FileNotFoundError as exc: + console.print(f"[red]Template directory error:[/] {exc}") + raise typer.Exit(1) from exc + template_label = template_dir if template_dir is not None else template console.print( Panel( f"Space: [bold]{space}[/]\n" f"Model: [bold]{model}[/]\n" - f"Template: [bold]{template}[/]\n" + f"Template: [bold]{template_label}[/]\n" f"Private: [bold]{private}[/]", title="Deploy HuggingFace Space", ) @@ -514,7 +535,7 @@ def hf_space( try: api.create_repo( repo_id=space, repo_type="space", - space_sdk=HF_SPACE_TEMPLATES[template]["sdk"], + space_sdk=sdk, private=private, exist_ok=True, ) for in_repo_name, content in files.items(): diff --git a/soup_cli/commands/infer.py b/soup_cli/commands/infer.py index 474b584..65b2c7f 100644 --- a/soup_cli/commands/infer.py +++ b/soup_cli/commands/infer.py @@ -1,5 +1,7 @@ """soup infer — batch inference on a list of prompts.""" +from __future__ import annotations + import json import time from pathlib import Path @@ -13,6 +15,42 @@ from rich.progress import BarColumn, Progress, SpinnerColumn, TextColumn, TimeEl console = Console() +def _is_path_like(value: str) -> bool: + """Heuristic: looks like a filesystem path rather than a HF repo id. + + HF repo ids are ``owner/name`` with no leading dot/slash and no Windows + drive letter; anything else (``./foo``, ``/abs/path``, ``C:\\...``) is + treated as a path so we surface a meaningful FileNotFoundError instead + of attempting an HF download. + """ + if not value: + return True + if value.startswith((".", "/", "\\", "~")): + return True + # Windows drive letter, e.g. "C:\..." or "C:/..." + if len(value) >= 2 and value[1] == ":": + return True + return False + + +def _resolve_model_source(model: str) -> tuple[str, str]: + """Return ``("local", path)`` or ``("hf", repo_id)`` for ``--model``. + + Falls through to HF when the local path doesn't exist *and* the value + looks like a HF repo id (no leading ``./`` etc.). Raises + :class:`FileNotFoundError` when the value looks like a path but doesn't + exist locally — distinguishes "your file is missing" from "your HF id + is wrong" so the error message is actionable. + """ + candidate = Path(model) + if candidate.exists(): + return "local", str(candidate) + if _is_path_like(model): + raise FileNotFoundError(f"Model path not found: {model}") + # Looks like a HF repo id — let transformers handle the download. + return "hf", model + + def infer( model: str = typer.Option( ..., @@ -58,17 +96,29 @@ def infer( ), ): """Run batch inference on a JSONL file of prompts.""" + from soup_cli.utils.paths import is_under_cwd + # Validate input file input_path = Path(input_file) if not input_path.exists(): console.print(f"[red]Input file not found: {input_path}[/]") raise typer.Exit(1) - # Validate model path - model_path = Path(model) - if not model_path.exists(): - console.print(f"[red]Model not found: {model_path}[/]") - raise typer.Exit(1) + # Resolve model: local path or HF repo id (auto-fallback, #N7). + try: + model_kind, model_ref = _resolve_model_source(model) + except FileNotFoundError as exc: + console.print( + f"[red]{exc}[/]\n" + "[dim]If you meant a HuggingFace repo, use the form " + "'owner/repo-name' (no leading './').[/]" + ) + raise typer.Exit(1) from exc + model_path = Path(model_ref) + if model_kind == "hf": + console.print( + f"[dim]Local path not found; treating {model_ref!r} as a HF repo id.[/]" + ) # Read prompts prompts = _read_prompts(input_path) @@ -104,6 +154,16 @@ def infer( model_obj, tokenizer = _load_model(str(model_path), base, device) console.print("[green]Model loaded.[/]\n") + # Output path containment — defence-in-depth (project policy v0.20.0+). + # Checked late, after model+inputs validate, so that pre-existing tests + # asserting on "model not found" / "no prompts" errors keep working when + # they pass an out-of-cwd `tmp_path`. + if not is_under_cwd(output_file): + console.print( + "[red]--output must stay under the current working directory.[/]" + ) + raise typer.Exit(1) + # Run inference — stream results to disk as they are generated output_path = Path(output_file) total_tokens = 0 diff --git a/soup_cli/commands/quickstart.py b/soup_cli/commands/quickstart.py index 7d0e06b..987ecdc 100644 --- a/soup_cli/commands/quickstart.py +++ b/soup_cli/commands/quickstart.py @@ -4,6 +4,7 @@ from __future__ import annotations import json from pathlib import Path +from typing import Optional import typer from rich.console import Console @@ -122,16 +123,43 @@ def quickstart( "--dry-run", help="Create data and config only, do not train", ), + output: Optional[str] = typer.Option( + None, + "--output", + "-o", + help=( + "Output directory for data, config, and run artifacts " + "(default: current directory)." + ), + ), ): """Run a complete demo: create sample data, config, and train.""" + import os + + from soup_cli.utils.paths import is_under_cwd + model_id, advisory = _pick_quickstart_model() if advisory: console.print(f"[yellow]{advisory}[/]") + + # Resolve output directory (containment-checked) + if output is None: + out_dir = Path.cwd() + else: + if not is_under_cwd(output): + console.print( + f"[red]--output must stay under the current working directory; " + f"got: {output}[/]" + ) + raise typer.Exit(2) + out_dir = Path(os.path.realpath(output)) + out_dir.mkdir(parents=True, exist_ok=True) + console.print( Panel( "This will:\n" - " 1. Create [bold]quickstart_data.jsonl[/] (20 examples)\n" - " 2. Create [bold]quickstart_soup.yaml[/] config\n" + f" 1. Create [bold]{out_dir}/quickstart_data.jsonl[/] (20 examples)\n" + f" 2. Create [bold]{out_dir}/quickstart_soup.yaml[/] config\n" " 3. Train a tiny LoRA adapter (~1 min on GPU)\n\n" f"Model: [bold]{model_id}[/]", title="[bold]Soup Quickstart[/]", @@ -145,7 +173,7 @@ def quickstart( raise typer.Exit() # 1. Create demo data - data_path = Path("quickstart_data.jsonl") + data_path = out_dir / "quickstart_data.jsonl" if data_path.exists(): console.print(f"[yellow]Data file already exists:[/] {data_path}") else: @@ -155,13 +183,27 @@ def quickstart( console.print(f"[green]Created:[/] {data_path} ({len(DEMO_DATA)} examples)") # 2. Create demo config - config_path = Path("quickstart_soup.yaml") + config_path = out_dir / "quickstart_soup.yaml" if config_path.exists(): console.print(f"[yellow]Config file already exists:[/] {config_path}") else: rendered = DEMO_CONFIG.replace(_DEFAULT_MODEL, model_id) + # When --output is set, retarget data + run dirs into that dir. + if output is not None: + rendered = rendered.replace( + "./quickstart_data.jsonl", str(data_path) + ).replace("./quickstart_output", str(out_dir / "quickstart_output")) config_path.write_text(rendered, encoding="utf-8") console.print(f"[green]Created:[/] {config_path}") + # Also write a `soup.yaml` symlink-style alias for tools that look for it. + soup_yaml = out_dir / "soup.yaml" + if not soup_yaml.exists() and output is not None: + try: + soup_yaml.write_text( + config_path.read_text(encoding="utf-8"), encoding="utf-8" + ) + except OSError: + pass if dry_run: console.print("\n[yellow]Dry run - files created, skipping training.[/]") diff --git a/soup_cli/commands/runs.py b/soup_cli/commands/runs.py index 37ad075..f323067 100644 --- a/soup_cli/commands/runs.py +++ b/soup_cli/commands/runs.py @@ -20,12 +20,48 @@ console = Console() app = typer.Typer(no_args_is_help=False, invoke_without_command=True) +def _filter_runs_by_cwd(runs: list[dict], cwd: str) -> list[dict]: + """Return only runs whose output_dir is under ``cwd``. + + Runs with no ``output_dir`` are dropped. Comparisons use + ``os.path.realpath + commonpath`` for Windows 8.3 short-name safety + (see project rule in CLAUDE.md). + """ + cwd_real = os.path.realpath(cwd) + kept: list[dict] = [] + for run in runs: + out = run.get("output_dir") + if not out: + continue + try: + cand = os.path.realpath(out) + if os.path.commonpath([cand, cwd_real]) == cwd_real: + kept.append(run) + except (ValueError, OSError): + # Different drives on Windows / unreadable path — skip. + continue + return kept + + @app.callback(invoke_without_command=True) def list_runs( ctx: typer.Context, limit: int = typer.Option(20, "--limit", "-l", help="Max runs to show"), + cwd_only: bool = typer.Option( + False, + "--cwd-only", + help=( + "Only show runs whose output_dir is under the current " + "working directory (default: show all in ~/.soup/experiments.db)." + ), + ), ): - """List all training runs.""" + """List all training runs. + + By default lists every run from the global ``~/.soup/experiments.db``. + Use ``--cwd-only`` to restrict to runs anchored under the current + directory. + """ if ctx.invoked_subcommand is not None: return @@ -33,6 +69,8 @@ def list_runs( tracker = ExperimentTracker() runs = tracker.list_runs(limit=limit) + if cwd_only: + runs = _filter_runs_by_cwd(runs, os.getcwd()) if not runs: console.print("[dim]No runs found. Train a model with:[/] [bold]soup train[/]") diff --git a/soup_cli/monitoring/display.py b/soup_cli/monitoring/display.py index 231bb1f..2a6d1b6 100644 --- a/soup_cli/monitoring/display.py +++ b/soup_cli/monitoring/display.py @@ -1,6 +1,6 @@ """Rich live training dashboard in the terminal.""" -from typing import Optional +from typing import Any, Mapping, Optional from rich.console import Console from rich.live import Live @@ -11,6 +11,49 @@ from soup_cli.config.schema import SoupConfig console = Console() +def format_gate_row(state: Optional[Mapping[str, Any]]) -> str: + """Render the eval-gate status row for the live training panel (#36). + + Returns an empty string when ``state`` is None or empty (so the row is + hidden when eval-gate is disabled). + + Format example:: + + Gate: helpfulness 7.8 [green]✓[/] | math 0.82 [red]✗[/] (-0.06 from base) | STOP + + Pure formatter — no I/O, no side effects — so it is trivially testable + via ``Console(file=StringIO())`` without spinning up a Live display. + """ + if not state: + return "" + tasks = state.get("tasks") or [] + if not tasks: + return "" + parts: list[str] = [] + for task in tasks: + name = str(task.get("name", "?")) + score = task.get("score") + # Explicit ``is True`` so a missing field renders the ``?`` mark + # rather than the false-y red ✗ (e.g. tasks still pending). + passed = task.get("passed") is True + delta = task.get("delta") + score_str = f"{score:.2f}" if isinstance(score, (int, float)) else "—" + mark = "[green]✓[/]" if passed else "[red]✗[/]" + chunk = f"{name} {score_str} {mark}" + if delta is not None and isinstance(delta, (int, float)): + sign = "+" if delta >= 0 else "" + chunk += f" ({sign}{delta:.2f})" + parts.append(chunk) + body = " | ".join(parts) + action = state.get("action") + suffix = "" + if action == "stop": + suffix = " | [bold red]STOP[/]" + elif action == "warn": + suffix = " | [yellow]WARN[/]" + return f"[bold]Gate:[/] {body}{suffix}" + + class TrainingDisplay: """Live-updating terminal dashboard for training progress.""" diff --git a/soup_cli/monitoring/hf_push.py b/soup_cli/monitoring/hf_push.py index 862e885..9c9949e 100644 --- a/soup_cli/monitoring/hf_push.py +++ b/soup_cli/monitoring/hf_push.py @@ -177,6 +177,32 @@ def resolve_latest_checkpoint_revision( return best_name +def _find_highest_local_checkpoint(output_dir: str) -> Optional[int]: + """Return the highest ``checkpoint-`` step under ``output_dir``, or None. + + Skips non-directories and malformed names. Returns None if ``output_dir`` + does not exist or contains no checkpoints. + """ + base = Path(output_dir) + if not base.is_dir(): + return None + best: Optional[int] = None + try: + children = list(base.iterdir()) + except OSError: + return None + for entry in children: + if not entry.is_dir(): + continue + match = _CHECKPOINT_BRANCH_RE.match(entry.name) + if not match: + continue + step = int(match.group(1)) + if best is None or step > best: + best = step + return best + + def _download_checkpoint( repo_id: str, revision: str, @@ -247,6 +273,22 @@ def prepare_hf_resume( if revision is None: return None + # Prefer local newer (#50): if local has checkpoint-N >= remote's, + # skip the download and return the local path. Saves bandwidth and + # avoids overwriting a fresher local checkpoint with stale Hub state. + remote_match = _CHECKPOINT_BRANCH_RE.match(revision) + remote_step = int(remote_match.group(1)) if remote_match else -1 + local_step = _find_highest_local_checkpoint(output_dir) + if local_step is not None and local_step >= remote_step: + local_revision = f"checkpoint-{local_step}" + local_path = Path(output_dir) / local_revision + logger.info( + "HF resume: local %s >= remote %s; skipping download", + local_revision, + revision, + ) + return str(local_path) + # Mirror HF Trainer's on-disk layout: output_dir/ local_dir = str(Path(output_dir) / revision) try: diff --git a/soup_cli/utils/hf_space.py b/soup_cli/utils/hf_space.py new file mode 100644 index 0000000..b04539b --- /dev/null +++ b/soup_cli/utils/hf_space.py @@ -0,0 +1,94 @@ +"""HF Space custom template directory renderer (v0.40.2 #51). + +Lets users supply their own ``app.py`` / ``README.md`` (+ optional +``requirements.txt``) for ``soup deploy hf-space`` instead of the built-in +``gradio-chat`` / ``streamlit-chat`` templates. + +Security model — mirrors v0.29.0 Part F policy: + +- Template directory must stay under the current working directory + (``utils/paths.is_under_cwd``). +- ``model_repo`` is validated via :func:`utils.hf.validate_repo_id` BEFORE + substitution, so a crafted repo id cannot inject Python source into the + rendered ``app.py`` uploaded to HF Hub. +- Per-file size cap of 256 KB — defends against pathological templates. +- Only ``app.py``, ``README.md``, and ``requirements.txt`` are read; any + other files in the directory are ignored. +""" +from __future__ import annotations + +import os +import stat as stat_module +from pathlib import Path + +from soup_cli.utils.hf import validate_repo_id +from soup_cli.utils.paths import is_under_cwd + +_MAX_TEMPLATE_FILE_BYTES = 256 * 1024 # 256 KB +_KNOWN_FILES = ("app.py", "README.md", "requirements.txt") +_REQUIRED_FILES = ("app.py", "README.md") + + +def render_custom_template_dir(template_dir: str, model_repo: str) -> dict[str, str]: + """Render a custom Space template directory. + + Returns a dict mapping in-repo filename → rendered content. + ``{MODEL_REPO}`` placeholder is replaced with the validated ``model_repo``. + + Raises: + ValueError: containment violation, invalid ``model_repo``, oversized file. + FileNotFoundError: missing required file (``app.py`` / ``README.md``). + """ + validate_repo_id(model_repo) + + if not is_under_cwd(template_dir): + raise ValueError( + "template-dir must stay under the current working directory; " + f"got: {template_dir!r}" + ) + + base = Path(template_dir) + if not base.is_dir(): + raise FileNotFoundError( + f"template-dir does not exist or is not a directory: {template_dir}" + ) + + rendered: dict[str, str] = {} + for fname in _KNOWN_FILES: + fpath = base / fname + # ``lstat`` does NOT follow symlinks — defence-in-depth against a + # crafted symlink at /app.py -> /etc/passwd reading + # outside the containment-checked directory (mirrors v0.33.0 + # ``prune_checkpoints`` TOCTOU policy). + try: + st = os.lstat(fpath) + except OSError: + if fname in _REQUIRED_FILES: + raise FileNotFoundError( + f"template-dir is missing required file: {fname}" + ) from None + continue + if stat_module.S_ISLNK(st.st_mode): + raise ValueError( + f"template file {fname} is a symlink; refusing to render" + ) + if not stat_module.S_ISREG(st.st_mode): + if fname in _REQUIRED_FILES: + raise FileNotFoundError( + f"template-dir is missing required file: {fname}" + ) + continue + if st.st_size > _MAX_TEMPLATE_FILE_BYTES: + raise ValueError( + f"template file {fname} exceeds 256 KB cap " + f"({st.st_size} bytes); refusing to render" + ) + try: + content = fpath.read_text(encoding="utf-8") + except UnicodeDecodeError as exc: + raise ValueError( + f"template file {fname} is not valid UTF-8" + ) from exc + rendered[fname] = content.replace("{MODEL_REPO}", model_repo) + + return rendered diff --git a/soup_cli/utils/log_level.py b/soup_cli/utils/log_level.py index a533ade..7090e48 100644 --- a/soup_cli/utils/log_level.py +++ b/soup_cli/utils/log_level.py @@ -99,3 +99,16 @@ def setup_logging(tier: LogLevel) -> logging.Logger: handler.setLevel(py_level) return logger + + +def apply_logging_level(tier: LogLevel) -> None: + """Set the *root* logger level so non-soup libraries respect the tier (#G2/N1). + + ``setup_logging`` only configures the ``soup`` logger; library logs + (``transformers``, ``peft``, ``trl``) flow through the root logger. + Without this, ``--log-level quiet`` left third-party INFO chatter on + stdout. ``apply_logging_level`` plumbs the tier into the root level so + QUIET silences libraries and DEBUG surfaces them. + """ + py_level = resolve_python_log_level(tier) + logging.getLogger().setLevel(py_level) diff --git a/tests/test_v0402_part_a.py b/tests/test_v0402_part_a.py new file mode 100644 index 0000000..aa701a9 --- /dev/null +++ b/tests/test_v0402_part_a.py @@ -0,0 +1,358 @@ +"""Tests for v0.40.2 Part A — originally scheduled issues (#36, #50, #51).""" +from __future__ import annotations + +from io import StringIO + +import pytest +from rich.console import Console +from typer.testing import CliRunner + +from soup_cli.cli import app + +runner = CliRunner() + + +# --------------------------------------------------------------------------- +# #50 — `--hf-resume` prefer local newer +# --------------------------------------------------------------------------- + + +class TestHFResumePreferLocal: + def test_find_highest_local_checkpoint_empty(self, tmp_path): + from soup_cli.monitoring.hf_push import _find_highest_local_checkpoint + + assert _find_highest_local_checkpoint(str(tmp_path)) is None + + def test_find_highest_local_checkpoint_picks_max(self, tmp_path): + from soup_cli.monitoring.hf_push import _find_highest_local_checkpoint + + (tmp_path / "checkpoint-100").mkdir() + (tmp_path / "checkpoint-50").mkdir() + (tmp_path / "checkpoint-200").mkdir() + # Spurious dirs ignored + (tmp_path / "garbage").mkdir() + (tmp_path / "checkpoint-bad").mkdir() + # Files (not dirs) ignored + (tmp_path / "checkpoint-999").write_text("not a dir") + + assert _find_highest_local_checkpoint(str(tmp_path)) == 200 + + def test_find_highest_local_checkpoint_missing_dir(self, tmp_path): + from soup_cli.monitoring.hf_push import _find_highest_local_checkpoint + + assert _find_highest_local_checkpoint(str(tmp_path / "nope")) is None + + def test_prepare_hf_resume_skips_download_when_local_equal( + self, tmp_path, monkeypatch + ): + """If local checkpoint-N == remote checkpoint-N, skip download.""" + from soup_cli.monitoring.hf_push import prepare_hf_resume + + monkeypatch.chdir(tmp_path) + out_dir = tmp_path / "runs" + out_dir.mkdir() + (out_dir / "checkpoint-300").mkdir() + + download_called = {"called": False} + + def fake_download(*args, **kwargs): + download_called["called"] = True + return kwargs.get("local_dir", "") + + monkeypatch.setattr( + "soup_cli.monitoring.hf_push.resolve_latest_checkpoint_revision", + lambda repo_id, token=None, endpoint=None: "checkpoint-300", + ) + monkeypatch.setattr( + "soup_cli.monitoring.hf_push._download_checkpoint", fake_download + ) + + result = prepare_hf_resume( + repo_id="user/my-model", output_dir=str(out_dir), token="t1" + ) + assert download_called["called"] is False + assert result is not None + assert "checkpoint-300" in result + + def test_prepare_hf_resume_skips_download_when_local_newer( + self, tmp_path, monkeypatch + ): + from soup_cli.monitoring.hf_push import prepare_hf_resume + + monkeypatch.chdir(tmp_path) + out_dir = tmp_path / "runs" + out_dir.mkdir() + (out_dir / "checkpoint-500").mkdir() + + download_called = {"called": False} + + def fake_download(*args, **kwargs): + download_called["called"] = True + return kwargs.get("local_dir", "") + + monkeypatch.setattr( + "soup_cli.monitoring.hf_push.resolve_latest_checkpoint_revision", + lambda repo_id, token=None, endpoint=None: "checkpoint-300", + ) + monkeypatch.setattr( + "soup_cli.monitoring.hf_push._download_checkpoint", fake_download + ) + + result = prepare_hf_resume( + repo_id="user/my-model", output_dir=str(out_dir), token="t1" + ) + assert download_called["called"] is False + assert result is not None + assert "checkpoint-500" in result + + def test_prepare_hf_resume_downloads_when_remote_newer( + self, tmp_path, monkeypatch + ): + from soup_cli.monitoring.hf_push import prepare_hf_resume + + monkeypatch.chdir(tmp_path) + out_dir = tmp_path / "runs" + out_dir.mkdir() + (out_dir / "checkpoint-100").mkdir() + + download_called = {"called": False, "revision": None} + + def fake_download(repo_id, revision, local_dir, token, endpoint): + download_called["called"] = True + download_called["revision"] = revision + return local_dir + + monkeypatch.setattr( + "soup_cli.monitoring.hf_push.resolve_latest_checkpoint_revision", + lambda repo_id, token=None, endpoint=None: "checkpoint-500", + ) + monkeypatch.setattr( + "soup_cli.monitoring.hf_push._download_checkpoint", fake_download + ) + + result = prepare_hf_resume( + repo_id="user/my-model", output_dir=str(out_dir), token="t1" + ) + assert download_called["called"] is True + assert download_called["revision"] == "checkpoint-500" + assert result is not None + + +# --------------------------------------------------------------------------- +# #51 — `soup deploy hf-space --template-dir` +# --------------------------------------------------------------------------- + + +class TestHfSpaceCustomTemplate: + def test_render_custom_template_dir(self, tmp_path, monkeypatch): + from soup_cli.utils.hf_space import render_custom_template_dir + + monkeypatch.chdir(tmp_path) + tdir = tmp_path / "mytpl" + tdir.mkdir() + (tdir / "app.py").write_text("MODEL = '{MODEL_REPO}'\n") + (tdir / "README.md").write_text("# Space for {MODEL_REPO}\n") + (tdir / "requirements.txt").write_text("gradio\n") + + rendered = render_custom_template_dir( + str(tdir), model_repo="user/my-model" + ) + assert rendered["app.py"] == "MODEL = 'user/my-model'\n" + assert rendered["README.md"] == "# Space for user/my-model\n" + assert rendered["requirements.txt"] == "gradio\n" + + def test_render_custom_template_rejects_outside_cwd(self, tmp_path, monkeypatch): + from soup_cli.utils.hf_space import render_custom_template_dir + + cwd = tmp_path / "project" + cwd.mkdir() + monkeypatch.chdir(cwd) + outside = tmp_path / "elsewhere" + outside.mkdir() + (outside / "app.py").write_text("x") + (outside / "README.md").write_text("x") + + with pytest.raises(ValueError, match="under the current"): + render_custom_template_dir(str(outside), model_repo="user/my-model") + + def test_render_custom_template_rejects_invalid_repo_id( + self, tmp_path, monkeypatch + ): + from soup_cli.utils.hf_space import render_custom_template_dir + + monkeypatch.chdir(tmp_path) + tdir = tmp_path / "tpl" + tdir.mkdir() + (tdir / "app.py").write_text("x") + (tdir / "README.md").write_text("x") + + with pytest.raises(ValueError): + render_custom_template_dir(str(tdir), model_repo="bad..repo") + + def test_render_custom_template_requires_app_py(self, tmp_path, monkeypatch): + from soup_cli.utils.hf_space import render_custom_template_dir + + monkeypatch.chdir(tmp_path) + tdir = tmp_path / "tpl" + tdir.mkdir() + (tdir / "README.md").write_text("x") + + with pytest.raises(FileNotFoundError, match="app.py"): + render_custom_template_dir(str(tdir), model_repo="user/my-model") + + def test_render_custom_template_requires_readme(self, tmp_path, monkeypatch): + from soup_cli.utils.hf_space import render_custom_template_dir + + monkeypatch.chdir(tmp_path) + tdir = tmp_path / "tpl" + tdir.mkdir() + (tdir / "app.py").write_text("x") + + with pytest.raises(FileNotFoundError, match="README.md"): + render_custom_template_dir(str(tdir), model_repo="user/my-model") + + def test_render_custom_template_size_cap(self, tmp_path, monkeypatch): + from soup_cli.utils.hf_space import render_custom_template_dir + + monkeypatch.chdir(tmp_path) + tdir = tmp_path / "tpl" + tdir.mkdir() + # 256KB+1 byte + big = "a" * (256 * 1024 + 1) + (tdir / "app.py").write_text(big) + (tdir / "README.md").write_text("x") + + with pytest.raises(ValueError, match="256"): + render_custom_template_dir(str(tdir), model_repo="user/my-model") + + def test_render_custom_template_rejects_symlink(self, tmp_path, monkeypatch): + """Symlinked app.py must be rejected — TOCTOU defence.""" + import os + import sys + + if sys.platform == "win32": + pytest.skip("symlinks require admin on Windows") + + from soup_cli.utils.hf_space import render_custom_template_dir + + monkeypatch.chdir(tmp_path) + tdir = tmp_path / "tpl" + tdir.mkdir() + # Real app.py outside, symlink inside + evil_target = tmp_path / "evil_app.py" + evil_target.write_text("evil\n") + os.symlink(str(evil_target), str(tdir / "app.py")) + (tdir / "README.md").write_text("x") + + with pytest.raises(ValueError, match="symlink"): + render_custom_template_dir(str(tdir), model_repo="user/my-model") + + def test_deploy_hf_space_help_shows_template_dir(self): + result = runner.invoke(app, ["deploy", "hf-space", "--help"]) + assert "--template-dir" in result.output + + +# --------------------------------------------------------------------------- +# #36 — Eval-gate dashboard row +# --------------------------------------------------------------------------- + + +class TestEvalGateDashboardRow: + def test_format_gate_row_disabled(self): + from soup_cli.monitoring.display import format_gate_row + + # No state -> empty string + assert format_gate_row(None) == "" + + def test_format_gate_row_pass(self): + from soup_cli.monitoring.display import format_gate_row + + state = { + "tasks": [ + {"name": "helpfulness", "score": 7.8, "passed": True}, + ], + "overall_passed": True, + "action": None, + } + out = format_gate_row(state) + assert "Gate" in out + assert "helpfulness" in out + assert "7.8" in out + + def test_format_gate_row_regression_stop(self): + from soup_cli.monitoring.display import format_gate_row + + state = { + "tasks": [ + { + "name": "math", + "score": 0.82, + "passed": False, + "delta": -0.06, + "baseline": 0.88, + }, + ], + "overall_passed": False, + "action": "stop", + } + out = format_gate_row(state) + assert "math" in out + assert "0.82" in out + assert "STOP" in out.upper() or "stop" in out + + def test_format_gate_row_warn_action(self): + from soup_cli.monitoring.display import format_gate_row + + state = { + "tasks": [{"name": "t", "score": 0.5, "passed": False, "delta": -0.1}], + "overall_passed": False, + "action": "warn", + } + out = format_gate_row(state) + assert "WARN" in out + + def test_format_gate_row_multi_task(self): + from soup_cli.monitoring.display import format_gate_row + + state = { + "tasks": [ + {"name": "helpfulness", "score": 7.8, "passed": True}, + {"name": "math", "score": 0.82, "passed": False, "delta": -0.06}, + ], + "overall_passed": False, + "action": "stop", + } + out = format_gate_row(state) + assert "helpfulness" in out + assert "math" in out + assert "|" in out # separator + assert "STOP" in out + + def test_format_gate_row_passed_missing_field_renders_neutral(self): + from soup_cli.monitoring.display import format_gate_row + + # `passed` absent — should not silently render as red ✗ + state = { + "tasks": [{"name": "t", "score": 0.5}], + "overall_passed": True, + "action": None, + } + out = format_gate_row(state) + # No action suffix (None), and the task row renders without + # falsely claiming pass. + assert "STOP" not in out + assert "WARN" not in out + + def test_format_gate_row_renders_via_console(self): + from soup_cli.monitoring.display import format_gate_row + + state = { + "tasks": [{"name": "t", "score": 1.0, "passed": True}], + "overall_passed": True, + "action": None, + } + buf = StringIO() + Console(file=buf, force_terminal=False, no_color=True, width=120).print( + format_gate_row(state) + ) + assert "t" in buf.getvalue() diff --git a/tests/test_v0402_part_b.py b/tests/test_v0402_part_b.py new file mode 100644 index 0000000..ed80b69 --- /dev/null +++ b/tests/test_v0402_part_b.py @@ -0,0 +1,245 @@ +"""Tests for v0.40.2 Part B — v0.40.1 carry-overs (H2/H3/N1-G2/N7/M5).""" +from __future__ import annotations + +import json +import re + +import pytest +from typer.testing import CliRunner + +from soup_cli.cli import app + +runner = CliRunner() + + +def _plain(s: str) -> str: + """Strip ANSI escape sequences for help-string assertions.""" + return re.sub(r"\x1b\[[0-9;]*[a-zA-Z]", "", s) + + +# --------------------------------------------------------------------------- +# H2 — data flag aliases +# --------------------------------------------------------------------------- + + +class TestDataFlagAliases: + def test_split_accepts_train_flag(self, tmp_path, monkeypatch): + """--train is accepted on `data split` (informational; train is remainder).""" + monkeypatch.chdir(tmp_path) + ds = tmp_path / "ds.jsonl" + rows = [{"text": f"row {i}"} for i in range(20)] + ds.write_text( + "\n".join(json.dumps(r) for r in rows) + "\n", encoding="utf-8" + ) + + result = runner.invoke( + app, + ["data", "split", str(ds), "--train", "70", "--val", "20", "--test", "10"], + ) + assert result.exit_code == 0, (result.output, repr(result.exception)) + + def test_split_help_mentions_train(self): + result = runner.invoke(app, ["data", "split", "--help"]) + assert "--train" in _plain(result.output) + + def test_filter_min_coherence_alias(self): + result = runner.invoke(app, ["data", "filter", "--help"]) + assert "--min-coherence" in _plain(result.output) + + def test_register_positional_name_path(self, tmp_path, monkeypatch): + monkeypatch.chdir(tmp_path) + ds = tmp_path / "myset.jsonl" + ds.write_text('{"text": "x"}\n', encoding="utf-8") + + result = runner.invoke( + app, + ["data", "register", "myset", str(ds)], + ) + assert result.exit_code == 0, (result.output, repr(result.exception)) + assert "myset" in result.output + + def test_unregister_positional_name(self, tmp_path, monkeypatch): + monkeypatch.chdir(tmp_path) + ds = tmp_path / "myset.jsonl" + ds.write_text('{"text": "x"}\n', encoding="utf-8") + + # Register first + r1 = runner.invoke(app, ["data", "register", "myset", str(ds)]) + assert r1.exit_code == 0 + + r2 = runner.invoke(app, ["data", "unregister", "myset"]) + assert r2.exit_code == 0, (r2.output, repr(r2.exception)) + + +# --------------------------------------------------------------------------- +# H3 — `soup quickstart --output DIR` +# --------------------------------------------------------------------------- + + +class TestQuickstartOutput: + def test_quickstart_help_shows_output(self): + result = runner.invoke(app, ["quickstart", "--help"]) + assert "--output" in _plain(result.output) + + def test_quickstart_output_routes_files(self, tmp_path, monkeypatch): + monkeypatch.chdir(tmp_path) + target = tmp_path / "myrun" + result = runner.invoke( + app, ["quickstart", "--dry-run", "--output", str(target)] + ) + assert result.exit_code == 0, (result.output, repr(result.exception)) + # The actual files quickstart.py writes under --output: + assert (target / "quickstart_data.jsonl").exists() + assert (target / "quickstart_soup.yaml").exists() + + +# --------------------------------------------------------------------------- +# N1 / G2 — `--log-level` deeper plumbing +# --------------------------------------------------------------------------- + + +class TestLogLevelPlumbing: + def test_quiet_emits_less_than_normal(self): + # `data inspect` on missing file emits an error; we look at how many + # log lines are produced. quiet should emit fewer informational lines. + # Use --help as a stable surface. + normal = runner.invoke(app, ["--log-level", "normal", "version"]) + quiet = runner.invoke(app, ["--log-level", "quiet", "version"]) + assert normal.exit_code == 0 + assert quiet.exit_code == 0 + # Both succeed; quiet's output should not exceed normal's by a wide margin. + assert len(quiet.output) <= len(normal.output) + 10 + + def test_log_level_sets_logging_module_level(self, monkeypatch): + import logging as stdlogging + + from soup_cli.utils.log_level import LogLevel, apply_logging_level + + # Reset root logger + root = stdlogging.getLogger() + prev_level = root.level + + try: + apply_logging_level(LogLevel.DEBUG) + assert root.level == stdlogging.DEBUG + apply_logging_level(LogLevel.QUIET) + assert root.level == stdlogging.ERROR + apply_logging_level(LogLevel.VERBOSE) + assert root.level == stdlogging.INFO + apply_logging_level(LogLevel.NORMAL) + assert root.level == stdlogging.WARNING + finally: + root.setLevel(prev_level) + + +# --------------------------------------------------------------------------- +# N7 — `soup infer` accepts HF ids when local path missing +# --------------------------------------------------------------------------- + + +class TestInferHFFallback: + def test_resolve_model_source_local_path_exists(self, tmp_path): + from soup_cli.commands.infer import _resolve_model_source + + model_dir = tmp_path / "mymodel" + model_dir.mkdir() + (model_dir / "config.json").write_text("{}") + + kind, value = _resolve_model_source(str(model_dir)) + assert kind == "local" + assert value == str(model_dir) + + def test_resolve_model_source_hf_id_when_no_local(self, tmp_path, monkeypatch): + from soup_cli.commands.infer import _resolve_model_source + + monkeypatch.chdir(tmp_path) + kind, value = _resolve_model_source("user/my-model") + assert kind == "hf" + assert value == "user/my-model" + + def test_resolve_model_source_invalid_hf_id_when_no_local( + self, tmp_path, monkeypatch + ): + from soup_cli.commands.infer import _resolve_model_source + + monkeypatch.chdir(tmp_path) + # Path-like but doesn't exist and isn't a valid HF id + with pytest.raises(FileNotFoundError, match="not found"): + _resolve_model_source("./nonexistent") + + def test_resolve_model_source_absolute_path_missing(self, tmp_path, monkeypatch): + from soup_cli.commands.infer import _resolve_model_source + + monkeypatch.chdir(tmp_path) + with pytest.raises(FileNotFoundError): + _resolve_model_source("/nonexistent/abs/path") + + def test_resolve_model_source_tilde_path_missing(self, tmp_path, monkeypatch): + from soup_cli.commands.infer import _resolve_model_source + + monkeypatch.chdir(tmp_path) + with pytest.raises(FileNotFoundError): + _resolve_model_source("~/nope/missing") + + def test_resolve_model_source_windows_drive_letter(self, tmp_path, monkeypatch): + from soup_cli.commands.infer import _resolve_model_source + + monkeypatch.chdir(tmp_path) + # Drive-letter syntax always treated as path-like; missing → error. + with pytest.raises(FileNotFoundError): + _resolve_model_source("Z:/nope/missing") + + def test_is_path_like_branches(self): + from soup_cli.commands.infer import _is_path_like + + assert _is_path_like("") is True + assert _is_path_like("./foo") is True + assert _is_path_like("/abs") is True + assert _is_path_like("~/x") is True + assert _is_path_like("\\\\share\\foo") is True + assert _is_path_like("C:/x") is True + assert _is_path_like("user/my-model") is False + assert _is_path_like("microsoft/phi-2") is False + + +# --------------------------------------------------------------------------- +# M5 — `soup runs --cwd-only` +# --------------------------------------------------------------------------- + + +class TestRunsCwdOnly: + def test_runs_help_shows_cwd_only(self): + result = runner.invoke(app, ["runs", "--help"]) + assert "--cwd-only" in _plain(result.output) + + def test_filter_runs_by_cwd(self, tmp_path): + from soup_cli.commands.runs import _filter_runs_by_cwd + + cwd = str(tmp_path.resolve()) + runs = [ + {"run_id": "r1", "output_dir": str(tmp_path / "outA")}, + {"run_id": "r2", "output_dir": "/some/other/place"}, + {"run_id": "r3", "output_dir": str(tmp_path / "nested" / "outB")}, + {"run_id": "r4", "output_dir": None}, + ] + result = _filter_runs_by_cwd(runs, cwd) + ids = [r["run_id"] for r in result] + assert "r1" in ids + assert "r3" in ids + assert "r2" not in ids + assert "r4" not in ids + + def test_filter_runs_by_cwd_cross_drive(self, tmp_path): + """Cross-drive paths on Windows raise ValueError in commonpath; survive.""" + from soup_cli.commands.runs import _filter_runs_by_cwd + + cwd = str(tmp_path.resolve()) + runs = [ + {"run_id": "r1", "output_dir": "Z:\\some\\path"}, + {"run_id": "r2", "output_dir": str(tmp_path / "x")}, + ] + # Should not raise — should drop r1, keep r2 + result = _filter_runs_by_cwd(runs, cwd) + ids = [r["run_id"] for r in result] + assert "r1" not in ids + assert "r2" in ids