Commit Graph

442 Commits

Author SHA1 Message Date
Alpamys c4a4639f84 chore: gitignore .claude/CLAUDE.md (local-only dev instructions)
CLAUDE.md is Claude Code's local project instructions file — it guides the
LLM's behavior during development sessions (conventions, release checklist,
internal Part terminology, test table, etc). Same category as .claude/plan.md,
which is already gitignored.

- Added .claude/CLAUDE.md to .gitignore under the same "Internal plan +
  Claude Code local dev instructions" block as plan.md / settings.json
- git rm --cached to stop tracking (local file preserved)

Rationale: this file has grown to ~620 lines of internal conventions that
don't belong in the public repo — users don't need to see our TDD workflow,
release checklist, review-agent instructions, or "Part X" internal labels.
What users DO need (coding conventions, contrib workflow) is already in
CONTRIBUTING.md.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 22:07:48 +05:00
Alpamys 0034628b03 docs: drop internal Part A/B/C/D/E labels from public docs
"Part A/B/C/D/E" is our internal decomposition (tracked in .claude/plan.md
and referenced in commit messages + GitHub issues). It leaked into
user-facing docs during v0.26.0 release prep. Users don't care about our
internal breakdown — they care about features and versions.

Cleanup:
- README.md: "New in v0.26.0" bullets now describe features by name only,
  (vX.Y.Z) version tags retained where present
- SECURITY.md: v0.26.0 hardening entries grouped by feature name, not Part
- CONTRIBUTING.md: module tree annotations use (v0.26.0) not (v0.26.0 Part X)

.claude/CLAUDE.md: added explicit rule under Release Checklist terminology
stating that Part X labels are internal-only and must NOT appear in public
docs. Prevents the same mistake next release.

.claude/plan.md + commit messages continue to use Part X — that's the
correct venue for internal dev decomposition.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 22:05:47 +05:00
Alpamys 899ad8edf7 test(eval_gate): strip ANSI escapes in train --help CI assertion
`test_train_gate_flag_accepted` asserted `"--gate" in result.output`, but
Typer/Click under CI emits ANSI color codes that split the flag name into
non-contiguous chars: `\x1b[1;36m-\x1b[0m\x1b[1;36m-gate\x1b[0m`. The literal
"--gate" substring is never present. All 9 OS × Python combos failed on the
v0.26.0 Parts B-E push.

Fix: strip ANSI via regex before checking. Also assert on "eval-gated" from
the option description to double-check the flag is wired to its help text.

CI-only / tests-only: no soup_cli/ changes, no version bump needed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 21:51:23 +05:00
Alpamys ddab34115c feat(v0.26.0): Parts B-E — Eval Gate, Trace-to-Pref, Quant-Check, Soup Cans
Closes the v0.26.0 "Red and Blue Ocean" flywheel after Part A (Registry):
Train (eval-gated) -> Registry -> Deploy (quant-check) -> Trace-to-Pref -> Train.

Part B — Eval-Gated Training:
- soup_cli/config/schema.py: EvalGateConfig (enabled/suite/every_n_epochs/
  regression_threshold/baseline/on_regression) + TrainingConfig.eval_gate field
- soup_cli/eval/gate.py: EvalSuite, GateTask, run_gate, resolve_baseline,
  load_suite; baselines from registry:// or file
- soup_cli/monitoring/callback.py: on_epoch_end + _run_eval_gate with fail-safe
  error handling (structured errors treated as regressions under on_regression=stop)
- soup_cli/commands/train.py: --gate <suite.yaml> shortcut flag
- soup_cli/commands/eval.py: gate subcommand (stub generator; live scoring v0.26.1)

Part C — Trace-to-Preference:
- soup_cli/data/traces/: parse_langchain, parse_openai, parse_soup_serve;
  build_pairs from thumbs_up / regenerations / user_edit
- soup_cli/commands/data.py: from-traces + review subcommands
- PII warning panel, 100,000-line cap, path containment, Literal validation

Part D — Quant-Lobotomy Checker:
- soup_cli/eval/quant_check.py: classify_delta (OK/MINOR/MAJOR), run_quant_check,
  resolve_model_ref with artifact kinds filter, table/json/markdown renderers
- soup_cli/commands/eval.py: quant-check subcommand

Part E — Soup Cans:
- soup_cli/cans/: Manifest + DataRef (Pydantic v2); pack_entry + fork_can
  (100MB cap, dunder-key guard); safe tar extraction (filter='data' on py3.12+,
  narrow fallback, manual symlink rejection + commonpath check)
- soup_cli/commands/can.py: pack/inspect/verify/fork subcommands

Shared utility:
- soup_cli/utils/paths.py: single is_under_cwd helper replacing 5 duplicates
  (os.path.realpath + commonpath — Windows 8.3 short-name safe)

Tests: 103 new (29 eval_gate + 24 trace_to_pref + 23 quant_check + 27 cans)
Full suite: 2511 passed on Windows Python 3.10.

Security hardening (review-driven, all severities fixed):
- EvalGateConfig bounds; GateTask null-byte + judge URL scheme allowlist
- Narrow except in _safe_extract so TarError from filter='data' is not swallowed
- resolve_model_ref artifact kinds filter (avoid wrong artifact)
- Manifest.author cap + null/newline rejection; created_at ISO-8601 validation
- fork_can dunder-key + null-byte rejection (prototype pollution prevention)
- fork_can size cap (100MB matches pack_entry)
- inspect_can/read_config refuse paths outside cwd

Docs:
- README.md: v0.26.0 "New in" block (flywheel); 43 recipes; all new commands
  in All Commands list; version examples bumped to 0.26.0; Windows-safe arrows
- CLAUDE.md: architecture + test table + schema + CLI + security section
  extended with B/C/D/E; phase vs Part terminology clarified; release
  checklist step 18 adds Known Limitations section; step 20 adds comment
  template; step 21 adds completeness check via gh issue list --milestone
- SECURITY.md: per-Part security notes (B/C/D/E) under v0.26.0
- CONTRIBUTING.md: test count + directory tree updates

Local smoke: version, eval gate, eval quant-check (table + json),
data from-traces, data review, can pack/inspect/verify/fork — all happy-path
end-to-end. Fixed Unicode arrows (U+2192) in can.py + gate.py that crashed on
Windows CP1252 consoles.

Deferred to v0.26.1 (known limitations, filed as issues post-release):
- eval gate/quant-check live model scoring (stub generator currently)
- data from-traces quality.py judge validation; serve --trace-log collector
- can run + can publish + orchestrator
- eval --attach-to-registry flag; export auto-artifact registration

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 21:37:05 +05:00
Alpamys 4cd4bab969 feat(registry): add Local Model Registry / Provenance Vault (v0.26.0 Part A)
Foundation of v0.26.0 "Red and Blue Ocean" — every fine-tune is now
tracked with lineage, config, eval baseline, and shippable artifacts.

New module soup_cli/registry/:
- hashing.py: deterministic SHA-256 of config (canonical JSON) + data
  (streamed) + base model; used as the entry_hash identity
- store.py: SQLite store (~/.soup/registry.db) with registry_entries,
  registry_artifacts, registry_lineage, registry_tags. Context-manager
  API, cycle-safe BFS walks, AmbiguousRefError on prefix collision,
  LIKE-wildcard-escaped search + resolve, FK ON DELETE CASCADE.
- diff.py: flat-walk ConfigChange diff + per-benchmark eval delta.

New CLI commands:
- soup registry push/list/show/search/diff/promote/delete
- soup history <name> — lineage DAG tree viewer

Security hardening (v0.26.0):
- name/tag validation: alphanumeric + _-. only, null-byte rejected,
  name ≤128, tag ≤64
- artifact path containment via os.path.realpath + commonpath
  (Windows 8.3 short-name safe); enforce_cwd=True default
- SQL parameterised; LIKE wildcards %/_ escaped with ESCAPE '\'
- DB 600 perms on POSIX; SOUP_REGISTRY_DB_PATH env override
- indirect-cycle detection in add_lineage via BFS ancestor walk
- Rich markup escaped in all CLI output
- resolve() raises AmbiguousRefError instead of silent None

Tests: 92 new tests in tests/test_registry.py (hashing, validation,
CRUD, artifacts, lineage + cycle, diff, CLI, history, security,
auto-register integration with ExperimentTracker). Full suite:
2409 passed (was 2313).

All review findings addressed (4 agents: python, code, security, tdd):
HIGH: context manager + try/finally cleanup, FK cascade (removed
manual cascade), cycle detection, LIKE wildcard escaping.
MEDIUM: ambiguous resolve raises, exit 0 on user cancel, cwd
captured at construction, enforce_cwd=True default, Windows
ASCII-safe error messages.

Deferred to v0.26.1: soup eval --attach-to-registry flag and
soup export auto-artifact registration.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-20 20:07:54 +05:00
Alpamys 1408ba744a refactor(bench): strengthen test assertions from PR #31
- happy_path: assert mocked VRAM value (4.00 GB) renders in table
- happy_path: assert 'Benchmarking Configuration' panel rendered
- happy_path: assert mock_generate call count (1 warmup + 3 prompts)
- cpu_warning: assert 'N/A' appears in VRAM column when no CUDA
- Improve docstrings to describe what each test verifies
- Use full exception repr in exit_code asserts for CI debugging
2026-04-19 22:13:35 +05:00
Salil M 543e14d3b2
test(bench): add happy path and cpu warning tests for soup bench (#31) 2026-04-19 22:12:00 +05:00
Alpamys f8a20eea14 refactor(bench): polish prompts-file feature from PR #30
- Narrow broad 'except Exception' to specific exceptions
  (OSError, UnicodeDecodeError, json.JSONDecodeError) + `raise ... from`
- Rename file handle `f` -> `fh` to avoid shadowing (ruff-friendly)
- Clarify comment on --num-prompts ignored-when-file semantics
- Strengthen test: assert actual prompts were passed to _generate
  (not just exit code + output substring)
- Add PEP 8 second blank line between test functions
2026-04-19 15:31:19 +05:00
Salil M 4dd09b132f
FEATURE: add --prompts-file option to bench command for custom test suites (#30)
* feat(bench): add --prompts-file option with path traversal security

* test(bench): add unit tests for custom prompts and path traversal

* docs(bench): document --prompts-file usage in README.md

* feat(bench): add --prompts-file support with path validation

* test(bench): add unit tests for custom prompts and security checks

* style: remove trailing whitespace to pass ruff linting

* test: fix mock patch targets for local imports in bench command

* refactor(bench): simplify prompts-file logic and clean up comment

* test(bench): update assertions to match new prompts-file semantics
2026-04-19 15:24:14 +05:00
Alpamys 8ea99d459f refactor(bench): polish soup bench from PR #25
- Add -> None return type annotation (project convention)
- Replace broad 'except Exception' with specific exceptions
  (OSError, ImportError, RuntimeError, ValueError) + `raise ... from exc`
- Use `_` for unused response variable in tuple unpacking
- Add CPU warning: TPS on CPU is 10-100x slower, misleading users
- Add warmup run (discarded) to avoid CUDA JIT skewing averages
- Document that VRAM scope includes model load (deployment planning)
- Apply ruff style (trailing commas, en-dash -> ASCII, etc.)
2026-04-15 22:06:45 +05:00
Salil M 3c339481d1
Add 'soup bench' command to measure model speed and VRAM usage #24 (#25)
* feat(cli): create 'soup bench' command for inference speed and VRAM measurement

* register 'bench' command into the main CLI router

* add test case for handling missing model paths gracefully

* add 'Inference Benchmarking' section explaining the 'soup bench' tool

* Added soup.yaml

* style: fix linting (unused imports, inconsistent spacing)

* style: sort imports in bench and test_bench to satisfy ruff

* style: final import sort and grouping fix for CI

* Update gitignore
2026-04-15 22:04:16 +05:00
Alpamys 14e86da0e4 a 2026-04-13 14:24:30 +05:00
Alpamys dd04679a2c chore(release): bump to 0.25.1 (Windows py3.9 autopilot fix)
Silent patch release. No behavior changes for the happy path — only
fixes a false-positive path-traversal error in 'soup autopilot' on
Windows + Python 3.9 (commit 670968e) and silences a flaky trl import
on windows-latest CI (commit e44e0bd).

Upgrade: pip install --upgrade soup-cli

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 13:38:22 +05:00
Alpamys 670968e2d5 fix(autopilot): Windows py3.9 path traversal false-positive
test_writes_config fails on windows-latest / Python 3.9 with exit code 1
because the path-traversal check in soup_cli/commands/autopilot.py was:

    data_path = Path(data).resolve()
    data_path.relative_to(Path.cwd().resolve())

On Windows + Python 3.9, Path.resolve() occasionally leaves 8.3 short
names (e.g. "C:\Users\RUNNER~1") in one of the two sides but not the
other, so relative_to raises ValueError even when both paths point to
the same location. GitHub Actions runner home dirs frequently trigger
this (the runneradmin account is created as "runneradmin" but short
names get generated as "RUNNER~1").

Fix: introduce _is_under_cwd(path) helper in soup_cli/commands/autopilot.py
that uses os.path.realpath on both sides (handles 8.3 expansion
consistently) plus os.path.commonpath for the containment check, with
case-insensitive comparison on NT. Apply it to both the --data and
--output path guards. The data_path / output_path locals are then
rebuilt from the realpath result so downstream logic sees the
canonical long-name path.

Also enriches the test assertion to print result.output and
result.exception on failure so future CI breaks are easier to diagnose
without needing to push a debug commit first.

Local verification: all 38 tests in tests/test_autopilot.py pass on
Python 3.10 Windows, full suite 2313 passed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 13:31:43 +05:00
Alpamys e44e0bd663 fix(ci): Windows encoding failure in TestGRPOCPUMinNewTokens
Two tests in tests/test_bugfixes.py::TestGRPOCPUMinNewTokens fail on
windows-latest / Python 3.11 when importing trl.trainer.grpo_trainer:

    RuntimeError: Failed to import trl.trainer.grpo_trainer because of
    the following error:
    'charmap' codec can't decode byte 0x90 in position 6555: character
    maps to <undefined>

Root cause: upstream trl reads an auxiliary file without an explicit
encoding, so Python uses the system default. On Windows that is cp1252
('charmap'), which chokes on non-ASCII bytes present in the file. This
is an upstream issue but Soup needs a green CI.

Two-layer fix:

1. .github/workflows/ci.yml — set PYTHONUTF8=1 and PYTHONIOENCODING=utf-8
   as job-level env. Python's UTF-8 mode makes all file I/O default to
   UTF-8 regardless of locale, which is the correct global fix for this
   class of bug.

2. tests/test_bugfixes.py — add a _trl_grpo_importable() helper that
   returns False on UnicodeDecodeError / ImportError / RuntimeError, and
   use it as a belt-and-braces skip in both TestGRPOCPUMinNewTokens
   tests. Ensures the tests skip cleanly instead of erroring out if a
   future CI change accidentally drops PYTHONUTF8.

Local verification: both tests pass with 'pytest tests/test_bugfixes.py::
TestGRPOCPUMinNewTokens -v' (Python 3.10, Windows).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 13:12:44 +05:00
Alpamys e4c3042a56 feat(v0.25.0): Beyond the Wrapper — 8 major features
Ships v0.25.0 with eight new capabilities (Parts A–H) that close every
competitive gap vs LLaMA-Factory/Axolotl/Unsloth and add unique differentiators:

Part A — 9 new model recipes: Llama 4 Scout (sft/dpo/grpo), Qwen 3 14B/32B/8B-grpo,
Gemma 3 12B/27B-dpo, DeepSeek V3 (MoE LoRA).

Part B — Tool-calling / agentic fine-tuning: new "tool-calling" data format with
detection + normalization, synth data template, init template, eval scoring
(tool_call_match / tool_call_name_match / tool_call_args_subset), plus
qwen3-8b-tools and llama4-scout-tools recipes.

Part C — RLVR (RL from Verifiable Rewards): reward_fn=verifiable routing to
math_verify_reward (regex-only, no eval), code_exec_reward (subprocess sandbox
with RLIMIT_AS/RLIMIT_CPU on POSIX, ephemeral tempdir cwd, concurrency cap,
one-time warning panel), and json_schema_reward. verifiable_domain Literal
validated via model_validator.

Part D — VeRA + OLoRA PEFT methods: LoraConfig.use_vera / use_olora with
mutual-exclusion validator and a unified peft_builder helper that returns
either LoraConfig or VeraConfig with the right init kwargs.

Part E — Apple Silicon MLX backend: detection + hardware profiling in utils/mlx,
MLXSFTTrainerWrapper via mlx-lm, scaffolding DPO/GRPO wrappers rejected at
config load time by SoupConfig._validate_mlx_task_support, lazy trainer
registry, doctor integration, 3 MLX SFT recipes, [mlx] extra in pyproject.

Part F — Data augmentation: soup data augment with rephrase / translate / style
strategies, path-traversal-protected input/output, count capped 1-10, lang/styles
lists bounded (10 entries × 32 chars), rate limiting, and optional --dedup.

Part G — Training intelligence: forgetting detection (ForgettingDetector with
3 built-in mini benchmarks and warning levels) and checkpoint intelligence
(CheckpointTracker with composite metric, early-stop on regression, safe
top-N pruning refusing symlinks and non-checkpoint dirs). SQLite schema
extended with checkpoint_quality + forgetting_eval tables.

Part H — Autopilot: soup autopilot command with dataset/model/hardware
profilers, decision engine (task/quant/peft/batch/lr/epochs/max_length/perf
flags), YAML generator, and full CLI with dry-run + --yes + path-traversal
protection + goal whitelist + gpu_budget bounds [1GB, 1TB]. Bakes forgetting
detection + checkpoint intelligence + early-stop into the generated config.

Totals:
- 2313 tests passing (183 new, up from 2130)
- 86 test files (8 new)
- 43 ready-made recipes (14 new)
- 16 built-in templates (tool-calling added)
- Review findings: all CRITICAL/HIGH/MEDIUM/LOW addressed (3 documented
  design limitations: code_exec best-effort sandbox, prune_checkpoints TOCTOU,
  MLX training integration test requires real hardware)

Docs: CLAUDE.md, README.md, SECURITY.md, CONTRIBUTING.md updated.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 12:58:11 +05:00
Alpamys 5d45b7d2ea fix(docker): correct image name and modernize compose file
- Fix image reference in README and docker-compose.yml
  (soupcli/soup -> ghcr.io/makazhanalpamys/soup) to match
  where the workflow actually publishes
- Remove deprecated 'version: 3.8' from docker-compose.yml
- Use 'docker compose' (v2) instead of legacy 'docker-compose'
- Document in Dockerfile that pip install pulls from PyPI,
  not local source
2026-04-10 23:06:04 +05:00
Salil M 0f5d42f706
Enhance #14 - Add official Docker support for easier onboarding (#20)
* feat: add Dockerfile with CUDA 12.1 and Python 3.11 for Soup

* feat: add docker-compose.yml with GPU passthrough and volume mounts

* chore: add .dockerignore to exclude caching and local data from image

* ci: add GitHub action to build and publish Docker image to GHCR

* docs: add Docker installation and usage instructions to README
2026-04-10 23:04:00 +05:00
Alpamys 19ae61a176 fix: security hardening patch release (v0.24.3)
Rolls up post-review fixes from v0.24.2:
- Validate Last-Event-ID header with isdigit() (prevents ValueError 500)
- Type ChatRequest.messages as list[ChatMessage] (enforces role+content)
- Protect _train_process reads with _train_lock (race condition)
- Narrow form_to_yaml exception to ValueError/TypeError
- Fix Optional[Literal] schema extraction (filter NoneType from args)
- SSRF: use ipaddress.is_loopback instead of string allowlist
- XSS: escapeHtml on all server-supplied innerHTML values
2026-04-07 20:08:24 +05:00
Alpamys 34d8b4f8ba docs: update CLAUDE.md test count to 2130 (post-review fix tests) 2026-04-07 20:02:52 +05:00
Alpamys 0f8a0da222 fix(security): apply escapeHtml to all server-supplied innerHTML injections
Wrap all server-returned values (run IDs, model names, experiment names,
task names, device names, error messages, PIDs, eval results, template
names, recipe names) with escapeHtml() before interpolating into innerHTML
template literals. Prevents stored XSS via malicious model/experiment names.
2026-04-07 19:56:51 +05:00
Alpamys b90f6e0c7a fix(security): harden chat proxy SSRF with ipaddress.is_loopback validation
Replace string-based hostname allowlist with ipaddress.ip_address().is_loopback
to properly handle 127.x.x.x range and IPv6 loopback. Blocks private/link-local
addresses (192.168.x.x, 10.x.x.x, ::ffff:127.0.0.1) for HTTP endpoints.

Adds 2 tests: reject private IP, allow 127.0.0.2 loopback.
2026-04-07 19:28:08 +05:00
Alpamys 69c4ee6a29 fix(ui): address review findings — input validation, race conditions, type safety
- Validate Last-Event-ID header with isdigit() before int() parse
- Type ChatRequest.messages as list[ChatMessage] (role+content model)
- Protect _train_process reads with _train_lock (race condition fix)
- Narrow form_to_yaml exception handling to ValueError/TypeError
- Fix _extract_field_info Optional[Literal] detection (filter NoneType)
2026-04-07 19:22:44 +05:00
Alpamys 54230f7bdd feat(ui): add Web UI Enhancement with live training monitor, enhanced metrics, chat upgrade, and config builder (v0.24.2)
Part A: Training Live Monitor — SSE log streaming (/api/train/logs with
Last-Event-ID reconnection), live metrics SSE (/api/train/metrics/live),
progress endpoint (/api/train/progress), frontend with auto-scroll log
panel, progress bar, and live indicator badge.

Part B: Enhanced Metrics & Eval Display — 2x2 chart grid (loss, LR,
grad_norm, throughput) + GPU memory chart, eval results table in run
detail modal, /api/runs/compare endpoint (max 5 runs).

Part C: Chat Upgrade — /api/chat/send SSE proxy with SSRF protection
(localhost-only HTTP, HTTPS for remote), streaming via ReadableStream,
typing indicator, cancel button, markdown renderer (bold/italic/code),
chat settings panel (temperature/max_tokens/top_p/system prompt/adapter),
chat export as JSON.

Part D: Visual Config Builder — /api/config/schema (Pydantic field
metadata extraction), /api/recipes (29 ready-made configs as JSON),
/api/config/from-form (form values to validated YAML), recipe dropdown.

Security: Chat proxy SSRF validation, max_tokens cap 16384, temperature
0-2, top_p 0-1, Bearer auth on POST, XSS prevention, compare max 5 runs.

Tests: 58 new tests across 4 files (2128 total, 78 files), 67% coverage.
2026-04-07 19:13:55 +05:00
Salil Mhatre d134abb008
Introduce 'soup runs clean' for smart checkpoint space management (#9)
* feat(cli): add 'soup runs clean' intelligent checkpoint cleanup to reclaim disk space

* update README

* feat(cli): add 'soup runs clean' intelligent checkpoint cleanup to reclaim disk space

* fixed whitespace trails

* style(cli): fix lints (line length and spacing) in runs.py

* style: fix all E501 line length lint errors

* fix test mismatch, improve deletion warnings, add path validation, and enforce argument exclusivity

* fix: break long message into multiple lines for Ruff compliance

* test: update runs clean test to use CWD-based output directory for security compliance
2026-04-06 22:35:16 +05:00
Alpamys d748afd7f1 refactor(doctor): clean up _check_resources from PR #7
- Extract _get_ram_gb() helper for testability
- Replace bare 'except Exception' with specific exceptions (OSError, ValueError, etc.)
- Add encoding='utf-8' to /proc/meminfo open (PEP 8)
- Add timeout=5 and check=False to sysctl subprocess call (reliability)
- Fix Linux RAM calc comment and use _GB constant instead of magic number
- Show 'free' suffix on disk space for clarity
2026-04-05 17:31:30 +05:00
Salil Mhatre de5505c7d8
feat(doctor): add RAM and disk space checks to soup doctor command wi… (#7)
* feat(doctor): add RAM and disk space checks to soup doctor command with tests and updated docs

* fix(doctor): resolve subprocess type checker error by manually validating macOS RAM query return code
2026-04-05 17:28:05 +05:00
Alpamys 7ed0b3225e fix: clean up --json flag from PR #6
- Fix trailing whitespace (ruff W293)
- Rename is_json -> json_output for clarity
- Use console.print() instead of bare print() for consistency
2026-04-04 20:46:30 +05:00
Salil 57041cb3c1
add --json flag to version command for machine-readable output in CI/… (#6)
* add --json flag to version command for machine-readable output in CI/scripts and include tests

* docs: update README with soup version --json flag examples
2026-04-04 20:42:46 +05:00
Alpamys 80210a209d fix: cross-platform output path validation tests for CI
- Replace Windows-only paths (C:/Windows/...) with tempdir-based
  paths that work on Linux/macOS CI runners
- Use monkeypatch.chdir for path traversal test isolation
- Fixes test_path_outside_cwd_raises failure on Ubuntu CI
2026-04-04 20:39:17 +05:00
Alpamys 02a2af4b83 fix: v0.24.1 — Windows Unicode fix, AWQ/GPTQ output path traversal
- Replace non-ASCII symbols (checkmarks, arrows, bullets, em-dashes)
  with ASCII equivalents in Rich console output to prevent
  UnicodeEncodeError on Windows without PYTHONIOENCODING=utf-8
- Add _validate_output_path() for AWQ/GPTQ export — output path
  traversal is now checked before import check (previously unreachable
  when autoawq/auto-gptq not installed)
- 4 new tests for output path validation (2065 total, 0 failures)
- Update SECURITY.md with v0.22.0–v0.24.1 hardening history
2026-04-03 23:41:44 +05:00
Alpamys d83dad0a3b docs: update CONTRIBUTING.md for v0.24.0, add CODEOWNERS
- Update test counts to 74 files / 2061 tests (was 62 / 1789)
- Add complete test file table matching CLAUDE.md
- Sync PR checklist with .github/pull_request_template.md
- Add Good First Issues section and New Recipe guide
- Add Conventional Commits format for commit messages
- Add CODEOWNERS for auto-reviewer assignment
2026-04-03 22:22:00 +05:00
Alpamys 1b6b428aaa feat: v0.24.0 — Dataset Hub, Freeze Training, Loss Watchdog, Dataset Registry
Part A: HuggingFace Dataset browser
- soup data search: search HF Hub for datasets (sort by downloads/likes)
- soup data preview: preview remote dataset metadata, splits, features
- soup data download: stream HF dataset to local JSONL (with format conversion)
- Security: trust_remote_code=False, path traversal protection, samples cap at 1M

Part B: Freeze training (like LLaMA-Factory finetuning_type: freeze)
- freeze_layers / freeze_ratio config fields
- soup_cli/utils/freeze.py: detect layers, freeze bottom N
- Wired into SFT trainer before LoRA application
- Supports LLaMA (layers.N) and GPT-2 (h.N) naming

Part C: Loss watchdog (like Axolotl loss_watchdog_threshold)
- loss_watchdog, loss_watchdog_threshold, loss_watchdog_patience config
- Implemented in SoupTrainerCallback with patience counter
- Rich warning panel (stops Live display first), fires only once
- Wired into all 11 trainers via callback kwargs

Part D: Dataset info registry
- soup data register/unregister/registry commands
- ~/.soup/datasets.json local name→path+format mapping
- Name validation, path traversal protection, Rich markup escaping

82 new tests (2061 total), 74 test files.
2026-04-03 16:35:23 +05:00
Alpamys ada4a078b6 fix: v0.23.1 — CI fix, security warnings, expanded test coverage
- Fix macOS CI: CLI help tests use inspect.signature (Rich truncation)
- Security: trust_remote_code warning panels for AWQ/GPTQ export
- Tests: packing trainer mock, curriculum fallback branch, empty list edge case
- 1979 tests across 70 test files
2026-04-03 14:20:21 +05:00
Alpamys 6db403f6c3 fix: CLI help tests use inspect.signature instead of Rich-rendered output
Rich/Typer truncates help panel on narrow terminals (macOS CI), causing
--bits and --group-size flags to not appear in rendered help text. Switch
to inspecting the function signature directly for cross-platform reliability.
2026-04-03 14:08:36 +05:00
Alpamys 50ccf15113 fix: v0.23.0 security — trust_remote_code warning panels for AWQ/GPTQ export 2026-04-03 14:01:16 +05:00
Alpamys f272ee2f4f feat: v0.23.0 — AWQ/GPTQ Export, Sample Packing, Data Split, Curriculum Learning
- AWQ export (`soup export --format awq`) via autoawq, with --bits, --group-size, --calibration-data
- GPTQ export (`soup export --format gptq`) via auto-gptq, with calibration data support
- Sample packing (`packing: true`) for SFT/Pretrain trainers via TRL's native packing
- `soup data split` — train/val/test splitting with random and stratified strategies
- Curriculum learning (`curriculum: true`) — sort dataset by difficulty for staged training
- New utility: soup_cli/utils/curriculum.py (sort_by_length, create_buckets)
- Security: calibration data path traversal protection, bits validation (4/8 only)
- 1970 tests across 70 test files
2026-04-03 13:55:01 +05:00
Alpamys 559203c2e3 fix: v0.22.1 — Python 3.9 compat (str | None → Optional[str]), version bump
The v0.22.0 release broke CI on Python 3.9 because serve.py used
PEP 604 union syntax (str | None) at module level, which requires 3.10+.
Fixed in previous commit; this bumps version to v0.22.1 for a clean PyPI release.
2026-04-03 13:09:16 +05:00
Alpamys f3c2dda9f2 fix: Python 3.9 compat — replace str | None with Optional[str] in serve.py
The `str | None` union syntax at module level requires Python 3.10+.
serve.py cannot use `from __future__ import annotations` because it
defines Pydantic models inside functions (FastAPI needs runtime types).
2026-04-03 13:02:20 +05:00
Alpamys dee9317dde feat: v0.22.0 — Training Profiler, Multi-Adapter Serving, Data Sampling, Adapter Management
New commands:
- `soup profile` — estimate memory, speed, GPU requirements before training
  (--config, --gpu, --json flags)
- `soup adapters list/info/compare` — LoRA adapter management
- `soup data sample` — intelligent dataset sampling (random/diverse/hard strategies)
- `soup serve --adapters` — multi-adapter serving with adapter selection

New files:
- soup_cli/utils/profiler.py — memory/speed estimation engine
- soup_cli/commands/profile.py — profile CLI command
- soup_cli/commands/adapters.py — adapter management CLI

Security:
- Multi-adapter: adapter path traversal protection (resolve + relative_to)
- Multi-adapter: adapter name validation (alphanumeric + hyphens only)
- Multi-adapter: unknown adapter → 404, no adapter name leakage in errors
- Multi-adapter: /v1/adapters returns names only (no filesystem paths)
- Multi-adapter: --adapters rejected for non-transformers backends
- Data sample: output path confinement (resolve + relative_to(cwd))

101 new tests (1890 total), 66 test files, 65.5% coverage, ruff clean.
2026-04-03 12:54:24 +05:00
Alpamys 0a4095eb0b fix: v0.21.1 — Windows UnicodeEncodeError, load_config str, recipe count
- fix: replace Unicode ⚠ (U+26A0) with ASCII [yellow]![/] in migrate
  warnings to prevent UnicodeEncodeError on Windows cp1251/cp866
- fix: load_config() now accepts str in addition to Path
- fix: recipe count in docs corrected from 30 to 29
- chore: bump version to v0.21.1
2026-04-02 14:31:17 +05:00
Alpamys eba63f2387 fix: allow exit code 2 for `soup recipes` no-args help (Typer compat)
Different Typer versions return exit code 0 or 2 for no_args_is_help.
Accept both in the test to fix CI on macOS/Python 3.11.
2026-04-02 14:12:32 +05:00
Alpamys 1b1d679141 feat: v0.21.0 — migrate, recipes, NEFTune, rsLoRA
- `soup migrate` — import configs from LLaMA-Factory, Axolotl, Unsloth
  notebooks (AST-only .ipynb parsing, path traversal protection)
- `soup recipes` — 30 ready-made configs for popular models
  (list/show/use/search with path traversal protection)
- NEFTune (`neftune_alpha`) — noisy embeddings for SFT/DPO/KTO/ORPO/SimPO/IPO
- rsLoRA (`use_rslora`) — rank-stabilized LoRA scaling in all 11 trainers
- Fix: `soup doctor` torchvision circular import crash
- Fix: `load_eval_tasks()` now accepts str in addition to Path
- Security: Rich markup injection prevention in migration warnings
- Security: 10 MB file size limit on migration input files
- 1789 tests, 62 test files, 64% coverage
2026-04-02 14:08:36 +05:00
Alpamys 7f0945c410 chore: bump version to v0.20.2 2026-04-01 18:40:48 +05:00
Alpamys 7aa390b760 fix: restore /static/ prefix for logo path in Web UI 2026-04-01 18:35:59 +05:00
Alpamys d6a7e3f816 chore: bump version to v0.20.1
Bugfix release: ANSI-safe CI test assertions (macOS fix), path
confinement hardening, circular import fix, rate limiting implementation,
trust_remote_code warning, new terracotta logo + Web UI color scheme.
2026-04-01 18:29:48 +05:00
Alpamys 3247dfb1b1 fix: use relative logo path in Web UI, add SVG logo to repo 2026-04-01 18:23:57 +05:00
Alpamys 4164c0ad80 fix: replace SVG logo with PNG (GitHub doesn't render SVG in README) 2026-04-01 18:22:06 +05:00
Alpamys 4ee6968d7e chore: rebrand to new terracotta logo, update Web UI color scheme
Replace purple/cyan cyberpunk theme with warm terracotta palette matching
new SVG logo. Update README to use soup_logo_svg.svg. Update chart colors
in app.js to match new palette (#C0512D primary, #E8975A warm accent).
2026-04-01 18:18:37 +05:00
Alpamys 4affc1a5c7 fix: use ANSI-safe assertions in synth data pro help tests (macOS CI fix) 2026-04-01 18:16:09 +05:00