mirror of https://github.com/razor-ai/soup.git
feat(v0.71.5): ingest/data/prompt/drift polish
Closes #157, #205, #207, #149, #164, #163. Defers #204 (live SaaS pull — paid accounts, infra-blocked, kept open). - #164: get_metric_series falls back to eval_results when metrics is empty - #163: build_verdict confidence biased by advise_history (same project+choice, >=3 precedents); decision never changes - #207: shared utils/webhooks.py (SSRF-hardened) + --slack-url/--discord-url on ingest/prune-prompt/ab/active-sample; ab fires only on a decision - #205: soup prune-prompt --tokenizer (token-prefix detect + decode remainder, boundary-safe) - #149: DynamicCurriculumCallback buckets by loss/perplexity percentile; length keeps round-robin - #157: soup data push/forge --hub modelscope|modelers (data score N/A) 107 new tests in tests/test_v0715.py (12474 -> 12581). ruff clean.
This commit is contained in:
parent
76fdd848cb
commit
1f63393421
45
CHANGELOG.md
45
CHANGELOG.md
|
|
@ -12,6 +12,51 @@ reproducing 70+ versions of notes.
|
|||
|
||||
## [Unreleased]
|
||||
|
||||
## [0.71.5] - 2026-06-02
|
||||
|
||||
### Added
|
||||
- **`soup eval against` now reads eval metrics** — `ExperimentTracker.get_metric_series`
|
||||
falls back to the `eval_results` table when the metric is not a per-step
|
||||
training column (`loss` / `lr` / `grad_norm` / `speed` / `gpu_mem`). So
|
||||
`soup eval against <base> --candidate <run> --metric task_accuracy` returns a
|
||||
real score series (benchmark scores live in `eval_results`, not `metrics`)
|
||||
instead of "Empty series". Per-step columns still read from `metrics` — no
|
||||
regression for existing callers.
|
||||
- **`soup advise` learns from past project outcomes** — `soup advise` now reads
|
||||
this project's accepted-verdict history (`~/.soup/advise_history.jsonl`) and
|
||||
biases the rubric: 3+ successful SFT precedents flip a marginal RAG call to
|
||||
SFT; 3+ negative GRPO outcomes suppress GRPO in favour of SFT-on-traces; an
|
||||
encouraged choice gets a small confidence nudge. Scoped per-project (one
|
||||
project's record never biases another). No history → identical to before.
|
||||
- **Slack/Discord webhooks on four more commands** — `--slack-url` / `--discord-url`
|
||||
(SSRF-hardened, loopback-only HTTP, RFC1918 rejected, never crashes the
|
||||
command) now ship on `soup ingest`, `soup prune-prompt`, `soup ab` (fires only
|
||||
on a `reject_h0` / `accept_h0` decision, not `continue`), and
|
||||
`soup data active-sample` — not just `soup drift-alarm`. The validator + sender
|
||||
moved to a shared `soup_cli/utils/webhooks.py`.
|
||||
- **Tokenizer-aware `soup prune-prompt`** — `--tokenizer <model_or_path>` detects
|
||||
and strips the shared system-prompt prefix on **token** boundaries instead of
|
||||
characters, so a multi-byte UTF-8 prefix can never be split mid-code-point.
|
||||
Default (no `--tokenizer`) keeps the whitespace-character behaviour.
|
||||
- **Curriculum bucketing by loss percentile** — `DynamicCurriculumCallback` now
|
||||
buckets samples by the percentile rank of the live loss (or perplexity)
|
||||
signal within a rolling window when `data.curriculum_metric` is `loss` /
|
||||
`perplexity`, so a consistently-hard sample is routed to the same difficulty
|
||||
bucket across recomputes. `length` and warm-up still use round-robin.
|
||||
- **`--hub` on `soup data push` and `soup data forge`** — `soup data push
|
||||
--hub modelscope|modelers` uploads a dataset via the matching SDK
|
||||
(`repo_type=dataset`, commit message sanitised); `soup data forge --hub
|
||||
<non-hf> --teacher owner/name` pre-fetches the teacher model from that hub
|
||||
(and warns when the teacher is not a repo id so `--hub` is never silently
|
||||
ignored). HF stays the default.
|
||||
|
||||
### Notes
|
||||
- Live SaaS *pull* adapters for `soup ingest` (Langfuse / LangSmith / Helicone /
|
||||
OpenPipe / OpenAI SDKs, issue #204) remain deferred: they need credentialed
|
||||
vendor accounts with populated trace data to validate honestly. Tracked as an
|
||||
open, `infra-blocked` (external-account) item. `soup ingest` continues to parse
|
||||
the JSONL export you pull from your dashboard.
|
||||
|
||||
## [0.71.4] - 2026-06-02
|
||||
|
||||
### Added
|
||||
|
|
|
|||
|
|
@ -120,7 +120,7 @@ src/soup_cli/
|
|||
templates/ - 17 built-in soup.yaml templates (YAML + manifest.json) with load_template loader (v0.39.0, +bco v0.40.0)
|
||||
ui/ - Web UI (FastAPI + HTML/JS SPA)
|
||||
|
||||
tests/ - Test suite (274 files, 12474 tests)
|
||||
tests/ - Test suite (275 files, 12581 tests)
|
||||
examples/ - Real-world config examples and datasets
|
||||
```
|
||||
|
||||
|
|
|
|||
29
README.md
29
README.md
|
|
@ -49,21 +49,22 @@ infrastructure instead of improving models. Soup fixes that.
|
|||
|
||||
## What's New
|
||||
|
||||
**v0.71.4 — Adapter lifecycle + loop wiring.** The merge, PR, and continuous-loop surfaces go live:
|
||||
**v0.71.5 — Ingest, data & prompt polish.** Sharper production-loop ergonomics:
|
||||
|
||||
- **Canary verdict on merge** — `soup adapters merge … --canary suite.json` scores the merged
|
||||
adapter and reports **OK / MINOR / MAJOR**; `--strict-verdict` exits non-zero on a MAJOR
|
||||
regression. Works with no model load using a pre-scored canary suite.
|
||||
- **Evolutionary merge for real** — `soup adapters merge --strategy cmaes --eval suite --budget 1h`
|
||||
now runs the full CMA-ES search (merge → score → optimise) and writes the best blend, instead of
|
||||
just printing a plan.
|
||||
- **Publish an adapter PR** — `soup adapters pr <title> --base-sha <hex> --adapter <path> --push
|
||||
owner/repo#42` posts the rendered PR straight to a GitHub PR comment.
|
||||
- **Continuous fine-tuning loop, wired up** — `soup loop watch --pre-wired` runs the real
|
||||
traces → DPO → eval-gate → canary pipeline; `--pack-cans` snapshots every iteration as a
|
||||
shareable Soup Can with Registry lineage (`soup loop replay <id> --extract dir`).
|
||||
- **Branches ↔ Registry** — `soup adapters branch <name> --attach-to-registry <id>` /
|
||||
`--from-registry <id>` links training-env snapshots into the Registry lineage DAG.
|
||||
- **Alternative model hubs for data** — `soup data push --hub modelscope|modelers` uploads a
|
||||
local JSONL to ModelScope / Modelers, and `soup data forge --hub … --teacher owner/name`
|
||||
pre-fetches the teacher from that hub.
|
||||
- **Tokenizer-aware prompt pruning** — `soup prune-prompt --tokenizer <id-or-path>` finds the
|
||||
shared *token* prefix and decodes only the remainder, so BPE multi-byte sequences never get
|
||||
truncated mid-token the way char-slicing can.
|
||||
- **Webhooks everywhere** — `--slack-url` / `--discord-url` now work on `soup ingest`,
|
||||
`soup prune-prompt`, `soup ab`, and `soup data active-sample` (same SSRF-hardened validator as
|
||||
`soup drift-alarm`). The A/B harness only pings when the sequential test actually decides.
|
||||
- **Curriculum by difficulty percentile** — dynamic curriculum can bucket by `loss` /
|
||||
`perplexity` percentile instead of length round-robin.
|
||||
- **Smarter pre-flight `advise`** — `soup advise` now nudges its confidence using your prior
|
||||
verdicts for the same project, and `soup runs replay` can plot a benchmark-score curve, not
|
||||
just the loss curve.
|
||||
|
||||
Full history: [CHANGELOG.md](CHANGELOG.md) · [GitHub Releases](https://github.com/MakazhanAlpamys/Soup/releases).
|
||||
|
||||
|
|
|
|||
|
|
@ -459,6 +459,11 @@ Every completed run also stores an estimated cost (`$` per run) computed from th
|
|||
captured GPU device name and duration. `soup runs show` renders `—` for CPU /
|
||||
MPS / unknown GPUs (no fabricated zeros).
|
||||
|
||||
As of v0.71.5, the metric-series lookup that powers replay (`ExperimentTracker.get_metric_series`)
|
||||
transparently falls back to the `eval_results` table when a metric has no per-step
|
||||
rows — so you can plot a benchmark-score curve (e.g. `mmlu`, `gsm8k`) the same way
|
||||
you plot `loss`, without caring which table holds the series.
|
||||
|
||||
### Tracker integrations (--tracker mlflow / swanlab / trackio)
|
||||
|
||||
```bash
|
||||
|
|
|
|||
|
|
@ -90,10 +90,12 @@ soup data download user/ds --samples 1000 Stream first 1000 samples
|
|||
soup data register --name my-ds --path d.jsonl --format alpaca Register dataset
|
||||
soup data unregister --name my-ds Remove from registry
|
||||
soup data push --input d.jsonl --hf-dataset user/name Upload local JSONL as HF dataset
|
||||
soup data push --input d.jsonl --hf-dataset u/n --hub modelscope|modelers Upload to an alternative hub
|
||||
soup data registry List all registered datasets
|
||||
soup data demo List bundled demo JSONL fixtures
|
||||
soup data demo alpaca_demo --output ./d.jsonl Copy a bundled demo JSONL fixture
|
||||
soup data forge --docs ./docs --task sft --target-rows 1000 Synthetic data pipeline + provenance
|
||||
soup data forge --docs ./docs --hub modelscope --teacher owner/name Pre-fetch the teacher from an alternative hub
|
||||
soup data score --input rows.jsonl Composite quality scorecard (PII + toxicity + lang + edu)
|
||||
soup data decontaminate --input rows.jsonl --benchmarks mmlu,gsm8k Drop benchmark-overlap rows
|
||||
soup data toxicity --input rows.jsonl -o tox.jsonl Flag toxic rows (keyword baseline)
|
||||
|
|
@ -137,7 +139,7 @@ soup can publish r.can --hf-hub user/name Publish .can to HF Hub as dataset
|
|||
soup runs List training runs
|
||||
soup runs show <run_id> Run details + loss graph + cost
|
||||
soup runs compare <run_1> <run_2> Compare two runs
|
||||
soup runs replay <run_id> Replay summary + loss curve from history
|
||||
soup runs replay <run_id> Replay summary + loss curve from history (also plots a benchmark-score curve when the metric lives in eval_results)
|
||||
soup why [run_id] Explain training anomalies (heuristic)
|
||||
soup tui Full-screen Textual dashboard (requires [tui] extra)
|
||||
soup train --config soup.yaml --profile Record torch.profiler trace to <output>/profiles/
|
||||
|
|
@ -169,8 +171,10 @@ soup edit set --base <m> --method rome|memit|alphaedit --subject "..." --target
|
|||
soup edit diff <before-run> <after-run> --probes p.jsonl Knowledge-injection diff visualizer
|
||||
soup ingest --source langfuse|langsmith|helicone|openpipe|otel|openai-stored --logs <jsonl> Universal trace importer (6 SaaS adapters → normalised JSONL)
|
||||
soup prune-prompt --input <jsonl> --output <jsonl> --min-frequency 0.95 Detect + strip shared system-prompt prefix
|
||||
soup prune-prompt ... --tokenizer <id-or-path> Tokenizer-aware prefix detection (decodes remaining ids, boundary-safe)
|
||||
soup data active-sample --input <jsonl> --output <jsonl> --budget N Top-N uncertain prod traces for human review
|
||||
soup ab --input <jsonl> --metric latency|judge_score|retry_rate mSPRT sequential A/B (decision: continue / reject_h0 / accept_h0)
|
||||
soup ingest|prune-prompt|ab|data active-sample ... --slack-url <https> | --discord-url <https> Shared SSRF-validated webhook on completion
|
||||
soup drift-alarm --reference <jsonl> --live <jsonl> --threshold 0.2 Rolling-KL drift alarm (exit 3 on drift)
|
||||
soup drift-alarm ... --slack-url <https> | --discord-url <https> Optional SSRF-validated webhook on drift detected
|
||||
soup tunability --list List built-in candidate-base catalogue
|
||||
|
|
|
|||
12
docs/data.md
12
docs/data.md
|
|
@ -94,6 +94,14 @@ soup prune-prompt --input traces.jsonl --output pruned.jsonl --min-frequency 0.9
|
|||
|
||||
Binary-search over up-to-32 candidate templates finds the longest qualifying prefix (a longer threshold-meeting prefix may exist beyond the universal one — Soup does not early-exit on the 100% match). Two-pass file read with a 100 000-row DoS cap.
|
||||
|
||||
**Tokenizer-aware mode (v0.71.5).** Pass `--tokenizer <id-or-path>` (a HuggingFace repo id, a local path, or anything `AutoTokenizer.from_pretrained` accepts) to detect the shared prefix in *token* space and decode only the remaining ids:
|
||||
|
||||
```bash
|
||||
soup prune-prompt --input traces.jsonl --output pruned.jsonl --tokenizer Qwen/Qwen2.5-0.5B
|
||||
```
|
||||
|
||||
Char-level stripping can cut a BPE multi-byte sequence in half when the shared prefix ends mid-token; token-aware pruning finds the longest shared *token-id* prefix and decodes the remainder, so the boundary always lands on a real token. Per-row encoding is capped at 50 000 tokens. Omit `--tokenizer` to keep the original character-level behaviour.
|
||||
|
||||
|
||||
## Active-Learning Sampler (`soup data active-sample`)
|
||||
|
||||
|
|
@ -108,6 +116,8 @@ soup data active-sample --input traces.jsonl --output for-review.jsonl --budget
|
|||
|
||||
The output JSONL is a drop-in prompt set for `soup eval human` (v0.19). Budget is bounded `[1, 100 000]`.
|
||||
|
||||
**Webhooks (v0.71.5).** `soup ingest`, `soup prune-prompt`, `soup ab`, and `soup data active-sample` all accept `--slack-url` / `--discord-url` and POST a one-line summary on completion through the same SSRF-hardened validator as `soup drift-alarm` (scheme allowlist, loopback-only HTTP, RFC1918 / link-local / reserved / multicast rejected; the post never raises, so a flaky webhook can't fail the command). `soup ab` only fires when the sequential test actually decides (`reject_h0` / `accept_h0`), not while it's still `continue`-ing.
|
||||
|
||||
|
||||
## Synthetic Data Generation
|
||||
|
||||
|
|
@ -479,6 +489,8 @@ Three tasks supported: `sft` (Q&A pairs), `preference` (chosen/rejected), `tool`
|
|||
|
||||
Document discovery is one level deep over `.txt` / `.md` / `.json` / `.jsonl`; dotfiles + symlinked directories are skipped. All paths are cwd-contained, all writes are atomic via staged-tempfile + `os.replace`, and write targets are rejected if they're symlinks. **Judge providers are live**: `--judge-provider ollama` (localhost-only), `--judge-provider anthropic` (env-only API key), `--judge-provider vllm` (scheme-validated). Per-call judge exceptions logged at DEBUG.
|
||||
|
||||
**Alternative teacher hubs (v0.71.5).** `--hub modelscope|modelers` pre-fetches the `--teacher` from that hub when the teacher is a routable repo id (`owner/name`); `--hub hf` (default) is a no-op and leaves the teacher as a provenance label. If `--hub` is non-HF but `--teacher` is not a repo id (e.g. the default `local-judge`), Soup prints a loud yellow warning rather than silently dropping the flag.
|
||||
|
||||
|
||||
## Data Quality Scorecard
|
||||
|
||||
|
|
|
|||
|
|
@ -82,6 +82,8 @@ soup advise compare
|
|||
4. Task is `factual_lookup` with high output variance → **RAG**.
|
||||
5. Otherwise → **SFT**.
|
||||
|
||||
**Cross-project confidence bias (v0.71.5).** When `~/.soup/advise_history.jsonl` holds ≥3 prior verdicts for the *same choice* in the *same project*, `soup advise` nudges its confidence (not its decision) toward what worked before: a net-positive precedent record (you accepted it AND its recorded outcome was good) bumps confidence up by a small constant; a net-negative one bumps it down. The rubric verdict itself never changes — only how sure Soup is. Verdicts must be `--record`ed for the bias to kick in.
|
||||
|
||||
**Why this command exists.** "Choose fine-tuning vs RAG vs prompt-engineering" is the most-mis-made decision in the space. Reddit, HN, IBM, and Google Cloud all converge on the same advice (start with prompts, escalate to RAG, fine-tune as last resort) and almost everyone ignores it because nobody has the data to prove their case is the exception. Soup `autopilot` picks hyperparameters AFTER you've decided to train; `soup advise` owns the layer above. No trainer library has an incentive to tell users *not to train* — Unsloth's funnel, Axolotl's hosted business, LLaMA-Factory's Alibaba alignment all monetise the training event.
|
||||
|
||||
|
||||
|
|
@ -213,6 +215,8 @@ soup ab --input ab.jsonl --metric judge_score --alpha 0.01 --beta 0.10 --effect-
|
|||
|
||||
Input rows look like `{"arm": "control", "latency": 1.23}` or `{"arm": "treatment", "judge_score": 0.91}`. Decision is one of `continue` (keep collecting samples), `reject_h0` (real difference detected), `accept_h0` (no significant difference). Composes with `soup loop canary` (v0.58) — promote or roll back as soon as the LLR clears a decision boundary.
|
||||
|
||||
`soup ab` accepts `--slack-url` / `--discord-url` (v0.71.5) and pings the webhook **only when the test actually decides** (`reject_h0` / `accept_h0`) — a still-running `continue` stays quiet so you're not paged on every peek. Same SSRF-hardened validator as `soup drift-alarm`.
|
||||
|
||||
|
||||
## Drift Alarm (`soup drift-alarm`)
|
||||
|
||||
|
|
|
|||
|
|
@ -907,12 +907,15 @@ Layer dynamic re-weighting on top of the static `curriculum` bucketer. Every N s
|
|||
training:
|
||||
curriculum: true # static bucketer (v0.23.0)
|
||||
curriculum_buckets: 4
|
||||
curriculum_metric: perplexity # length (default) | loss | perplexity
|
||||
curriculum_dynamic: true # NEW — dynamic re-weighting
|
||||
curriculum_dynamic_recompute_steps: 50 # refresh every 50 global steps
|
||||
curriculum_dynamic_floor: 0.05 # min weight per bucket
|
||||
curriculum_dynamic_temperature: 1.0 # softmax temp on uncertainty
|
||||
```
|
||||
|
||||
**Bucketing by difficulty percentile (v0.71.5).** When `curriculum_metric` is `loss` or `perplexity`, the dynamic callback assigns each step's sample to a bucket by its *rank* within a rolling 512-step window of the difficulty signal (perplexity = `exp(min(loss, 50))`), instead of the round-robin fallback used for `length`. This keeps the buckets calibrated to the live loss distribution rather than a static length sort. `length` (the default) keeps the round-robin assignment.
|
||||
|
||||
Visualise the recorded bucket-weight evolution with `soup runs curriculum-curve <run_id>`.
|
||||
|
||||
DDP / grad-accum safety: multi-rank launches must wire an `all_reduce` hook on per-bucket stats (a cross-validator rejects un-coordinated multi-rank runs upfront). Multi-trainer expansion beyond `sft` / `pretrain` is tracked for v0.48.1.
|
||||
|
|
|
|||
|
|
@ -4,7 +4,7 @@ build-backend = "hatchling.build"
|
|||
|
||||
[project]
|
||||
name = "soup-cli"
|
||||
version = "0.71.4"
|
||||
version = "0.71.5"
|
||||
description = "Fine-tune LLMs in one command. No SSH, no config hell."
|
||||
readme = "README.md"
|
||||
license = "Apache-2.0"
|
||||
|
|
|
|||
|
|
@ -1,3 +1,3 @@
|
|||
"""Soup CLI — Fine-tune LLMs in one command."""
|
||||
|
||||
__version__ = "0.71.4"
|
||||
__version__ = "0.71.5"
|
||||
|
|
|
|||
|
|
@ -0,0 +1,71 @@
|
|||
"""Shared CLI glue for the --slack-url / --discord-url webhook flags (v0.71.5 #207).
|
||||
|
||||
Keeps Typer + Rich Console concerns in the commands layer (``utils/webhooks``
|
||||
stays import-light + framework-free). Used by ``ingest`` / ``prune-prompt`` /
|
||||
``ab`` / ``data active-sample`` so the validate-then-deliver pattern is defined
|
||||
once.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Optional, Tuple
|
||||
|
||||
import typer
|
||||
from rich.console import Console
|
||||
from rich.markup import escape
|
||||
|
||||
|
||||
def validate_webhook_flags(
|
||||
slack_url: Optional[str],
|
||||
discord_url: Optional[str],
|
||||
*,
|
||||
console: Console,
|
||||
) -> Tuple[Optional[str], Optional[str]]:
|
||||
"""Validate webhook URLs at the CLI boundary (``typer.Exit(2)`` on bad).
|
||||
|
||||
Returns the (canonical) URLs. Mirrors the v0.63.0 ``drift-alarm``
|
||||
early-rejection pattern so a typo'd / SSRF-y URL fails fast with a
|
||||
friendly message instead of being silently swallowed at delivery time.
|
||||
"""
|
||||
from soup_cli.utils.webhooks import validate_webhook_url
|
||||
|
||||
out = []
|
||||
for label, value in (("--slack-url", slack_url), ("--discord-url", discord_url)):
|
||||
if value is None:
|
||||
out.append(None)
|
||||
continue
|
||||
try:
|
||||
out.append(validate_webhook_url(value))
|
||||
except (TypeError, ValueError) as exc:
|
||||
console.print(f"[red]{label}: {escape(str(exc))}[/]")
|
||||
raise typer.Exit(2) from exc
|
||||
return out[0], out[1]
|
||||
|
||||
|
||||
def emit_webhooks(
|
||||
slack_url: Optional[str],
|
||||
discord_url: Optional[str],
|
||||
*,
|
||||
payload: dict,
|
||||
console: Console,
|
||||
) -> None:
|
||||
"""POST the completion payload to any configured webhooks (best-effort).
|
||||
|
||||
Never raises (delegates to the never-raising
|
||||
:func:`soup_cli.utils.webhooks.send_webhooks`). Prints a per-target
|
||||
delivered/failed line so the operator sees whether the alert landed.
|
||||
"""
|
||||
if slack_url is None and discord_url is None:
|
||||
return
|
||||
from soup_cli.utils.webhooks import send_webhooks
|
||||
|
||||
for label, ok in send_webhooks(
|
||||
payload, slack_url=slack_url, discord_url=discord_url
|
||||
):
|
||||
colour = "green" if ok else "yellow"
|
||||
console.print(
|
||||
f"[{colour}]{label} webhook: {'delivered' if ok else 'failed'}[/]"
|
||||
)
|
||||
|
||||
|
||||
__all__ = ["emit_webhooks", "validate_webhook_flags"]
|
||||
|
|
@ -2,12 +2,15 @@
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Optional
|
||||
|
||||
import typer
|
||||
from rich.console import Console
|
||||
from rich.markup import escape
|
||||
from rich.panel import Panel
|
||||
from rich.table import Table
|
||||
|
||||
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
|
||||
from soup_cli.utils.ab_test import (
|
||||
MsprtConfig,
|
||||
run_msprt,
|
||||
|
|
@ -35,6 +38,20 @@ def ab(
|
|||
0.1, "--effect-size",
|
||||
help="Minimum detectable difference in means.",
|
||||
),
|
||||
slack_url: Optional[str] = typer.Option(
|
||||
None, "--slack-url",
|
||||
help=(
|
||||
"Optional Slack webhook URL — POSTed on a reject_h0 / accept_h0 "
|
||||
"decision (not on continue). SSRF-validated."
|
||||
),
|
||||
),
|
||||
discord_url: Optional[str] = typer.Option(
|
||||
None, "--discord-url",
|
||||
help=(
|
||||
"Optional Discord webhook URL — POSTed on a reject_h0 / accept_h0 "
|
||||
"decision (not on continue). SSRF-validated."
|
||||
),
|
||||
),
|
||||
) -> None:
|
||||
"""Sequential A/B test with early-stop guarantees (mSPRT)."""
|
||||
try:
|
||||
|
|
@ -43,6 +60,10 @@ def ab(
|
|||
console.print(f"[red]{escape(str(exc))}[/]")
|
||||
raise typer.Exit(2) from exc
|
||||
|
||||
slack_url, discord_url = validate_webhook_flags(
|
||||
slack_url, discord_url, console=console
|
||||
)
|
||||
|
||||
try:
|
||||
cfg = MsprtConfig(
|
||||
metric=canonical, alpha=alpha, beta=beta, effect_size=effect_size,
|
||||
|
|
@ -101,5 +122,24 @@ def ab(
|
|||
)
|
||||
)
|
||||
|
||||
# Webhook only fires on a terminal decision (reject_h0 / accept_h0) —
|
||||
# a `continue` verdict carries no actionable signal (issue #207).
|
||||
if verdict.decision != "continue":
|
||||
emit_webhooks(
|
||||
slack_url,
|
||||
discord_url,
|
||||
payload={
|
||||
"command": "ab",
|
||||
"metric": canonical,
|
||||
"decision": verdict.decision,
|
||||
"log_likelihood_ratio": verdict.log_likelihood_ratio,
|
||||
"n_control": verdict.n_control,
|
||||
"n_treatment": verdict.n_treatment,
|
||||
"mean_control": verdict.mean_control,
|
||||
"mean_treatment": verdict.mean_treatment,
|
||||
},
|
||||
console=console,
|
||||
)
|
||||
|
||||
|
||||
__all__ = ["ab"]
|
||||
|
|
|
|||
|
|
@ -2,11 +2,14 @@
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Optional
|
||||
|
||||
import typer
|
||||
from rich.console import Console
|
||||
from rich.markup import escape
|
||||
from rich.panel import Panel
|
||||
|
||||
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
|
||||
from soup_cli.utils.active_sampler import sample_uncertain_rows, validate_budget
|
||||
|
||||
console = Console()
|
||||
|
|
@ -24,6 +27,14 @@ def active_sample(
|
|||
100, "--budget",
|
||||
help="Max rows to surface for human review (1 - 100_000).",
|
||||
),
|
||||
slack_url: Optional[str] = typer.Option(
|
||||
None, "--slack-url",
|
||||
help="Optional Slack webhook URL — POSTed on completion. SSRF-validated.",
|
||||
),
|
||||
discord_url: Optional[str] = typer.Option(
|
||||
None, "--discord-url",
|
||||
help="Optional Discord webhook URL — POSTed on completion. SSRF-validated.",
|
||||
),
|
||||
) -> None:
|
||||
"""Surface the most uncertain prod traces for human review."""
|
||||
try:
|
||||
|
|
@ -32,6 +43,10 @@ def active_sample(
|
|||
console.print(f"[red]{escape(str(exc))}[/]")
|
||||
raise typer.Exit(2) from exc
|
||||
|
||||
slack_url, discord_url = validate_webhook_flags(
|
||||
slack_url, discord_url, console=console
|
||||
)
|
||||
|
||||
try:
|
||||
plan = sample_uncertain_rows(
|
||||
input_path,
|
||||
|
|
@ -55,5 +70,18 @@ def active_sample(
|
|||
)
|
||||
)
|
||||
|
||||
emit_webhooks(
|
||||
slack_url,
|
||||
discord_url,
|
||||
payload={
|
||||
"command": "active-sample",
|
||||
"rows_selected": plan.rows_selected,
|
||||
"rows_in": plan.rows_in,
|
||||
"mean_uncertainty": plan.mean_uncertainty,
|
||||
"budget": plan.budget,
|
||||
},
|
||||
console=console,
|
||||
)
|
||||
|
||||
|
||||
__all__ = ["active_sample"]
|
||||
|
|
|
|||
|
|
@ -39,6 +39,7 @@ from soup_cli.utils.advise import (
|
|||
synth_probe_lora_delta,
|
||||
)
|
||||
from soup_cli.utils.advise_history import (
|
||||
current_project_name,
|
||||
history_path,
|
||||
load_history,
|
||||
record_verdict,
|
||||
|
|
@ -269,8 +270,27 @@ def advise_run(
|
|||
sft_wall_clock_secs=wall_clock,
|
||||
)
|
||||
|
||||
# v0.71.5 #163 — bias the rubric by this project's past accepted-verdict
|
||||
# outcomes. Best-effort: a missing / unreadable history must never block
|
||||
# a verdict, so any failure falls back to the un-biased rubric.
|
||||
history = None
|
||||
project = None
|
||||
try:
|
||||
verdict = build_verdict(profile, task_category, goal=goal, roi=roi)
|
||||
history = load_history(limit=20)
|
||||
project = current_project_name()
|
||||
except (TypeError, ValueError, OSError):
|
||||
history = None
|
||||
project = None
|
||||
|
||||
try:
|
||||
verdict = build_verdict(
|
||||
profile,
|
||||
task_category,
|
||||
goal=goal,
|
||||
roi=roi,
|
||||
history=history,
|
||||
project=project,
|
||||
)
|
||||
except (TypeError, ValueError) as exc:
|
||||
console.print(f"[red]Verdict build failed:[/] {escape(str(exc))}")
|
||||
raise typer.Exit(1) from exc
|
||||
|
|
|
|||
|
|
@ -1914,16 +1914,32 @@ def push_dataset_cmd(
|
|||
"--message",
|
||||
help="Commit message for the dataset upload",
|
||||
),
|
||||
hub: str = typer.Option(
|
||||
"hf",
|
||||
"--hub",
|
||||
help=(
|
||||
"Target hub: hf (default) / modelscope / modelers. Non-HF hubs "
|
||||
"upload via the matching SDK (v0.71.5 #157) — repo_type=dataset, "
|
||||
"commit message sanitised to first line + 200 chars."
|
||||
),
|
||||
),
|
||||
):
|
||||
"""Upload a local JSONL dataset to HuggingFace Hub as a dataset repo."""
|
||||
"""Upload a local JSONL dataset to a model hub as a dataset repo."""
|
||||
from soup_cli.utils.hf import (
|
||||
get_hf_api,
|
||||
resolve_endpoint,
|
||||
resolve_token,
|
||||
validate_repo_id,
|
||||
)
|
||||
from soup_cli.utils.hubs import validate_hub_name
|
||||
from soup_cli.utils.paths import is_under_cwd
|
||||
|
||||
try:
|
||||
hub_canonical = validate_hub_name(hub)
|
||||
except (TypeError, ValueError) as exc:
|
||||
console.print(f"[red]{exc}[/]")
|
||||
raise typer.Exit(2) from exc
|
||||
|
||||
file_path = Path(input_path)
|
||||
if not file_path.exists():
|
||||
console.print(f"[red]Dataset file not found: {file_path}[/]")
|
||||
|
|
@ -1940,9 +1956,18 @@ def push_dataset_cmd(
|
|||
try:
|
||||
validate_repo_id(hf_dataset)
|
||||
except ValueError as exc:
|
||||
console.print(f"[red]Invalid --hf-dataset repo id:[/] {exc}")
|
||||
console.print(f"[red]Invalid --{hub_canonical}-dataset repo id:[/] {exc}")
|
||||
raise typer.Exit(1) from exc
|
||||
|
||||
# v0.71.5 #157 — non-HF hubs route through the shared upload_repo adapter.
|
||||
# ModelScope / Modelers SDKs upload a folder, so the single JSONL is
|
||||
# staged into a temp dir and uploaded as a dataset repo.
|
||||
if hub_canonical != "hf":
|
||||
_push_dataset_non_hf(
|
||||
hub_canonical, hf_dataset, file_path, commit_message
|
||||
)
|
||||
return
|
||||
|
||||
token = resolve_token()
|
||||
if token is None:
|
||||
console.print(
|
||||
|
|
@ -1987,6 +2012,53 @@ def push_dataset_cmd(
|
|||
)
|
||||
|
||||
|
||||
def _push_dataset_non_hf(
|
||||
hub: str, repo_id: str, file_path: Path, commit_message: str
|
||||
) -> None:
|
||||
"""Upload a single JSONL to a non-HF hub as a dataset (v0.71.5 #157).
|
||||
|
||||
``upload_repo`` uploads a folder, so the file is staged into a temp dir
|
||||
first. ``upload_repo`` already sanitises the commit message (first line +
|
||||
200 chars) and validates the repo id shape.
|
||||
"""
|
||||
import shutil
|
||||
import tempfile
|
||||
|
||||
from rich.markup import escape
|
||||
|
||||
from soup_cli.utils.hubs import upload_repo
|
||||
|
||||
# Stage UNDER cwd — `upload_repo` enforces cwd-containment on folder_path
|
||||
# (the system tempdir would be rejected).
|
||||
staging = tempfile.mkdtemp(prefix=".soup_dataset_push.", dir=os.getcwd())
|
||||
try:
|
||||
shutil.copy2(str(file_path), os.path.join(staging, file_path.name))
|
||||
try:
|
||||
upload_repo(
|
||||
hub,
|
||||
repo_id,
|
||||
folder_path=staging,
|
||||
commit_message=commit_message,
|
||||
repo_type="dataset",
|
||||
)
|
||||
except ImportError as exc:
|
||||
console.print(f"[red]{exc}[/]")
|
||||
raise typer.Exit(1) from exc
|
||||
except (TypeError, ValueError) as exc:
|
||||
console.print(f"[red]Upload failed:[/] {exc}")
|
||||
raise typer.Exit(1) from exc
|
||||
except Exception as exc: # noqa: BLE001 — surface SDK errors generically
|
||||
console.print(f"[red]Upload failed:[/] {exc}")
|
||||
raise typer.Exit(1) from exc
|
||||
finally:
|
||||
shutil.rmtree(staging, ignore_errors=True)
|
||||
|
||||
console.print(
|
||||
f"[green]Uploaded[/] {escape(file_path.name)} to {escape(hub)} "
|
||||
f"dataset [bold]{escape(repo_id)}[/]"
|
||||
)
|
||||
|
||||
|
||||
# --- v0.42.0 Part C / F: AOT preprocess + document ingestion ---------------
|
||||
|
||||
@app.command(name="preprocess")
|
||||
|
|
|
|||
|
|
@ -77,6 +77,14 @@ def forge(
|
|||
"(scheme allowlist + loopback). Ignored for Anthropic."
|
||||
),
|
||||
),
|
||||
hub: str = typer.Option(
|
||||
"hf", "--hub",
|
||||
help=(
|
||||
"Teacher hub: hf (default) / modelscope / modelers. When non-HF "
|
||||
"and --teacher is a repo id (owner/name), the teacher is "
|
||||
"pre-fetched from that hub (v0.71.5 #157)."
|
||||
),
|
||||
),
|
||||
):
|
||||
"""Run the multi-stage synthetic data pipeline with provenance.
|
||||
|
||||
|
|
@ -94,9 +102,45 @@ def forge(
|
|||
write_forge_dataset,
|
||||
write_provenance,
|
||||
)
|
||||
from soup_cli.utils.hubs import validate_hub_name
|
||||
|
||||
try:
|
||||
hub_canonical = validate_hub_name(hub)
|
||||
except (TypeError, ValueError) as exc:
|
||||
console.print(f"[red]{escape(str(exc))}[/]")
|
||||
raise typer.Exit(2) from exc
|
||||
|
||||
effective_teacher = teacher
|
||||
# v0.71.5 #157 — non-HF hub + repo-id teacher → pre-fetch the teacher from
|
||||
# that hub and record the resolved local path in provenance. HF (default)
|
||||
# is a no-op: the teacher stays a provenance label.
|
||||
if hub_canonical != "hf":
|
||||
if "/" in teacher:
|
||||
from soup_cli.utils.hubs import prefetch_model_from_hub
|
||||
|
||||
try:
|
||||
effective_teacher = prefetch_model_from_hub(
|
||||
teacher, hub_canonical, console=console
|
||||
)
|
||||
except ImportError as exc:
|
||||
console.print(f"[red]{escape(str(exc))}[/]")
|
||||
raise typer.Exit(1) from exc
|
||||
except (TypeError, ValueError) as exc:
|
||||
console.print(
|
||||
f"[red]Teacher pre-fetch failed:[/] {escape(str(exc))}"
|
||||
)
|
||||
raise typer.Exit(1) from exc
|
||||
else:
|
||||
# Non-HF hub requested but the teacher is not a routable repo id
|
||||
# (owner/name) — warn loudly instead of silently ignoring --hub
|
||||
# (code-review MEDIUM fix v0.71.5 #157).
|
||||
console.print(
|
||||
f"[yellow]--hub {escape(hub_canonical)} ignored:[/] --teacher "
|
||||
f"{escape(teacher)!r} is not a repo id (owner/name), so there "
|
||||
"is nothing to pre-fetch."
|
||||
)
|
||||
|
||||
judge_fn = _default_judge
|
||||
effective_teacher = teacher
|
||||
if judge_provider is not None:
|
||||
canonical = judge_provider.strip().lower()
|
||||
if canonical not in JUDGE_PROVIDERS:
|
||||
|
|
|
|||
|
|
@ -21,6 +21,7 @@ from rich.console import Console
|
|||
from rich.markup import escape
|
||||
from rich.panel import Panel
|
||||
|
||||
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
|
||||
from soup_cli.utils import ingest_sources as _ingest_sources
|
||||
from soup_cli.utils.ingest_sources import (
|
||||
SUPPORTED_INGEST_SOURCES,
|
||||
|
|
@ -53,6 +54,14 @@ def ingest(
|
|||
"-o",
|
||||
help="Output JSONL (default: traces.jsonl in cwd).",
|
||||
),
|
||||
slack_url: Optional[str] = typer.Option(
|
||||
None, "--slack-url",
|
||||
help="Optional Slack webhook URL — POSTed on completion. SSRF-validated.",
|
||||
),
|
||||
discord_url: Optional[str] = typer.Option(
|
||||
None, "--discord-url",
|
||||
help="Optional Discord webhook URL — POSTed on completion. SSRF-validated.",
|
||||
),
|
||||
) -> None:
|
||||
"""Import production traces from a SaaS observability vendor (v0.63.0).
|
||||
|
||||
|
|
@ -66,6 +75,10 @@ def ingest(
|
|||
console.print(f"[red]{escape(str(exc))}[/]")
|
||||
raise typer.Exit(2) from exc
|
||||
|
||||
slack_url, discord_url = validate_webhook_flags(
|
||||
slack_url, discord_url, console=console
|
||||
)
|
||||
|
||||
if not is_under_cwd(logs):
|
||||
console.print(f"[red]--logs '{escape(logs)}' is outside cwd — refusing[/]")
|
||||
raise typer.Exit(1)
|
||||
|
|
@ -111,6 +124,18 @@ def ingest(
|
|||
f"{escape(output_path.name)}[/]"
|
||||
)
|
||||
|
||||
emit_webhooks(
|
||||
slack_url,
|
||||
discord_url,
|
||||
payload={
|
||||
"command": "ingest",
|
||||
"source": canonical,
|
||||
"traces_written": count,
|
||||
"auth_env_set": auth_value is not None,
|
||||
},
|
||||
console=console,
|
||||
)
|
||||
|
||||
|
||||
def _env_label(source: str) -> str:
|
||||
"""Return the env-var name that authenticates ``source``.
|
||||
|
|
|
|||
|
|
@ -7,11 +7,14 @@ signature trick — shipped OSS for v0.63.0 Part B).
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Optional
|
||||
|
||||
import typer
|
||||
from rich.console import Console
|
||||
from rich.markup import escape
|
||||
from rich.panel import Panel
|
||||
|
||||
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
|
||||
from soup_cli.utils.prune_prompt import prune_traces, validate_min_frequency
|
||||
|
||||
console = Console()
|
||||
|
|
@ -35,6 +38,22 @@ def prune_prompt_cmd(
|
|||
"--min-frequency",
|
||||
help="Prefix must appear in >= this fraction of rows to be stripped (0.0 - 1.0).",
|
||||
),
|
||||
tokenizer: Optional[str] = typer.Option(
|
||||
None,
|
||||
"--tokenizer",
|
||||
help=(
|
||||
"Optional HF tokenizer (model id or local path) for token-aware "
|
||||
"prefix detection. Default: whitespace-character level."
|
||||
),
|
||||
),
|
||||
slack_url: Optional[str] = typer.Option(
|
||||
None, "--slack-url",
|
||||
help="Optional Slack webhook URL — POSTed on completion. SSRF-validated.",
|
||||
),
|
||||
discord_url: Optional[str] = typer.Option(
|
||||
None, "--discord-url",
|
||||
help="Optional Discord webhook URL — POSTed on completion. SSRF-validated.",
|
||||
),
|
||||
) -> None:
|
||||
"""Detect + strip a shared system-prompt prefix (v0.63.0 Part B)."""
|
||||
try:
|
||||
|
|
@ -43,11 +62,16 @@ def prune_prompt_cmd(
|
|||
console.print(f"[red]{escape(str(exc))}[/]")
|
||||
raise typer.Exit(2) from exc
|
||||
|
||||
slack_url, discord_url = validate_webhook_flags(
|
||||
slack_url, discord_url, console=console
|
||||
)
|
||||
|
||||
try:
|
||||
report = prune_traces(
|
||||
input_path,
|
||||
output_path=output_path,
|
||||
min_frequency=min_frequency,
|
||||
tokenizer=tokenizer,
|
||||
)
|
||||
except FileNotFoundError:
|
||||
console.print(f"[red]Input not found: {escape(input_path)}[/]")
|
||||
|
|
@ -56,6 +80,14 @@ def prune_prompt_cmd(
|
|||
console.print(f"[red]{escape(str(exc))}[/]")
|
||||
raise typer.Exit(1) from exc
|
||||
|
||||
payload = {
|
||||
"command": "prune-prompt",
|
||||
"prefix_found": bool(report.prefix),
|
||||
"prefix_chars": report.prefix_chars,
|
||||
"rows_pruned": report.rows_pruned,
|
||||
"rows_total": report.rows_total,
|
||||
}
|
||||
|
||||
if not report.prefix:
|
||||
console.print(
|
||||
Panel(
|
||||
|
|
@ -65,6 +97,7 @@ def prune_prompt_cmd(
|
|||
border_style="yellow",
|
||||
)
|
||||
)
|
||||
emit_webhooks(slack_url, discord_url, payload=payload, console=console)
|
||||
return
|
||||
|
||||
snippet = report.prefix if len(report.prefix) <= 200 else report.prefix[:200] + "..."
|
||||
|
|
@ -77,6 +110,7 @@ def prune_prompt_cmd(
|
|||
border_style="green",
|
||||
)
|
||||
)
|
||||
emit_webhooks(slack_url, discord_url, payload=payload, console=console)
|
||||
|
||||
|
||||
__all__ = ["prune_prompt_cmd"]
|
||||
|
|
|
|||
|
|
@ -318,6 +318,15 @@ class ExperimentTracker:
|
|||
Used by ``soup eval against`` for run-vs-run paired-bootstrap CI.
|
||||
Returns an empty list when the metric does not appear in any row
|
||||
— the caller treats that as "no signal, do not gate".
|
||||
|
||||
v0.71.5 #164: the per-step ``metrics`` table only carries training
|
||||
columns (``loss`` / ``lr`` / ``grad_norm`` / ``speed`` / ``gpu_mem``).
|
||||
Eval metrics like ``task_accuracy`` / ``refusal_rate`` live in the
|
||||
``eval_results`` table instead. So when the per-step pass yields no
|
||||
rows we fall back to the per-benchmark scores in ``eval_results``.
|
||||
Querying ``metrics`` first preserves the established behaviour for
|
||||
every training-loop column (no regression for existing callers);
|
||||
the fallback only fires when the column path is empty.
|
||||
"""
|
||||
if not isinstance(run_id, str) or not run_id:
|
||||
raise ValueError("run_id must be a non-empty string")
|
||||
|
|
@ -335,7 +344,34 @@ class ExperimentTracker:
|
|||
# Skip non-numeric cells silently — same-run inconsistency
|
||||
# is not the caller's problem; they get a shorter series.
|
||||
continue
|
||||
return series
|
||||
if series:
|
||||
return series
|
||||
# Bridge to eval_results (v0.71.5 #164) — benchmark scores for
|
||||
# `soup eval against`. Empty when neither table has data.
|
||||
return self._eval_score_series(run_id, metric)
|
||||
|
||||
def _eval_score_series(self, run_id: str, benchmark: str) -> list[float]:
|
||||
"""Return the per-row ``score`` series from ``eval_results``.
|
||||
|
||||
Ordered by insertion (``id``) for deterministic pairing in the
|
||||
paired-bootstrap CI. Non-numeric cells are skipped silently.
|
||||
"""
|
||||
conn = self._get_conn()
|
||||
rows = conn.execute(
|
||||
"SELECT score FROM eval_results "
|
||||
"WHERE run_id = ? AND benchmark = ? ORDER BY id",
|
||||
(run_id, benchmark),
|
||||
).fetchall()
|
||||
out: list[float] = []
|
||||
for row in rows:
|
||||
value = row["score"]
|
||||
if value is None:
|
||||
continue
|
||||
try:
|
||||
out.append(float(value))
|
||||
except (TypeError, ValueError):
|
||||
continue
|
||||
return out
|
||||
|
||||
def save_eval_result(
|
||||
self,
|
||||
|
|
|
|||
|
|
@ -31,14 +31,17 @@ from __future__ import annotations
|
|||
|
||||
import json
|
||||
import logging
|
||||
import math
|
||||
import os
|
||||
import stat
|
||||
import tempfile
|
||||
from typing import Any, Dict, Optional, Tuple
|
||||
from collections import deque
|
||||
from typing import Any, Deque, Dict, Optional, Tuple
|
||||
|
||||
from soup_cli.utils.curriculum_dynamic import (
|
||||
DynamicCurriculumPolicy,
|
||||
compute_bucket_weights,
|
||||
percentile_bucket,
|
||||
)
|
||||
from soup_cli.utils.paths import is_under_cwd
|
||||
|
||||
|
|
@ -46,6 +49,14 @@ logger = logging.getLogger(__name__)
|
|||
|
||||
_MAX_PATH_LEN = 4096
|
||||
_HISTORY_FILENAME = "curriculum_history.jsonl"
|
||||
# Curriculum metrics that drive percentile bucketing from the live loss
|
||||
# signal (v0.71.5 #149). ``length`` has no per-step signal in HF logs, so it
|
||||
# falls back to round-robin (length bucketing is the static curriculum's
|
||||
# data-prep-time job — see utils/curriculum.py).
|
||||
_PERCENTILE_METRICS = frozenset({"loss", "perplexity"})
|
||||
_VALID_CURRICULUM_METRICS = frozenset({"length", "perplexity", "loss"})
|
||||
# Rolling-window size for the percentile reference distribution.
|
||||
_SIGNAL_WINDOW = 512
|
||||
|
||||
__all__ = [
|
||||
"DynamicCurriculumCallback",
|
||||
|
|
@ -149,14 +160,25 @@ class DynamicCurriculumCallback(_try_import_callback_base()): # type: ignore[mi
|
|||
self,
|
||||
policy: DynamicCurriculumPolicy,
|
||||
output_dir: str,
|
||||
curriculum_metric: str = "length",
|
||||
) -> None:
|
||||
if not isinstance(policy, DynamicCurriculumPolicy):
|
||||
raise TypeError(
|
||||
"policy must be DynamicCurriculumPolicy, got "
|
||||
f"{type(policy).__name__}"
|
||||
)
|
||||
if curriculum_metric not in _VALID_CURRICULUM_METRICS:
|
||||
raise ValueError(
|
||||
"curriculum_metric must be one of "
|
||||
f"{sorted(_VALID_CURRICULUM_METRICS)}, got {curriculum_metric!r}"
|
||||
)
|
||||
self._policy = policy
|
||||
self._curriculum_metric = curriculum_metric
|
||||
self._output_dir = _validate_output_dir(output_dir)
|
||||
# Rolling difficulty-signal window for percentile bucketing (v0.71.5
|
||||
# #149). Persists across recomputes so a consistently-hard sample
|
||||
# keeps landing in the same bucket.
|
||||
self._signal_window: Deque[float] = deque(maxlen=_SIGNAL_WINDOW)
|
||||
# Per-bucket accumulator: bucket_id -> {num_samples, loss_sum, grad_norm_sum}
|
||||
self._stats: Dict[int, Dict[str, float]] = {}
|
||||
# Most recently computed weights (read by external sampler hook).
|
||||
|
|
@ -173,6 +195,10 @@ class DynamicCurriculumCallback(_try_import_callback_base()): # type: ignore[mi
|
|||
def policy(self) -> DynamicCurriculumPolicy:
|
||||
return self._policy
|
||||
|
||||
@property
|
||||
def curriculum_metric(self) -> str:
|
||||
return self._curriculum_metric
|
||||
|
||||
@property
|
||||
def output_dir(self) -> str:
|
||||
return self._output_dir
|
||||
|
|
@ -215,13 +241,55 @@ class DynamicCurriculumCallback(_try_import_callback_base()): # type: ignore[mi
|
|||
# HF emits "loss" + (optionally) "grad_norm" in `logs`.
|
||||
loss = logs.get("loss")
|
||||
grad_norm = logs.get("grad_norm")
|
||||
# Bucket id derived from the global step (round-robin BETA strategy).
|
||||
try:
|
||||
bucket_id = _pick_bucket(global_step, self._policy.num_buckets)
|
||||
except (TypeError, ValueError):
|
||||
return
|
||||
nb = self._policy.num_buckets
|
||||
# v0.71.5 #149: percentile bucketing on the difficulty signal for
|
||||
# loss / perplexity once the rolling window has warmed up; otherwise
|
||||
# round-robin (warm-up + the `length` metric, which has no per-step
|
||||
# signal in HF logs).
|
||||
signal = self._difficulty_signal(loss)
|
||||
bucket_id: Optional[int] = None
|
||||
if (
|
||||
self._curriculum_metric in _PERCENTILE_METRICS
|
||||
and signal is not None
|
||||
and len(self._signal_window) > 0
|
||||
):
|
||||
try:
|
||||
bucket_id = percentile_bucket(
|
||||
signal, list(self._signal_window), nb
|
||||
)
|
||||
except (TypeError, ValueError):
|
||||
bucket_id = None
|
||||
if bucket_id is None:
|
||||
try:
|
||||
bucket_id = _pick_bucket(global_step, nb)
|
||||
except (TypeError, ValueError):
|
||||
return
|
||||
if signal is not None:
|
||||
self._signal_window.append(signal)
|
||||
self._record_sample(bucket_id, loss, grad_norm)
|
||||
|
||||
def _difficulty_signal(self, loss: object) -> Optional[float]:
|
||||
"""Map the logged loss to the configured difficulty signal.
|
||||
|
||||
Returns ``None`` when no usable signal is available (metric is
|
||||
``length`` — no per-step length in HF logs — or the loss is
|
||||
missing / non-finite), which routes the step through the
|
||||
round-robin fallback.
|
||||
"""
|
||||
if self._curriculum_metric not in _PERCENTILE_METRICS:
|
||||
return None
|
||||
try:
|
||||
loss_f = float(loss) if loss is not None else None
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
if loss_f is None or not math.isfinite(loss_f):
|
||||
return None
|
||||
if self._curriculum_metric == "perplexity":
|
||||
# exp is monotonic in loss, so percentile ranks are identical;
|
||||
# clamp the exponent to avoid overflow on a stray loss spike.
|
||||
return math.exp(min(loss_f, 50.0))
|
||||
return loss_f
|
||||
|
||||
def on_step_end(
|
||||
self,
|
||||
args: Any,
|
||||
|
|
|
|||
|
|
@ -26,7 +26,10 @@ import re
|
|||
import stat
|
||||
from collections.abc import Iterable
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Dict, List, Mapping, Optional, Sequence, Tuple
|
||||
from typing import TYPE_CHECKING, Dict, List, Mapping, Optional, Sequence, Tuple
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover — annotation only, avoids circular import.
|
||||
from soup_cli.utils.advise_history import HistoryEntry
|
||||
|
||||
from soup_cli.utils.paths import is_under_cwd
|
||||
|
||||
|
|
@ -59,6 +62,17 @@ _MIN_ROWS_FOR_TRAINING = 50
|
|||
# routed into a GPU-intensive RL run).
|
||||
_MIN_ROWS_FOR_GRPO = 500
|
||||
|
||||
# History-bias thresholds (v0.71.5 #163). A choice is "encouraged" when the
|
||||
# project has >= _HISTORY_MIN_PRECEDENTS accepted verdicts whose mean measured
|
||||
# outcome is >= _HISTORY_POSITIVE_OUTCOME; "discouraged" when the mean is
|
||||
# < _HISTORY_NEGATIVE_OUTCOME over the same minimum count. Encouraged choices
|
||||
# get a small confidence nudge and can flip a marginal tie; discouraged choices
|
||||
# are suppressed (their effective row-floor is raised).
|
||||
_HISTORY_MIN_PRECEDENTS = 3
|
||||
_HISTORY_POSITIVE_OUTCOME = 0.3
|
||||
_HISTORY_NEGATIVE_OUTCOME = 0.0
|
||||
_HISTORY_CONFIDENCE_NUDGE = 0.05
|
||||
|
||||
# Probe defaults — tiny, no GPU required for the heuristic stubs.
|
||||
_PROBE_HOLDOUT_DEFAULT = 100
|
||||
_PROBE_LORA_STEPS_DEFAULT = 100
|
||||
|
|
@ -460,7 +474,7 @@ def _confidence_from_signals(*, row_count: int, diversity: float) -> float:
|
|||
return max(0.2, min(0.95, 0.4 + 0.4 * size_score + 0.2 * diversity))
|
||||
|
||||
|
||||
def build_verdict(
|
||||
def _base_verdict(
|
||||
profile: DatasetProfile,
|
||||
task_category: str,
|
||||
*,
|
||||
|
|
@ -584,6 +598,184 @@ def build_verdict(
|
|||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# History bias (v0.71.5 #163) — tune the rubric from past project outcomes
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _summarise_history_outcomes(
|
||||
history: Optional[Sequence["HistoryEntry"]],
|
||||
*,
|
||||
project: Optional[str] = None,
|
||||
) -> Dict[str, Tuple[float, int]]:
|
||||
"""Aggregate accepted-verdict outcomes per choice from prior history.
|
||||
|
||||
Duck-typed (reads ``.choice`` / ``.accepted`` / ``.outcome`` / ``.project``
|
||||
attributes) so :mod:`advise` never imports :mod:`advise_history` at runtime
|
||||
— that import direction is owned by ``advise_history`` and reversing it
|
||||
would create a cycle.
|
||||
|
||||
Filtering:
|
||||
- only entries whose ``accepted is True`` (a rejected verdict carries no
|
||||
endorsement signal),
|
||||
- only entries with a finite ``outcome`` in ``[-1, 1]`` (bool / None /
|
||||
out-of-range skipped — defensive even though ``HistoryEntry`` validates
|
||||
on construction),
|
||||
- only entries matching ``project`` when supplied (per-project scoping:
|
||||
one project's SFT wins must not bias another project's verdict).
|
||||
|
||||
Returns ``{choice: (mean_outcome, count)}`` for every choice with >= 1
|
||||
qualifying entry.
|
||||
"""
|
||||
if history is None:
|
||||
return {}
|
||||
if isinstance(history, (str, bytes)) or not isinstance(history, Sequence):
|
||||
raise TypeError("history must be a non-string Sequence or None")
|
||||
acc: Dict[str, Tuple[float, int]] = {}
|
||||
for entry in history:
|
||||
choice = getattr(entry, "choice", None)
|
||||
if choice not in CHOICES:
|
||||
continue
|
||||
if getattr(entry, "accepted", None) is not True:
|
||||
continue
|
||||
if project is not None and getattr(entry, "project", None) != project:
|
||||
continue
|
||||
outcome = getattr(entry, "outcome", None)
|
||||
if outcome is None or isinstance(outcome, bool):
|
||||
continue
|
||||
if not isinstance(outcome, (int, float)):
|
||||
continue
|
||||
f_out = float(outcome)
|
||||
if not math.isfinite(f_out) or not (-1.0 <= f_out <= 1.0):
|
||||
continue
|
||||
total, count = acc.get(choice, (0.0, 0))
|
||||
acc[choice] = (total + f_out, count + 1)
|
||||
return {ch: (total / count, count) for ch, (total, count) in acc.items()}
|
||||
|
||||
|
||||
def _is_encouraged(bias: Mapping[str, Tuple[float, int]], choice: str) -> bool:
|
||||
mean_count = bias.get(choice)
|
||||
if mean_count is None:
|
||||
return False
|
||||
mean, count = mean_count
|
||||
return count >= _HISTORY_MIN_PRECEDENTS and mean >= _HISTORY_POSITIVE_OUTCOME
|
||||
|
||||
|
||||
def _is_discouraged(bias: Mapping[str, Tuple[float, int]], choice: str) -> bool:
|
||||
mean_count = bias.get(choice)
|
||||
if mean_count is None:
|
||||
return False
|
||||
mean, count = mean_count
|
||||
return count >= _HISTORY_MIN_PRECEDENTS and mean < _HISTORY_NEGATIVE_OUTCOME
|
||||
|
||||
|
||||
def _precedent_count(bias: Mapping[str, Tuple[float, int]], choice: str) -> int:
|
||||
mean_count = bias.get(choice)
|
||||
return mean_count[1] if mean_count is not None else 0
|
||||
|
||||
|
||||
def build_verdict(
|
||||
profile: DatasetProfile,
|
||||
task_category: str,
|
||||
*,
|
||||
goal: Optional[str] = None,
|
||||
roi: Optional[ROIEstimate] = None,
|
||||
history: Optional[Sequence["HistoryEntry"]] = None,
|
||||
project: Optional[str] = None,
|
||||
) -> Verdict:
|
||||
"""Combine profile + task into a recommendation, optionally history-biased.
|
||||
|
||||
Without ``history`` this is byte-identical to the v0.54.0 rubric
|
||||
(regression-guarded). When ``history`` is supplied the per-project
|
||||
outcome record nudges marginal decisions (v0.71.5 #163):
|
||||
|
||||
- A choice with >= 3 accepted verdicts averaging >= +0.3 outcome is
|
||||
"encouraged": its confidence is nudged up, and it can flip a marginal
|
||||
RAG-vs-SFT tie toward SFT.
|
||||
- A choice with >= 3 accepted verdicts averaging < 0.0 outcome is
|
||||
"discouraged": it is suppressed (e.g. a project that keeps regressing on
|
||||
GRPO falls back to SFT-on-traces).
|
||||
|
||||
The confidence FLOOR is unchanged; only the tie-break + per-choice
|
||||
suppression shift. Per-project scoping is enforced in
|
||||
:func:`_summarise_history_outcomes`.
|
||||
"""
|
||||
base = _base_verdict(profile, task_category, goal=goal, roi=roi)
|
||||
if history is None:
|
||||
return base
|
||||
bias = _summarise_history_outcomes(history, project=project)
|
||||
if not bias:
|
||||
return base
|
||||
return _apply_history_bias(base, bias)
|
||||
|
||||
|
||||
def _apply_history_bias(
|
||||
base: Verdict,
|
||||
bias: Mapping[str, Tuple[float, int]],
|
||||
) -> Verdict:
|
||||
"""Adjust a base verdict using per-project history outcomes."""
|
||||
roi = base.estimated_roi
|
||||
task_category = base.task_category
|
||||
|
||||
# Marginal RAG → SFT flip: strong SFT track record, no comparable RAG
|
||||
# track record. RAG is the marginal call (it fired on a heuristic
|
||||
# variance threshold), so prior SFT success is decisive.
|
||||
if (
|
||||
base.choice == "RAG"
|
||||
and _is_encouraged(bias, "SFT")
|
||||
and not _is_encouraged(bias, "RAG")
|
||||
):
|
||||
n_sft = _precedent_count(bias, "SFT")
|
||||
return Verdict(
|
||||
choice="SFT",
|
||||
confidence=min(0.95, base.confidence + _HISTORY_CONFIDENCE_NUDGE),
|
||||
reason=(
|
||||
f"Base rubric leaned RAG, but {n_sft} prior SFT verdicts in "
|
||||
"this project averaged a positive outcome (precedent) — "
|
||||
"routing to SFT over RAG."
|
||||
),
|
||||
reverse_when=(
|
||||
"the answer space is small and stable and RAG's freshness "
|
||||
"outweighs the historical SFT lift — re-run after measuring."
|
||||
),
|
||||
task_category=task_category,
|
||||
estimated_roi=roi,
|
||||
)
|
||||
|
||||
# Discouraged GRPO: a project that keeps regressing on RL falls back to
|
||||
# SFT-on-traces (raises GRPO's effective floor for this project).
|
||||
if base.choice == "GRPO" and _is_discouraged(bias, "GRPO"):
|
||||
n_grpo = _precedent_count(bias, "GRPO")
|
||||
return Verdict(
|
||||
choice="SFT",
|
||||
confidence=base.confidence,
|
||||
reason=(
|
||||
f"Base rubric leaned GRPO, but {n_grpo} prior GRPO verdicts in "
|
||||
"this project averaged a negative outcome (precedent) — "
|
||||
"falling back to SFT on the reasoning traces."
|
||||
),
|
||||
reverse_when=(
|
||||
"a reliable programmatic reward is now available and the "
|
||||
"earlier GRPO regressions were reward-shaping bugs, not a "
|
||||
"fundamental mismatch."
|
||||
),
|
||||
task_category=task_category,
|
||||
estimated_roi=roi,
|
||||
)
|
||||
|
||||
# No flip — nudge confidence when the chosen path has a positive track
|
||||
# record (DPO keeps its own +0.1; we only ever raise, never lower).
|
||||
if _is_encouraged(bias, base.choice):
|
||||
return Verdict(
|
||||
choice=base.choice,
|
||||
confidence=min(0.95, base.confidence + _HISTORY_CONFIDENCE_NUDGE),
|
||||
reason=base.reason,
|
||||
reverse_when=base.reverse_when,
|
||||
task_category=task_category,
|
||||
estimated_roi=roi,
|
||||
)
|
||||
return base
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Probe runner (Part B) — heuristic stubs; real model loading is opt-in
|
||||
# ---------------------------------------------------------------------------
|
||||
|
|
|
|||
|
|
@ -100,6 +100,16 @@ def _project_name() -> str:
|
|||
return name[:128]
|
||||
|
||||
|
||||
def current_project_name() -> str:
|
||||
"""Public accessor for the current project label (v0.71.5 #163).
|
||||
|
||||
The CLI passes this to ``build_verdict(..., project=...)`` so the history
|
||||
bias is scoped to the project the verdict was recorded under. Returns the
|
||||
same string :func:`record_verdict` stamps into the ``project`` field.
|
||||
"""
|
||||
return _project_name()
|
||||
|
||||
|
||||
def record_verdict(
|
||||
verdict: Verdict,
|
||||
*,
|
||||
|
|
|
|||
|
|
@ -36,6 +36,7 @@ __all__ = [
|
|||
"DynamicCurriculumPolicy",
|
||||
"BucketStats",
|
||||
"compute_bucket_weights",
|
||||
"percentile_bucket",
|
||||
"validate_distributed_curriculum",
|
||||
]
|
||||
|
||||
|
|
@ -179,6 +180,48 @@ def _softmax(values: Sequence[float], temperature: float) -> List[float]:
|
|||
return [e / total for e in exps]
|
||||
|
||||
|
||||
def percentile_bucket(
|
||||
value: float,
|
||||
window: Sequence[float],
|
||||
num_buckets: int,
|
||||
) -> int:
|
||||
"""Bucket ``value`` by its percentile rank within a rolling ``window``.
|
||||
|
||||
v0.71.5 #149 — replaces step-mod round-robin with a difficulty-signal
|
||||
bucketing for ``curriculum_metric in {loss, perplexity}``. A value at or
|
||||
above every window member lands in the top (hardest) bucket; a value
|
||||
below every member lands in bucket 0. Because the bucket is a function of
|
||||
the value's rank (not the step), a consistently-high-loss sample is
|
||||
routed to the same bucket on every recompute.
|
||||
|
||||
Args:
|
||||
value: The current sample's difficulty signal (e.g. loss).
|
||||
window: Recent difficulty signals (rolling reference distribution).
|
||||
An empty / ``None`` window returns bucket 0 (warm-up — the caller
|
||||
should fall back to round-robin until the window fills).
|
||||
num_buckets: Number of difficulty buckets.
|
||||
|
||||
Returns:
|
||||
Bucket id in ``[0, num_buckets - 1]``.
|
||||
"""
|
||||
nb = _reject_bool_int("num_buckets", num_buckets)
|
||||
if nb < _MIN_BUCKETS or nb > _MAX_BUCKETS:
|
||||
raise ValueError(
|
||||
f"num_buckets must be in [{_MIN_BUCKETS}, {_MAX_BUCKETS}], got {nb}"
|
||||
)
|
||||
fv = _reject_bool_float("value", value)
|
||||
if nb == 1:
|
||||
return 0
|
||||
if not window:
|
||||
return 0
|
||||
le = sum(
|
||||
1 for w in window if _reject_bool_float("window value", w) <= fv
|
||||
)
|
||||
rank = le / len(window)
|
||||
bucket = int(rank * nb)
|
||||
return min(nb - 1, max(0, bucket))
|
||||
|
||||
|
||||
def compute_bucket_weights(
|
||||
stats: Mapping[int, Mapping[str, float]],
|
||||
policy: DynamicCurriculumPolicy,
|
||||
|
|
|
|||
|
|
@ -23,20 +23,22 @@ quant-check thresholds; operators can tune via --threshold.
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import ipaddress
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
from dataclasses import dataclass
|
||||
from typing import Iterable, Mapping, Optional, Tuple
|
||||
from urllib.parse import urlparse
|
||||
from typing import Iterable, Mapping, Tuple
|
||||
|
||||
from soup_cli.utils.paths import is_under_cwd
|
||||
|
||||
# Webhook helpers were lifted into the shared ``utils/webhooks`` module in
|
||||
# v0.71.5 #207 so every production-trace command can offer --slack-url /
|
||||
# --discord-url. Re-exported here for back-compat (callers + tests that import
|
||||
# ``validate_webhook_url`` / ``post_webhook`` from ``drift_alarm``).
|
||||
from soup_cli.utils.webhooks import post_webhook, validate_webhook_url
|
||||
|
||||
_MAX_REFERENCE_ROWS = 1_000_000
|
||||
_MAX_WEBHOOK_URL_LEN = 4096
|
||||
_MAX_TEXT_LEN = 1_000_000 # 1 MB / row
|
||||
_LOOPBACK_HOSTS = frozenset({"localhost", "127.0.0.1", "::1"})
|
||||
# Default smoothing constant for the KL kernel (Laplace-style add-epsilon
|
||||
# to defend against `log(0)` when a token in `p` is absent from `q`).
|
||||
_EPS = 1e-9
|
||||
|
|
@ -84,73 +86,6 @@ def validate_threshold(value: object) -> float:
|
|||
return f_val
|
||||
|
||||
|
||||
def _is_private_or_link_local(host: str) -> bool:
|
||||
"""Return True iff ``host`` resolves to a non-loopback private/reserved IP.
|
||||
|
||||
Explicit parentheses on the final clause (code-review MEDIUM fix
|
||||
v0.63.0): Python binds `and` tighter than `or`, but the SSRF gate is
|
||||
safety-critical and a future edit should not need to re-derive the
|
||||
precedence rules to verify the logic.
|
||||
"""
|
||||
try:
|
||||
ip = ipaddress.ip_address(host)
|
||||
except ValueError:
|
||||
return False
|
||||
return (
|
||||
ip.is_private
|
||||
or ip.is_link_local
|
||||
or (ip.is_loopback is False and (ip.is_reserved or ip.is_multicast))
|
||||
)
|
||||
|
||||
|
||||
def validate_webhook_url(url: object) -> str:
|
||||
"""SSRF-hardened webhook URL validator.
|
||||
|
||||
Mirrors v0.29.0 `HF_ENDPOINT` / v0.30.0 OTLP / v0.51.0 `validate_hub_endpoint`
|
||||
policy:
|
||||
- scheme allowlist {http, https}
|
||||
- null-byte / control-char rejection
|
||||
- ``0.0.0.0`` rejected
|
||||
- plain HTTP only permitted for loopback hosts
|
||||
- private / link-local / cloud-metadata IPs rejected
|
||||
"""
|
||||
if isinstance(url, bool):
|
||||
raise TypeError("webhook URL must be str, not bool")
|
||||
if not isinstance(url, str):
|
||||
raise TypeError(f"webhook URL must be str, got {type(url).__name__}")
|
||||
if not url:
|
||||
raise ValueError("webhook URL must be non-empty")
|
||||
if "\x00" in url:
|
||||
raise ValueError("webhook URL must not contain null bytes")
|
||||
if any(ord(c) < 0x20 for c in url):
|
||||
raise ValueError("webhook URL must not contain control characters")
|
||||
if len(url) > _MAX_WEBHOOK_URL_LEN:
|
||||
raise ValueError(f"webhook URL must be <= {_MAX_WEBHOOK_URL_LEN} chars")
|
||||
stripped = url.rstrip("/")
|
||||
parsed = urlparse(stripped)
|
||||
if parsed.scheme not in ("http", "https"):
|
||||
raise ValueError(
|
||||
f"webhook URL must use http/https scheme, got {parsed.scheme!r}"
|
||||
)
|
||||
if not parsed.netloc:
|
||||
raise ValueError("webhook URL is missing a host")
|
||||
host = parsed.hostname or ""
|
||||
if host == "0.0.0.0":
|
||||
raise ValueError(
|
||||
"webhook URL 0.0.0.0 is ambiguous; use 127.0.0.1 or localhost"
|
||||
)
|
||||
if parsed.scheme == "http" and host not in _LOOPBACK_HOSTS:
|
||||
if _is_private_or_link_local(host):
|
||||
raise ValueError(
|
||||
"webhook URL plain HTTP is only allowed for loopback; "
|
||||
"private/link-local hosts require HTTPS"
|
||||
)
|
||||
raise ValueError(
|
||||
"webhook URL for remote hosts must use HTTPS"
|
||||
)
|
||||
return stripped
|
||||
|
||||
|
||||
def compute_token_distribution(rows: Iterable[object]) -> Mapping[str, float]:
|
||||
"""Compute a normalised whitespace-token frequency distribution.
|
||||
|
||||
|
|
@ -305,39 +240,6 @@ def run_drift_check(
|
|||
)
|
||||
|
||||
|
||||
def post_webhook(
|
||||
*,
|
||||
url: Optional[str],
|
||||
payload: Mapping[str, object],
|
||||
timeout_seconds: float = 5.0,
|
||||
) -> bool:
|
||||
"""POST ``payload`` as JSON to ``url``. Returns True on 2xx, False otherwise.
|
||||
|
||||
Never raises — webhook delivery must NOT crash the drift-check run.
|
||||
Lazy-imports ``httpx`` so the runtime cost is paid only when an alarm
|
||||
actually fires.
|
||||
"""
|
||||
if url is None:
|
||||
return False
|
||||
try:
|
||||
validated = validate_webhook_url(url)
|
||||
except (TypeError, ValueError):
|
||||
return False
|
||||
try:
|
||||
import httpx # type: ignore[import-untyped]
|
||||
except ImportError:
|
||||
return False
|
||||
try:
|
||||
response = httpx.post(
|
||||
validated,
|
||||
json=dict(payload),
|
||||
timeout=timeout_seconds,
|
||||
)
|
||||
return 200 <= response.status_code < 300
|
||||
except Exception: # noqa: BLE001 — webhook must never crash drift check
|
||||
return False
|
||||
|
||||
|
||||
__all__ = [
|
||||
"DriftReport",
|
||||
"compute_token_distribution",
|
||||
|
|
|
|||
|
|
@ -113,8 +113,20 @@ def attach_curriculum_callback(
|
|||
getattr(tcfg, "curriculum_dynamic_temperature", 1.0) or 1.0
|
||||
),
|
||||
)
|
||||
# v0.71.5 #149 — thread curriculum_metric so the callback can bucket by
|
||||
# loss / perplexity percentile (round-robin fallback for `length`). Any
|
||||
# value that is not one of the three valid metrics (e.g. a missing field
|
||||
# or a test MagicMock) falls back to `length` so the callback always
|
||||
# constructs.
|
||||
curriculum_metric = getattr(tcfg, "curriculum_metric", "length")
|
||||
if curriculum_metric not in ("length", "perplexity", "loss"):
|
||||
curriculum_metric = "length"
|
||||
try:
|
||||
callback = DynamicCurriculumCallback(policy=policy, output_dir=output_dir)
|
||||
callback = DynamicCurriculumCallback(
|
||||
policy=policy,
|
||||
output_dir=output_dir,
|
||||
curriculum_metric=curriculum_metric,
|
||||
)
|
||||
except (TypeError, ValueError) as exc:
|
||||
logger.debug("attach_curriculum_callback rejected: %s", exc)
|
||||
return False
|
||||
|
|
|
|||
|
|
@ -27,7 +27,7 @@ import json
|
|||
import math
|
||||
import os
|
||||
from dataclasses import dataclass
|
||||
from typing import Sequence
|
||||
from typing import Any, List, Optional, Sequence, Union
|
||||
|
||||
from soup_cli.utils.paths import is_under_cwd
|
||||
|
||||
|
|
@ -35,6 +35,7 @@ from soup_cli.utils.paths import is_under_cwd
|
|||
_MAX_SCAN_ROWS = 100_000
|
||||
_MAX_ROW_CHARS = 1_000_000 # 1 MB / row
|
||||
_MAX_PREFIX_LEN = 100_000 # hard cap on returned prefix length
|
||||
_MAX_TOKENS_PER_ROW = 50_000 # token-aware mode (v0.71.5 #205) per-row cap
|
||||
|
||||
# Tunable: a frequency below this is meaningless (we want a *near-universal*
|
||||
# prefix). Operator can pick anything in [0, 1] via --min-frequency.
|
||||
|
|
@ -163,11 +164,112 @@ def detect_common_prefix(
|
|||
return best_prefix
|
||||
|
||||
|
||||
def detect_common_prefix_tokens(
|
||||
token_rows: Sequence[Sequence[int]],
|
||||
*,
|
||||
min_frequency: float,
|
||||
) -> List[int]:
|
||||
"""Token-level analogue of :func:`detect_common_prefix` (v0.71.5 #205).
|
||||
|
||||
Returns the longest token-id prefix shared by >= ``min_frequency`` of
|
||||
rows. Operating on token IDs (not characters) guarantees the prefix
|
||||
always ends on a token boundary — a multi-byte UTF-8 sequence can never
|
||||
be split mid-code-point.
|
||||
"""
|
||||
threshold = validate_min_frequency(min_frequency)
|
||||
if isinstance(token_rows, (str, bytes)) or not hasattr(token_rows, "__iter__"):
|
||||
raise TypeError(
|
||||
f"token_rows must be an iterable of int sequences, "
|
||||
f"got {type(token_rows).__name__}"
|
||||
)
|
||||
|
||||
materialised: List[List[int]] = []
|
||||
for idx, row in enumerate(token_rows):
|
||||
if isinstance(row, (str, bytes)) or not hasattr(row, "__iter__"):
|
||||
raise TypeError(
|
||||
f"token_rows[{idx}] must be a sequence of ints, "
|
||||
f"got {type(row).__name__}"
|
||||
)
|
||||
materialised.append(list(row)[:_MAX_TOKENS_PER_ROW])
|
||||
if len(materialised) >= _MAX_SCAN_ROWS:
|
||||
break
|
||||
|
||||
if not materialised:
|
||||
return []
|
||||
if len(materialised) == 1:
|
||||
return list(materialised[0]) if threshold >= 1.0 else []
|
||||
|
||||
need = max(1, int(math.ceil(threshold * len(materialised))))
|
||||
best_prefix: List[int] = []
|
||||
sample_templates = materialised[: min(32, len(materialised))]
|
||||
for template in sample_templates:
|
||||
lo, hi = 0, len(template)
|
||||
best_len = 0
|
||||
while lo <= hi:
|
||||
mid = (lo + hi) // 2
|
||||
if mid == 0:
|
||||
lo = mid + 1
|
||||
continue
|
||||
pfx = template[:mid]
|
||||
count = sum(1 for r in materialised if r[:mid] == pfx)
|
||||
if count >= need:
|
||||
best_len = mid
|
||||
lo = mid + 1
|
||||
else:
|
||||
hi = mid - 1
|
||||
if best_len > len(best_prefix):
|
||||
best_prefix = template[:best_len]
|
||||
return best_prefix
|
||||
|
||||
|
||||
def _resolve_tokenizer(tokenizer: Union[str, Any]) -> Any:
|
||||
"""Return a tokenizer object from a name (lazy AutoTokenizer) or object.
|
||||
|
||||
A pre-built tokenizer-like object (duck-typed ``encode`` / ``decode``)
|
||||
is returned as-is — this is the injectable test seam + lets advanced
|
||||
callers pass an already-loaded tokenizer. A string is treated as an HF
|
||||
model id / local path and lazy-loaded via ``transformers.AutoTokenizer``
|
||||
(so importing this module never pulls transformers).
|
||||
"""
|
||||
if hasattr(tokenizer, "encode") and hasattr(tokenizer, "decode"):
|
||||
return tokenizer
|
||||
if not isinstance(tokenizer, str):
|
||||
raise TypeError(
|
||||
"tokenizer must be a model id / path string or a tokenizer object"
|
||||
)
|
||||
if not tokenizer:
|
||||
raise ValueError("tokenizer name must be non-empty")
|
||||
try:
|
||||
from transformers import AutoTokenizer # noqa: PLC0415
|
||||
except ImportError as exc:
|
||||
raise ValueError(
|
||||
"tokenizer-aware prune-prompt needs transformers — "
|
||||
"install with: pip install 'soup-cli[train]'"
|
||||
) from exc
|
||||
try:
|
||||
return AutoTokenizer.from_pretrained(tokenizer)
|
||||
except Exception as exc: # noqa: BLE001 — surface a friendly message.
|
||||
raise ValueError(
|
||||
f"could not load tokenizer {tokenizer!r}: {type(exc).__name__}: {exc}"
|
||||
) from exc
|
||||
|
||||
|
||||
def _encode(tok: Any, text: str) -> List[int]:
|
||||
"""Encode ``text`` to token IDs (no special tokens), capped per-row."""
|
||||
try:
|
||||
ids = tok.encode(text, add_special_tokens=False)
|
||||
except TypeError:
|
||||
# Tokenizers / fakes without the kwarg.
|
||||
ids = tok.encode(text)
|
||||
return list(ids)[:_MAX_TOKENS_PER_ROW]
|
||||
|
||||
|
||||
def prune_traces(
|
||||
input_path: str,
|
||||
*,
|
||||
output_path: str,
|
||||
min_frequency: float = _DEFAULT_MIN_FREQUENCY,
|
||||
tokenizer: Optional[Union[str, Any]] = None,
|
||||
) -> PrunePromptReport:
|
||||
"""Read a JSONL of {prompt, output} rows, strip shared prefix, write.
|
||||
|
||||
|
|
@ -176,6 +278,12 @@ def prune_traces(
|
|||
``prompt`` field (other fields untouched). When no prefix clears the
|
||||
threshold, the output is byte-identical to the input plus a
|
||||
``rows_pruned=0`` report.
|
||||
|
||||
v0.71.5 #205: pass ``tokenizer`` (an HF model id / local path string, or
|
||||
a pre-built tokenizer object) to detect + strip the prefix on token
|
||||
boundaries instead of characters — the prefix can then never end
|
||||
mid-UTF-8-code-point. Default (``None``) keeps the whitespace-character
|
||||
behaviour.
|
||||
"""
|
||||
threshold = validate_min_frequency(min_frequency)
|
||||
|
||||
|
|
@ -198,6 +306,10 @@ def prune_traces(
|
|||
if not os.path.isfile(input_path):
|
||||
raise FileNotFoundError(input_path)
|
||||
|
||||
# Resolve the tokenizer up front (fails fast on a bad name even on an
|
||||
# empty input file) — None keeps the legacy character path.
|
||||
tok = _resolve_tokenizer(tokenizer) if tokenizer is not None else None
|
||||
|
||||
# First pass: collect prompts (capped).
|
||||
prompts: list[str] = []
|
||||
rows_total = 0
|
||||
|
|
@ -233,9 +345,24 @@ def prune_traces(
|
|||
min_frequency=threshold,
|
||||
)
|
||||
|
||||
prefix = detect_common_prefix(prompts, min_frequency=threshold)
|
||||
if tok is None:
|
||||
return _prune_char_level(
|
||||
input_path, output_path, prompts, rows_total, threshold
|
||||
)
|
||||
return _prune_token_level(
|
||||
tok, input_path, output_path, prompts, rows_total, threshold
|
||||
)
|
||||
|
||||
# Second pass: write output with prefix stripped where applicable.
|
||||
|
||||
def _prune_char_level(
|
||||
input_path: str,
|
||||
output_path: str,
|
||||
prompts: List[str],
|
||||
rows_total: int,
|
||||
threshold: float,
|
||||
) -> PrunePromptReport:
|
||||
"""Character-level prefix strip (the v0.63.0 default behaviour)."""
|
||||
prefix = detect_common_prefix(prompts, min_frequency=threshold)
|
||||
rows_pruned = 0
|
||||
with open(input_path, encoding="utf-8") as fh_in, \
|
||||
open(output_path, "w", encoding="utf-8") as fh_out:
|
||||
|
|
@ -253,7 +380,6 @@ def prune_traces(
|
|||
row["prompt"] = row["prompt"][len(prefix):]
|
||||
rows_pruned += 1
|
||||
fh_out.write(json.dumps(row, ensure_ascii=False) + "\n")
|
||||
|
||||
return PrunePromptReport(
|
||||
prefix=prefix,
|
||||
prefix_chars=len(prefix),
|
||||
|
|
@ -263,9 +389,58 @@ def prune_traces(
|
|||
)
|
||||
|
||||
|
||||
def _prune_token_level(
|
||||
tok: Any,
|
||||
input_path: str,
|
||||
output_path: str,
|
||||
prompts: List[str],
|
||||
rows_total: int,
|
||||
threshold: float,
|
||||
) -> PrunePromptReport:
|
||||
"""Token-aware prefix strip (v0.71.5 #205).
|
||||
|
||||
The detected prefix is a list of token IDs; stripping a row decodes the
|
||||
REMAINING token IDs so the boundary is always a clean token break.
|
||||
"""
|
||||
token_rows = [_encode(tok, p) for p in prompts]
|
||||
prefix_ids = detect_common_prefix_tokens(token_rows, min_frequency=threshold)
|
||||
prefix_text = tok.decode(prefix_ids) if prefix_ids else ""
|
||||
plen = len(prefix_ids)
|
||||
|
||||
rows_pruned = 0
|
||||
with open(input_path, encoding="utf-8") as fh_in, \
|
||||
open(output_path, "w", encoding="utf-8") as fh_out:
|
||||
for line in fh_in:
|
||||
line = line.strip()
|
||||
if not line:
|
||||
continue
|
||||
try:
|
||||
row = json.loads(line)
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
if not isinstance(row, dict):
|
||||
continue
|
||||
prompt = row.get("prompt")
|
||||
if prefix_ids and isinstance(prompt, str):
|
||||
ids = _encode(tok, prompt)
|
||||
if ids[:plen] == prefix_ids:
|
||||
row["prompt"] = tok.decode(ids[plen:])
|
||||
rows_pruned += 1
|
||||
fh_out.write(json.dumps(row, ensure_ascii=False) + "\n")
|
||||
|
||||
return PrunePromptReport(
|
||||
prefix=prefix_text,
|
||||
prefix_chars=len(prefix_text),
|
||||
rows_total=rows_total,
|
||||
rows_pruned=rows_pruned,
|
||||
min_frequency=threshold,
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"PrunePromptReport",
|
||||
"detect_common_prefix",
|
||||
"detect_common_prefix_tokens",
|
||||
"prune_traces",
|
||||
"validate_min_frequency",
|
||||
]
|
||||
|
|
|
|||
|
|
@ -0,0 +1,152 @@
|
|||
"""Shared Slack / Discord webhook helpers (v0.71.5 #207).
|
||||
|
||||
Lifts the SSRF-hardened ``validate_webhook_url`` + best-effort ``post_webhook``
|
||||
out of ``utils/drift_alarm.py`` (v0.63.0 Part E) so every production-trace
|
||||
command can offer ``--slack-url`` / ``--discord-url`` without re-implementing
|
||||
the SSRF gate. ``drift_alarm`` now re-exports these for back-compat.
|
||||
|
||||
SSRF policy — full parity with v0.29.0 ``HF_ENDPOINT`` / v0.30.0 OTLP /
|
||||
v0.51.0 ``validate_hub_endpoint`` / v0.63.0 drift-alarm:
|
||||
- scheme allowlist {http, https}
|
||||
- null-byte / control-char rejection
|
||||
- ``0.0.0.0`` rejected
|
||||
- plain HTTP only permitted for loopback hosts
|
||||
- private / link-local / reserved / multicast IPs rejected
|
||||
|
||||
``post_webhook`` NEVER raises — webhook delivery must not crash the command
|
||||
that triggered it. ``httpx`` is lazy-imported so the runtime cost is paid only
|
||||
when an alarm actually fires.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import ipaddress
|
||||
from typing import List, Mapping, Optional, Tuple
|
||||
from urllib.parse import urlparse
|
||||
|
||||
_MAX_WEBHOOK_URL_LEN = 4096
|
||||
_LOOPBACK_HOSTS = frozenset({"localhost", "127.0.0.1", "::1"})
|
||||
|
||||
|
||||
def _is_private_or_link_local(host: str) -> bool:
|
||||
"""Return True iff ``host`` resolves to a non-loopback private/reserved IP.
|
||||
|
||||
Explicit parentheses on the final clause (mirrors v0.63.0 drift-alarm
|
||||
code-review MEDIUM fix): Python binds ``and`` tighter than ``or``, but
|
||||
the SSRF gate is safety-critical and a future edit should not need to
|
||||
re-derive the precedence rules to verify the logic.
|
||||
"""
|
||||
try:
|
||||
ip = ipaddress.ip_address(host)
|
||||
except ValueError:
|
||||
return False
|
||||
return (
|
||||
ip.is_private
|
||||
or ip.is_link_local
|
||||
or (ip.is_loopback is False and (ip.is_reserved or ip.is_multicast))
|
||||
)
|
||||
|
||||
|
||||
def validate_webhook_url(url: object) -> str:
|
||||
"""SSRF-hardened webhook URL validator (returns the canonical URL)."""
|
||||
if isinstance(url, bool):
|
||||
raise TypeError("webhook URL must be str, not bool")
|
||||
if not isinstance(url, str):
|
||||
raise TypeError(f"webhook URL must be str, got {type(url).__name__}")
|
||||
if not url:
|
||||
raise ValueError("webhook URL must be non-empty")
|
||||
if "\x00" in url:
|
||||
raise ValueError("webhook URL must not contain null bytes")
|
||||
if any(ord(c) < 0x20 for c in url):
|
||||
raise ValueError("webhook URL must not contain control characters")
|
||||
if len(url) > _MAX_WEBHOOK_URL_LEN:
|
||||
raise ValueError(f"webhook URL must be <= {_MAX_WEBHOOK_URL_LEN} chars")
|
||||
stripped = url.rstrip("/")
|
||||
parsed = urlparse(stripped)
|
||||
if parsed.scheme not in ("http", "https"):
|
||||
raise ValueError(
|
||||
f"webhook URL must use http/https scheme, got {parsed.scheme!r}"
|
||||
)
|
||||
if not parsed.netloc:
|
||||
raise ValueError("webhook URL is missing a host")
|
||||
host = parsed.hostname or ""
|
||||
if host == "0.0.0.0":
|
||||
raise ValueError(
|
||||
"webhook URL 0.0.0.0 is ambiguous; use 127.0.0.1 or localhost"
|
||||
)
|
||||
if parsed.scheme == "http" and host not in _LOOPBACK_HOSTS:
|
||||
if _is_private_or_link_local(host):
|
||||
raise ValueError(
|
||||
"webhook URL plain HTTP is only allowed for loopback; "
|
||||
"private/link-local hosts require HTTPS"
|
||||
)
|
||||
raise ValueError("webhook URL for remote hosts must use HTTPS")
|
||||
return stripped
|
||||
|
||||
|
||||
def post_webhook(
|
||||
*,
|
||||
url: Optional[str],
|
||||
payload: Mapping[str, object],
|
||||
timeout_seconds: float = 5.0,
|
||||
) -> bool:
|
||||
"""POST ``payload`` as JSON to ``url``. Returns True on 2xx, False otherwise.
|
||||
|
||||
Never raises — webhook delivery must NOT crash the calling command.
|
||||
"""
|
||||
if url is None:
|
||||
return False
|
||||
try:
|
||||
validated = validate_webhook_url(url)
|
||||
except (TypeError, ValueError):
|
||||
return False
|
||||
try:
|
||||
import httpx # type: ignore[import-untyped]
|
||||
except ImportError:
|
||||
return False
|
||||
try:
|
||||
response = httpx.post(
|
||||
validated,
|
||||
json=dict(payload),
|
||||
timeout=timeout_seconds,
|
||||
)
|
||||
return 200 <= response.status_code < 300
|
||||
except Exception: # noqa: BLE001 — webhook must never crash the command
|
||||
return False
|
||||
|
||||
|
||||
def send_webhooks(
|
||||
payload: Mapping[str, object],
|
||||
*,
|
||||
slack_url: Optional[str] = None,
|
||||
discord_url: Optional[str] = None,
|
||||
timeout_seconds: float = 5.0,
|
||||
) -> List[Tuple[str, bool]]:
|
||||
"""POST ``payload`` to each provided webhook; return per-target delivery.
|
||||
|
||||
Returns a list of ``(label, delivered)`` for every non-``None`` URL,
|
||||
in ``slack`` then ``discord`` order. ``None`` URLs are skipped (not
|
||||
attempted). Never raises (delegates to the never-raising
|
||||
:func:`post_webhook`).
|
||||
"""
|
||||
results: List[Tuple[str, bool]] = []
|
||||
for label, url in (("slack", slack_url), ("discord", discord_url)):
|
||||
if url:
|
||||
results.append(
|
||||
(
|
||||
label,
|
||||
post_webhook(
|
||||
url=url,
|
||||
payload=payload,
|
||||
timeout_seconds=timeout_seconds,
|
||||
),
|
||||
)
|
||||
)
|
||||
return results
|
||||
|
||||
|
||||
__all__ = [
|
||||
"post_webhook",
|
||||
"send_webhooks",
|
||||
"validate_webhook_url",
|
||||
]
|
||||
File diff suppressed because it is too large
Load Diff
Loading…
Reference in New Issue