feat(v0.71.5): ingest/data/prompt/drift polish

Closes #157, #205, #207, #149, #164, #163. Defers #204 (live SaaS pull —
paid accounts, infra-blocked, kept open).

- #164: get_metric_series falls back to eval_results when metrics is empty
- #163: build_verdict confidence biased by advise_history (same project+choice,
  >=3 precedents); decision never changes
- #207: shared utils/webhooks.py (SSRF-hardened) + --slack-url/--discord-url on
  ingest/prune-prompt/ab/active-sample; ab fires only on a decision
- #205: soup prune-prompt --tokenizer (token-prefix detect + decode remainder,
  boundary-safe)
- #149: DynamicCurriculumCallback buckets by loss/perplexity percentile;
  length keeps round-robin
- #157: soup data push/forge --hub modelscope|modelers (data score N/A)

107 new tests in tests/test_v0715.py (12474 -> 12581). ruff clean.
This commit is contained in:
Alpamys 2026-06-02 14:34:36 +05:00
parent 76fdd848cb
commit 1f63393421
28 changed files with 2361 additions and 141 deletions

View File

@ -12,6 +12,51 @@ reproducing 70+ versions of notes.
## [Unreleased]
## [0.71.5] - 2026-06-02
### Added
- **`soup eval against` now reads eval metrics** — `ExperimentTracker.get_metric_series`
falls back to the `eval_results` table when the metric is not a per-step
training column (`loss` / `lr` / `grad_norm` / `speed` / `gpu_mem`). So
`soup eval against <base> --candidate <run> --metric task_accuracy` returns a
real score series (benchmark scores live in `eval_results`, not `metrics`)
instead of "Empty series". Per-step columns still read from `metrics` — no
regression for existing callers.
- **`soup advise` learns from past project outcomes** — `soup advise` now reads
this project's accepted-verdict history (`~/.soup/advise_history.jsonl`) and
biases the rubric: 3+ successful SFT precedents flip a marginal RAG call to
SFT; 3+ negative GRPO outcomes suppress GRPO in favour of SFT-on-traces; an
encouraged choice gets a small confidence nudge. Scoped per-project (one
project's record never biases another). No history → identical to before.
- **Slack/Discord webhooks on four more commands** — `--slack-url` / `--discord-url`
(SSRF-hardened, loopback-only HTTP, RFC1918 rejected, never crashes the
command) now ship on `soup ingest`, `soup prune-prompt`, `soup ab` (fires only
on a `reject_h0` / `accept_h0` decision, not `continue`), and
`soup data active-sample` — not just `soup drift-alarm`. The validator + sender
moved to a shared `soup_cli/utils/webhooks.py`.
- **Tokenizer-aware `soup prune-prompt`** — `--tokenizer <model_or_path>` detects
and strips the shared system-prompt prefix on **token** boundaries instead of
characters, so a multi-byte UTF-8 prefix can never be split mid-code-point.
Default (no `--tokenizer`) keeps the whitespace-character behaviour.
- **Curriculum bucketing by loss percentile** — `DynamicCurriculumCallback` now
buckets samples by the percentile rank of the live loss (or perplexity)
signal within a rolling window when `data.curriculum_metric` is `loss` /
`perplexity`, so a consistently-hard sample is routed to the same difficulty
bucket across recomputes. `length` and warm-up still use round-robin.
- **`--hub` on `soup data push` and `soup data forge`** — `soup data push
--hub modelscope|modelers` uploads a dataset via the matching SDK
(`repo_type=dataset`, commit message sanitised); `soup data forge --hub
<non-hf> --teacher owner/name` pre-fetches the teacher model from that hub
(and warns when the teacher is not a repo id so `--hub` is never silently
ignored). HF stays the default.
### Notes
- Live SaaS *pull* adapters for `soup ingest` (Langfuse / LangSmith / Helicone /
OpenPipe / OpenAI SDKs, issue #204) remain deferred: they need credentialed
vendor accounts with populated trace data to validate honestly. Tracked as an
open, `infra-blocked` (external-account) item. `soup ingest` continues to parse
the JSONL export you pull from your dashboard.
## [0.71.4] - 2026-06-02
### Added

View File

@ -120,7 +120,7 @@ src/soup_cli/
templates/ - 17 built-in soup.yaml templates (YAML + manifest.json) with load_template loader (v0.39.0, +bco v0.40.0)
ui/ - Web UI (FastAPI + HTML/JS SPA)
tests/ - Test suite (274 files, 12474 tests)
tests/ - Test suite (275 files, 12581 tests)
examples/ - Real-world config examples and datasets
```

View File

@ -49,21 +49,22 @@ infrastructure instead of improving models. Soup fixes that.
## What's New
**v0.71.4 — Adapter lifecycle + loop wiring.** The merge, PR, and continuous-loop surfaces go live:
**v0.71.5 — Ingest, data & prompt polish.** Sharper production-loop ergonomics:
- **Canary verdict on merge** — `soup adapters merge … --canary suite.json` scores the merged
adapter and reports **OK / MINOR / MAJOR**; `--strict-verdict` exits non-zero on a MAJOR
regression. Works with no model load using a pre-scored canary suite.
- **Evolutionary merge for real** — `soup adapters merge --strategy cmaes --eval suite --budget 1h`
now runs the full CMA-ES search (merge → score → optimise) and writes the best blend, instead of
just printing a plan.
- **Publish an adapter PR** — `soup adapters pr <title> --base-sha <hex> --adapter <path> --push
owner/repo#42` posts the rendered PR straight to a GitHub PR comment.
- **Continuous fine-tuning loop, wired up** — `soup loop watch --pre-wired` runs the real
traces → DPO → eval-gate → canary pipeline; `--pack-cans` snapshots every iteration as a
shareable Soup Can with Registry lineage (`soup loop replay <id> --extract dir`).
- **Branches ↔ Registry** — `soup adapters branch <name> --attach-to-registry <id>` /
`--from-registry <id>` links training-env snapshots into the Registry lineage DAG.
- **Alternative model hubs for data** — `soup data push --hub modelscope|modelers` uploads a
local JSONL to ModelScope / Modelers, and `soup data forge --hub … --teacher owner/name`
pre-fetches the teacher from that hub.
- **Tokenizer-aware prompt pruning** — `soup prune-prompt --tokenizer <id-or-path>` finds the
shared *token* prefix and decodes only the remainder, so BPE multi-byte sequences never get
truncated mid-token the way char-slicing can.
- **Webhooks everywhere** — `--slack-url` / `--discord-url` now work on `soup ingest`,
`soup prune-prompt`, `soup ab`, and `soup data active-sample` (same SSRF-hardened validator as
`soup drift-alarm`). The A/B harness only pings when the sequential test actually decides.
- **Curriculum by difficulty percentile** — dynamic curriculum can bucket by `loss` /
`perplexity` percentile instead of length round-robin.
- **Smarter pre-flight `advise`** — `soup advise` now nudges its confidence using your prior
verdicts for the same project, and `soup runs replay` can plot a benchmark-score curve, not
just the loss curve.
Full history: [CHANGELOG.md](CHANGELOG.md) &middot; [GitHub Releases](https://github.com/MakazhanAlpamys/Soup/releases).

View File

@ -459,6 +459,11 @@ Every completed run also stores an estimated cost (`$` per run) computed from th
captured GPU device name and duration. `soup runs show` renders `—` for CPU /
MPS / unknown GPUs (no fabricated zeros).
As of v0.71.5, the metric-series lookup that powers replay (`ExperimentTracker.get_metric_series`)
transparently falls back to the `eval_results` table when a metric has no per-step
rows — so you can plot a benchmark-score curve (e.g. `mmlu`, `gsm8k`) the same way
you plot `loss`, without caring which table holds the series.
### Tracker integrations (--tracker mlflow / swanlab / trackio)
```bash

View File

@ -90,10 +90,12 @@ soup data download user/ds --samples 1000 Stream first 1000 samples
soup data register --name my-ds --path d.jsonl --format alpaca Register dataset
soup data unregister --name my-ds Remove from registry
soup data push --input d.jsonl --hf-dataset user/name Upload local JSONL as HF dataset
soup data push --input d.jsonl --hf-dataset u/n --hub modelscope|modelers Upload to an alternative hub
soup data registry List all registered datasets
soup data demo List bundled demo JSONL fixtures
soup data demo alpaca_demo --output ./d.jsonl Copy a bundled demo JSONL fixture
soup data forge --docs ./docs --task sft --target-rows 1000 Synthetic data pipeline + provenance
soup data forge --docs ./docs --hub modelscope --teacher owner/name Pre-fetch the teacher from an alternative hub
soup data score --input rows.jsonl Composite quality scorecard (PII + toxicity + lang + edu)
soup data decontaminate --input rows.jsonl --benchmarks mmlu,gsm8k Drop benchmark-overlap rows
soup data toxicity --input rows.jsonl -o tox.jsonl Flag toxic rows (keyword baseline)
@ -137,7 +139,7 @@ soup can publish r.can --hf-hub user/name Publish .can to HF Hub as dataset
soup runs List training runs
soup runs show <run_id> Run details + loss graph + cost
soup runs compare <run_1> <run_2> Compare two runs
soup runs replay <run_id> Replay summary + loss curve from history
soup runs replay <run_id> Replay summary + loss curve from history (also plots a benchmark-score curve when the metric lives in eval_results)
soup why [run_id] Explain training anomalies (heuristic)
soup tui Full-screen Textual dashboard (requires [tui] extra)
soup train --config soup.yaml --profile Record torch.profiler trace to <output>/profiles/
@ -169,8 +171,10 @@ soup edit set --base <m> --method rome|memit|alphaedit --subject "..." --target
soup edit diff <before-run> <after-run> --probes p.jsonl Knowledge-injection diff visualizer
soup ingest --source langfuse|langsmith|helicone|openpipe|otel|openai-stored --logs <jsonl> Universal trace importer (6 SaaS adapters → normalised JSONL)
soup prune-prompt --input <jsonl> --output <jsonl> --min-frequency 0.95 Detect + strip shared system-prompt prefix
soup prune-prompt ... --tokenizer <id-or-path> Tokenizer-aware prefix detection (decodes remaining ids, boundary-safe)
soup data active-sample --input <jsonl> --output <jsonl> --budget N Top-N uncertain prod traces for human review
soup ab --input <jsonl> --metric latency|judge_score|retry_rate mSPRT sequential A/B (decision: continue / reject_h0 / accept_h0)
soup ingest|prune-prompt|ab|data active-sample ... --slack-url <https> | --discord-url <https> Shared SSRF-validated webhook on completion
soup drift-alarm --reference <jsonl> --live <jsonl> --threshold 0.2 Rolling-KL drift alarm (exit 3 on drift)
soup drift-alarm ... --slack-url <https> | --discord-url <https> Optional SSRF-validated webhook on drift detected
soup tunability --list List built-in candidate-base catalogue

View File

@ -94,6 +94,14 @@ soup prune-prompt --input traces.jsonl --output pruned.jsonl --min-frequency 0.9
Binary-search over up-to-32 candidate templates finds the longest qualifying prefix (a longer threshold-meeting prefix may exist beyond the universal one — Soup does not early-exit on the 100% match). Two-pass file read with a 100 000-row DoS cap.
**Tokenizer-aware mode (v0.71.5).** Pass `--tokenizer <id-or-path>` (a HuggingFace repo id, a local path, or anything `AutoTokenizer.from_pretrained` accepts) to detect the shared prefix in *token* space and decode only the remaining ids:
```bash
soup prune-prompt --input traces.jsonl --output pruned.jsonl --tokenizer Qwen/Qwen2.5-0.5B
```
Char-level stripping can cut a BPE multi-byte sequence in half when the shared prefix ends mid-token; token-aware pruning finds the longest shared *token-id* prefix and decodes the remainder, so the boundary always lands on a real token. Per-row encoding is capped at 50 000 tokens. Omit `--tokenizer` to keep the original character-level behaviour.
## Active-Learning Sampler (`soup data active-sample`)
@ -108,6 +116,8 @@ soup data active-sample --input traces.jsonl --output for-review.jsonl --budget
The output JSONL is a drop-in prompt set for `soup eval human` (v0.19). Budget is bounded `[1, 100 000]`.
**Webhooks (v0.71.5).** `soup ingest`, `soup prune-prompt`, `soup ab`, and `soup data active-sample` all accept `--slack-url` / `--discord-url` and POST a one-line summary on completion through the same SSRF-hardened validator as `soup drift-alarm` (scheme allowlist, loopback-only HTTP, RFC1918 / link-local / reserved / multicast rejected; the post never raises, so a flaky webhook can't fail the command). `soup ab` only fires when the sequential test actually decides (`reject_h0` / `accept_h0`), not while it's still `continue`-ing.
## Synthetic Data Generation
@ -479,6 +489,8 @@ Three tasks supported: `sft` (Q&A pairs), `preference` (chosen/rejected), `tool`
Document discovery is one level deep over `.txt` / `.md` / `.json` / `.jsonl`; dotfiles + symlinked directories are skipped. All paths are cwd-contained, all writes are atomic via staged-tempfile + `os.replace`, and write targets are rejected if they're symlinks. **Judge providers are live**: `--judge-provider ollama` (localhost-only), `--judge-provider anthropic` (env-only API key), `--judge-provider vllm` (scheme-validated). Per-call judge exceptions logged at DEBUG.
**Alternative teacher hubs (v0.71.5).** `--hub modelscope|modelers` pre-fetches the `--teacher` from that hub when the teacher is a routable repo id (`owner/name`); `--hub hf` (default) is a no-op and leaves the teacher as a provenance label. If `--hub` is non-HF but `--teacher` is not a repo id (e.g. the default `local-judge`), Soup prints a loud yellow warning rather than silently dropping the flag.
## Data Quality Scorecard

View File

@ -82,6 +82,8 @@ soup advise compare
4. Task is `factual_lookup` with high output variance → **RAG**.
5. Otherwise → **SFT**.
**Cross-project confidence bias (v0.71.5).** When `~/.soup/advise_history.jsonl` holds ≥3 prior verdicts for the *same choice* in the *same project*, `soup advise` nudges its confidence (not its decision) toward what worked before: a net-positive precedent record (you accepted it AND its recorded outcome was good) bumps confidence up by a small constant; a net-negative one bumps it down. The rubric verdict itself never changes — only how sure Soup is. Verdicts must be `--record`ed for the bias to kick in.
**Why this command exists.** "Choose fine-tuning vs RAG vs prompt-engineering" is the most-mis-made decision in the space. Reddit, HN, IBM, and Google Cloud all converge on the same advice (start with prompts, escalate to RAG, fine-tune as last resort) and almost everyone ignores it because nobody has the data to prove their case is the exception. Soup `autopilot` picks hyperparameters AFTER you've decided to train; `soup advise` owns the layer above. No trainer library has an incentive to tell users *not to train* — Unsloth's funnel, Axolotl's hosted business, LLaMA-Factory's Alibaba alignment all monetise the training event.
@ -213,6 +215,8 @@ soup ab --input ab.jsonl --metric judge_score --alpha 0.01 --beta 0.10 --effect-
Input rows look like `{"arm": "control", "latency": 1.23}` or `{"arm": "treatment", "judge_score": 0.91}`. Decision is one of `continue` (keep collecting samples), `reject_h0` (real difference detected), `accept_h0` (no significant difference). Composes with `soup loop canary` (v0.58) — promote or roll back as soon as the LLR clears a decision boundary.
`soup ab` accepts `--slack-url` / `--discord-url` (v0.71.5) and pings the webhook **only when the test actually decides** (`reject_h0` / `accept_h0`) — a still-running `continue` stays quiet so you're not paged on every peek. Same SSRF-hardened validator as `soup drift-alarm`.
## Drift Alarm (`soup drift-alarm`)

View File

@ -907,12 +907,15 @@ Layer dynamic re-weighting on top of the static `curriculum` bucketer. Every N s
training:
curriculum: true # static bucketer (v0.23.0)
curriculum_buckets: 4
curriculum_metric: perplexity # length (default) | loss | perplexity
curriculum_dynamic: true # NEW — dynamic re-weighting
curriculum_dynamic_recompute_steps: 50 # refresh every 50 global steps
curriculum_dynamic_floor: 0.05 # min weight per bucket
curriculum_dynamic_temperature: 1.0 # softmax temp on uncertainty
```
**Bucketing by difficulty percentile (v0.71.5).** When `curriculum_metric` is `loss` or `perplexity`, the dynamic callback assigns each step's sample to a bucket by its *rank* within a rolling 512-step window of the difficulty signal (perplexity = `exp(min(loss, 50))`), instead of the round-robin fallback used for `length`. This keeps the buckets calibrated to the live loss distribution rather than a static length sort. `length` (the default) keeps the round-robin assignment.
Visualise the recorded bucket-weight evolution with `soup runs curriculum-curve <run_id>`.
DDP / grad-accum safety: multi-rank launches must wire an `all_reduce` hook on per-bucket stats (a cross-validator rejects un-coordinated multi-rank runs upfront). Multi-trainer expansion beyond `sft` / `pretrain` is tracked for v0.48.1.

View File

@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "soup-cli"
version = "0.71.4"
version = "0.71.5"
description = "Fine-tune LLMs in one command. No SSH, no config hell."
readme = "README.md"
license = "Apache-2.0"

View File

@ -1,3 +1,3 @@
"""Soup CLI — Fine-tune LLMs in one command."""
__version__ = "0.71.4"
__version__ = "0.71.5"

View File

@ -0,0 +1,71 @@
"""Shared CLI glue for the --slack-url / --discord-url webhook flags (v0.71.5 #207).
Keeps Typer + Rich Console concerns in the commands layer (``utils/webhooks``
stays import-light + framework-free). Used by ``ingest`` / ``prune-prompt`` /
``ab`` / ``data active-sample`` so the validate-then-deliver pattern is defined
once.
"""
from __future__ import annotations
from typing import Optional, Tuple
import typer
from rich.console import Console
from rich.markup import escape
def validate_webhook_flags(
slack_url: Optional[str],
discord_url: Optional[str],
*,
console: Console,
) -> Tuple[Optional[str], Optional[str]]:
"""Validate webhook URLs at the CLI boundary (``typer.Exit(2)`` on bad).
Returns the (canonical) URLs. Mirrors the v0.63.0 ``drift-alarm``
early-rejection pattern so a typo'd / SSRF-y URL fails fast with a
friendly message instead of being silently swallowed at delivery time.
"""
from soup_cli.utils.webhooks import validate_webhook_url
out = []
for label, value in (("--slack-url", slack_url), ("--discord-url", discord_url)):
if value is None:
out.append(None)
continue
try:
out.append(validate_webhook_url(value))
except (TypeError, ValueError) as exc:
console.print(f"[red]{label}: {escape(str(exc))}[/]")
raise typer.Exit(2) from exc
return out[0], out[1]
def emit_webhooks(
slack_url: Optional[str],
discord_url: Optional[str],
*,
payload: dict,
console: Console,
) -> None:
"""POST the completion payload to any configured webhooks (best-effort).
Never raises (delegates to the never-raising
:func:`soup_cli.utils.webhooks.send_webhooks`). Prints a per-target
delivered/failed line so the operator sees whether the alert landed.
"""
if slack_url is None and discord_url is None:
return
from soup_cli.utils.webhooks import send_webhooks
for label, ok in send_webhooks(
payload, slack_url=slack_url, discord_url=discord_url
):
colour = "green" if ok else "yellow"
console.print(
f"[{colour}]{label} webhook: {'delivered' if ok else 'failed'}[/]"
)
__all__ = ["emit_webhooks", "validate_webhook_flags"]

View File

@ -2,12 +2,15 @@
from __future__ import annotations
from typing import Optional
import typer
from rich.console import Console
from rich.markup import escape
from rich.panel import Panel
from rich.table import Table
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
from soup_cli.utils.ab_test import (
MsprtConfig,
run_msprt,
@ -35,6 +38,20 @@ def ab(
0.1, "--effect-size",
help="Minimum detectable difference in means.",
),
slack_url: Optional[str] = typer.Option(
None, "--slack-url",
help=(
"Optional Slack webhook URL — POSTed on a reject_h0 / accept_h0 "
"decision (not on continue). SSRF-validated."
),
),
discord_url: Optional[str] = typer.Option(
None, "--discord-url",
help=(
"Optional Discord webhook URL — POSTed on a reject_h0 / accept_h0 "
"decision (not on continue). SSRF-validated."
),
),
) -> None:
"""Sequential A/B test with early-stop guarantees (mSPRT)."""
try:
@ -43,6 +60,10 @@ def ab(
console.print(f"[red]{escape(str(exc))}[/]")
raise typer.Exit(2) from exc
slack_url, discord_url = validate_webhook_flags(
slack_url, discord_url, console=console
)
try:
cfg = MsprtConfig(
metric=canonical, alpha=alpha, beta=beta, effect_size=effect_size,
@ -101,5 +122,24 @@ def ab(
)
)
# Webhook only fires on a terminal decision (reject_h0 / accept_h0) —
# a `continue` verdict carries no actionable signal (issue #207).
if verdict.decision != "continue":
emit_webhooks(
slack_url,
discord_url,
payload={
"command": "ab",
"metric": canonical,
"decision": verdict.decision,
"log_likelihood_ratio": verdict.log_likelihood_ratio,
"n_control": verdict.n_control,
"n_treatment": verdict.n_treatment,
"mean_control": verdict.mean_control,
"mean_treatment": verdict.mean_treatment,
},
console=console,
)
__all__ = ["ab"]

View File

@ -2,11 +2,14 @@
from __future__ import annotations
from typing import Optional
import typer
from rich.console import Console
from rich.markup import escape
from rich.panel import Panel
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
from soup_cli.utils.active_sampler import sample_uncertain_rows, validate_budget
console = Console()
@ -24,6 +27,14 @@ def active_sample(
100, "--budget",
help="Max rows to surface for human review (1 - 100_000).",
),
slack_url: Optional[str] = typer.Option(
None, "--slack-url",
help="Optional Slack webhook URL — POSTed on completion. SSRF-validated.",
),
discord_url: Optional[str] = typer.Option(
None, "--discord-url",
help="Optional Discord webhook URL — POSTed on completion. SSRF-validated.",
),
) -> None:
"""Surface the most uncertain prod traces for human review."""
try:
@ -32,6 +43,10 @@ def active_sample(
console.print(f"[red]{escape(str(exc))}[/]")
raise typer.Exit(2) from exc
slack_url, discord_url = validate_webhook_flags(
slack_url, discord_url, console=console
)
try:
plan = sample_uncertain_rows(
input_path,
@ -55,5 +70,18 @@ def active_sample(
)
)
emit_webhooks(
slack_url,
discord_url,
payload={
"command": "active-sample",
"rows_selected": plan.rows_selected,
"rows_in": plan.rows_in,
"mean_uncertainty": plan.mean_uncertainty,
"budget": plan.budget,
},
console=console,
)
__all__ = ["active_sample"]

View File

@ -39,6 +39,7 @@ from soup_cli.utils.advise import (
synth_probe_lora_delta,
)
from soup_cli.utils.advise_history import (
current_project_name,
history_path,
load_history,
record_verdict,
@ -269,8 +270,27 @@ def advise_run(
sft_wall_clock_secs=wall_clock,
)
# v0.71.5 #163 — bias the rubric by this project's past accepted-verdict
# outcomes. Best-effort: a missing / unreadable history must never block
# a verdict, so any failure falls back to the un-biased rubric.
history = None
project = None
try:
verdict = build_verdict(profile, task_category, goal=goal, roi=roi)
history = load_history(limit=20)
project = current_project_name()
except (TypeError, ValueError, OSError):
history = None
project = None
try:
verdict = build_verdict(
profile,
task_category,
goal=goal,
roi=roi,
history=history,
project=project,
)
except (TypeError, ValueError) as exc:
console.print(f"[red]Verdict build failed:[/] {escape(str(exc))}")
raise typer.Exit(1) from exc

View File

@ -1914,16 +1914,32 @@ def push_dataset_cmd(
"--message",
help="Commit message for the dataset upload",
),
hub: str = typer.Option(
"hf",
"--hub",
help=(
"Target hub: hf (default) / modelscope / modelers. Non-HF hubs "
"upload via the matching SDK (v0.71.5 #157) — repo_type=dataset, "
"commit message sanitised to first line + 200 chars."
),
),
):
"""Upload a local JSONL dataset to HuggingFace Hub as a dataset repo."""
"""Upload a local JSONL dataset to a model hub as a dataset repo."""
from soup_cli.utils.hf import (
get_hf_api,
resolve_endpoint,
resolve_token,
validate_repo_id,
)
from soup_cli.utils.hubs import validate_hub_name
from soup_cli.utils.paths import is_under_cwd
try:
hub_canonical = validate_hub_name(hub)
except (TypeError, ValueError) as exc:
console.print(f"[red]{exc}[/]")
raise typer.Exit(2) from exc
file_path = Path(input_path)
if not file_path.exists():
console.print(f"[red]Dataset file not found: {file_path}[/]")
@ -1940,9 +1956,18 @@ def push_dataset_cmd(
try:
validate_repo_id(hf_dataset)
except ValueError as exc:
console.print(f"[red]Invalid --hf-dataset repo id:[/] {exc}")
console.print(f"[red]Invalid --{hub_canonical}-dataset repo id:[/] {exc}")
raise typer.Exit(1) from exc
# v0.71.5 #157 — non-HF hubs route through the shared upload_repo adapter.
# ModelScope / Modelers SDKs upload a folder, so the single JSONL is
# staged into a temp dir and uploaded as a dataset repo.
if hub_canonical != "hf":
_push_dataset_non_hf(
hub_canonical, hf_dataset, file_path, commit_message
)
return
token = resolve_token()
if token is None:
console.print(
@ -1987,6 +2012,53 @@ def push_dataset_cmd(
)
def _push_dataset_non_hf(
hub: str, repo_id: str, file_path: Path, commit_message: str
) -> None:
"""Upload a single JSONL to a non-HF hub as a dataset (v0.71.5 #157).
``upload_repo`` uploads a folder, so the file is staged into a temp dir
first. ``upload_repo`` already sanitises the commit message (first line +
200 chars) and validates the repo id shape.
"""
import shutil
import tempfile
from rich.markup import escape
from soup_cli.utils.hubs import upload_repo
# Stage UNDER cwd — `upload_repo` enforces cwd-containment on folder_path
# (the system tempdir would be rejected).
staging = tempfile.mkdtemp(prefix=".soup_dataset_push.", dir=os.getcwd())
try:
shutil.copy2(str(file_path), os.path.join(staging, file_path.name))
try:
upload_repo(
hub,
repo_id,
folder_path=staging,
commit_message=commit_message,
repo_type="dataset",
)
except ImportError as exc:
console.print(f"[red]{exc}[/]")
raise typer.Exit(1) from exc
except (TypeError, ValueError) as exc:
console.print(f"[red]Upload failed:[/] {exc}")
raise typer.Exit(1) from exc
except Exception as exc: # noqa: BLE001 — surface SDK errors generically
console.print(f"[red]Upload failed:[/] {exc}")
raise typer.Exit(1) from exc
finally:
shutil.rmtree(staging, ignore_errors=True)
console.print(
f"[green]Uploaded[/] {escape(file_path.name)} to {escape(hub)} "
f"dataset [bold]{escape(repo_id)}[/]"
)
# --- v0.42.0 Part C / F: AOT preprocess + document ingestion ---------------
@app.command(name="preprocess")

View File

@ -77,6 +77,14 @@ def forge(
"(scheme allowlist + loopback). Ignored for Anthropic."
),
),
hub: str = typer.Option(
"hf", "--hub",
help=(
"Teacher hub: hf (default) / modelscope / modelers. When non-HF "
"and --teacher is a repo id (owner/name), the teacher is "
"pre-fetched from that hub (v0.71.5 #157)."
),
),
):
"""Run the multi-stage synthetic data pipeline with provenance.
@ -94,9 +102,45 @@ def forge(
write_forge_dataset,
write_provenance,
)
from soup_cli.utils.hubs import validate_hub_name
try:
hub_canonical = validate_hub_name(hub)
except (TypeError, ValueError) as exc:
console.print(f"[red]{escape(str(exc))}[/]")
raise typer.Exit(2) from exc
effective_teacher = teacher
# v0.71.5 #157 — non-HF hub + repo-id teacher → pre-fetch the teacher from
# that hub and record the resolved local path in provenance. HF (default)
# is a no-op: the teacher stays a provenance label.
if hub_canonical != "hf":
if "/" in teacher:
from soup_cli.utils.hubs import prefetch_model_from_hub
try:
effective_teacher = prefetch_model_from_hub(
teacher, hub_canonical, console=console
)
except ImportError as exc:
console.print(f"[red]{escape(str(exc))}[/]")
raise typer.Exit(1) from exc
except (TypeError, ValueError) as exc:
console.print(
f"[red]Teacher pre-fetch failed:[/] {escape(str(exc))}"
)
raise typer.Exit(1) from exc
else:
# Non-HF hub requested but the teacher is not a routable repo id
# (owner/name) — warn loudly instead of silently ignoring --hub
# (code-review MEDIUM fix v0.71.5 #157).
console.print(
f"[yellow]--hub {escape(hub_canonical)} ignored:[/] --teacher "
f"{escape(teacher)!r} is not a repo id (owner/name), so there "
"is nothing to pre-fetch."
)
judge_fn = _default_judge
effective_teacher = teacher
if judge_provider is not None:
canonical = judge_provider.strip().lower()
if canonical not in JUDGE_PROVIDERS:

View File

@ -21,6 +21,7 @@ from rich.console import Console
from rich.markup import escape
from rich.panel import Panel
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
from soup_cli.utils import ingest_sources as _ingest_sources
from soup_cli.utils.ingest_sources import (
SUPPORTED_INGEST_SOURCES,
@ -53,6 +54,14 @@ def ingest(
"-o",
help="Output JSONL (default: traces.jsonl in cwd).",
),
slack_url: Optional[str] = typer.Option(
None, "--slack-url",
help="Optional Slack webhook URL — POSTed on completion. SSRF-validated.",
),
discord_url: Optional[str] = typer.Option(
None, "--discord-url",
help="Optional Discord webhook URL — POSTed on completion. SSRF-validated.",
),
) -> None:
"""Import production traces from a SaaS observability vendor (v0.63.0).
@ -66,6 +75,10 @@ def ingest(
console.print(f"[red]{escape(str(exc))}[/]")
raise typer.Exit(2) from exc
slack_url, discord_url = validate_webhook_flags(
slack_url, discord_url, console=console
)
if not is_under_cwd(logs):
console.print(f"[red]--logs '{escape(logs)}' is outside cwd — refusing[/]")
raise typer.Exit(1)
@ -111,6 +124,18 @@ def ingest(
f"{escape(output_path.name)}[/]"
)
emit_webhooks(
slack_url,
discord_url,
payload={
"command": "ingest",
"source": canonical,
"traces_written": count,
"auth_env_set": auth_value is not None,
},
console=console,
)
def _env_label(source: str) -> str:
"""Return the env-var name that authenticates ``source``.

View File

@ -7,11 +7,14 @@ signature trick — shipped OSS for v0.63.0 Part B).
from __future__ import annotations
from typing import Optional
import typer
from rich.console import Console
from rich.markup import escape
from rich.panel import Panel
from soup_cli.commands._webhook_cli import emit_webhooks, validate_webhook_flags
from soup_cli.utils.prune_prompt import prune_traces, validate_min_frequency
console = Console()
@ -35,6 +38,22 @@ def prune_prompt_cmd(
"--min-frequency",
help="Prefix must appear in >= this fraction of rows to be stripped (0.0 - 1.0).",
),
tokenizer: Optional[str] = typer.Option(
None,
"--tokenizer",
help=(
"Optional HF tokenizer (model id or local path) for token-aware "
"prefix detection. Default: whitespace-character level."
),
),
slack_url: Optional[str] = typer.Option(
None, "--slack-url",
help="Optional Slack webhook URL — POSTed on completion. SSRF-validated.",
),
discord_url: Optional[str] = typer.Option(
None, "--discord-url",
help="Optional Discord webhook URL — POSTed on completion. SSRF-validated.",
),
) -> None:
"""Detect + strip a shared system-prompt prefix (v0.63.0 Part B)."""
try:
@ -43,11 +62,16 @@ def prune_prompt_cmd(
console.print(f"[red]{escape(str(exc))}[/]")
raise typer.Exit(2) from exc
slack_url, discord_url = validate_webhook_flags(
slack_url, discord_url, console=console
)
try:
report = prune_traces(
input_path,
output_path=output_path,
min_frequency=min_frequency,
tokenizer=tokenizer,
)
except FileNotFoundError:
console.print(f"[red]Input not found: {escape(input_path)}[/]")
@ -56,6 +80,14 @@ def prune_prompt_cmd(
console.print(f"[red]{escape(str(exc))}[/]")
raise typer.Exit(1) from exc
payload = {
"command": "prune-prompt",
"prefix_found": bool(report.prefix),
"prefix_chars": report.prefix_chars,
"rows_pruned": report.rows_pruned,
"rows_total": report.rows_total,
}
if not report.prefix:
console.print(
Panel(
@ -65,6 +97,7 @@ def prune_prompt_cmd(
border_style="yellow",
)
)
emit_webhooks(slack_url, discord_url, payload=payload, console=console)
return
snippet = report.prefix if len(report.prefix) <= 200 else report.prefix[:200] + "..."
@ -77,6 +110,7 @@ def prune_prompt_cmd(
border_style="green",
)
)
emit_webhooks(slack_url, discord_url, payload=payload, console=console)
__all__ = ["prune_prompt_cmd"]

View File

@ -318,6 +318,15 @@ class ExperimentTracker:
Used by ``soup eval against`` for run-vs-run paired-bootstrap CI.
Returns an empty list when the metric does not appear in any row
— the caller treats that as "no signal, do not gate".
v0.71.5 #164: the per-step ``metrics`` table only carries training
columns (``loss`` / ``lr`` / ``grad_norm`` / ``speed`` / ``gpu_mem``).
Eval metrics like ``task_accuracy`` / ``refusal_rate`` live in the
``eval_results`` table instead. So when the per-step pass yields no
rows we fall back to the per-benchmark scores in ``eval_results``.
Querying ``metrics`` first preserves the established behaviour for
every training-loop column (no regression for existing callers);
the fallback only fires when the column path is empty.
"""
if not isinstance(run_id, str) or not run_id:
raise ValueError("run_id must be a non-empty string")
@ -335,7 +344,34 @@ class ExperimentTracker:
# Skip non-numeric cells silently — same-run inconsistency
# is not the caller's problem; they get a shorter series.
continue
return series
if series:
return series
# Bridge to eval_results (v0.71.5 #164) — benchmark scores for
# `soup eval against`. Empty when neither table has data.
return self._eval_score_series(run_id, metric)
def _eval_score_series(self, run_id: str, benchmark: str) -> list[float]:
"""Return the per-row ``score`` series from ``eval_results``.
Ordered by insertion (``id``) for deterministic pairing in the
paired-bootstrap CI. Non-numeric cells are skipped silently.
"""
conn = self._get_conn()
rows = conn.execute(
"SELECT score FROM eval_results "
"WHERE run_id = ? AND benchmark = ? ORDER BY id",
(run_id, benchmark),
).fetchall()
out: list[float] = []
for row in rows:
value = row["score"]
if value is None:
continue
try:
out.append(float(value))
except (TypeError, ValueError):
continue
return out
def save_eval_result(
self,

View File

@ -31,14 +31,17 @@ from __future__ import annotations
import json
import logging
import math
import os
import stat
import tempfile
from typing import Any, Dict, Optional, Tuple
from collections import deque
from typing import Any, Deque, Dict, Optional, Tuple
from soup_cli.utils.curriculum_dynamic import (
DynamicCurriculumPolicy,
compute_bucket_weights,
percentile_bucket,
)
from soup_cli.utils.paths import is_under_cwd
@ -46,6 +49,14 @@ logger = logging.getLogger(__name__)
_MAX_PATH_LEN = 4096
_HISTORY_FILENAME = "curriculum_history.jsonl"
# Curriculum metrics that drive percentile bucketing from the live loss
# signal (v0.71.5 #149). ``length`` has no per-step signal in HF logs, so it
# falls back to round-robin (length bucketing is the static curriculum's
# data-prep-time job — see utils/curriculum.py).
_PERCENTILE_METRICS = frozenset({"loss", "perplexity"})
_VALID_CURRICULUM_METRICS = frozenset({"length", "perplexity", "loss"})
# Rolling-window size for the percentile reference distribution.
_SIGNAL_WINDOW = 512
__all__ = [
"DynamicCurriculumCallback",
@ -149,14 +160,25 @@ class DynamicCurriculumCallback(_try_import_callback_base()): # type: ignore[mi
self,
policy: DynamicCurriculumPolicy,
output_dir: str,
curriculum_metric: str = "length",
) -> None:
if not isinstance(policy, DynamicCurriculumPolicy):
raise TypeError(
"policy must be DynamicCurriculumPolicy, got "
f"{type(policy).__name__}"
)
if curriculum_metric not in _VALID_CURRICULUM_METRICS:
raise ValueError(
"curriculum_metric must be one of "
f"{sorted(_VALID_CURRICULUM_METRICS)}, got {curriculum_metric!r}"
)
self._policy = policy
self._curriculum_metric = curriculum_metric
self._output_dir = _validate_output_dir(output_dir)
# Rolling difficulty-signal window for percentile bucketing (v0.71.5
# #149). Persists across recomputes so a consistently-hard sample
# keeps landing in the same bucket.
self._signal_window: Deque[float] = deque(maxlen=_SIGNAL_WINDOW)
# Per-bucket accumulator: bucket_id -> {num_samples, loss_sum, grad_norm_sum}
self._stats: Dict[int, Dict[str, float]] = {}
# Most recently computed weights (read by external sampler hook).
@ -173,6 +195,10 @@ class DynamicCurriculumCallback(_try_import_callback_base()): # type: ignore[mi
def policy(self) -> DynamicCurriculumPolicy:
return self._policy
@property
def curriculum_metric(self) -> str:
return self._curriculum_metric
@property
def output_dir(self) -> str:
return self._output_dir
@ -215,13 +241,55 @@ class DynamicCurriculumCallback(_try_import_callback_base()): # type: ignore[mi
# HF emits "loss" + (optionally) "grad_norm" in `logs`.
loss = logs.get("loss")
grad_norm = logs.get("grad_norm")
# Bucket id derived from the global step (round-robin BETA strategy).
try:
bucket_id = _pick_bucket(global_step, self._policy.num_buckets)
except (TypeError, ValueError):
return
nb = self._policy.num_buckets
# v0.71.5 #149: percentile bucketing on the difficulty signal for
# loss / perplexity once the rolling window has warmed up; otherwise
# round-robin (warm-up + the `length` metric, which has no per-step
# signal in HF logs).
signal = self._difficulty_signal(loss)
bucket_id: Optional[int] = None
if (
self._curriculum_metric in _PERCENTILE_METRICS
and signal is not None
and len(self._signal_window) > 0
):
try:
bucket_id = percentile_bucket(
signal, list(self._signal_window), nb
)
except (TypeError, ValueError):
bucket_id = None
if bucket_id is None:
try:
bucket_id = _pick_bucket(global_step, nb)
except (TypeError, ValueError):
return
if signal is not None:
self._signal_window.append(signal)
self._record_sample(bucket_id, loss, grad_norm)
def _difficulty_signal(self, loss: object) -> Optional[float]:
"""Map the logged loss to the configured difficulty signal.
Returns ``None`` when no usable signal is available (metric is
``length`` — no per-step length in HF logs — or the loss is
missing / non-finite), which routes the step through the
round-robin fallback.
"""
if self._curriculum_metric not in _PERCENTILE_METRICS:
return None
try:
loss_f = float(loss) if loss is not None else None
except (TypeError, ValueError):
return None
if loss_f is None or not math.isfinite(loss_f):
return None
if self._curriculum_metric == "perplexity":
# exp is monotonic in loss, so percentile ranks are identical;
# clamp the exponent to avoid overflow on a stray loss spike.
return math.exp(min(loss_f, 50.0))
return loss_f
def on_step_end(
self,
args: Any,

View File

@ -26,7 +26,10 @@ import re
import stat
from collections.abc import Iterable
from dataclasses import dataclass, field
from typing import Dict, List, Mapping, Optional, Sequence, Tuple
from typing import TYPE_CHECKING, Dict, List, Mapping, Optional, Sequence, Tuple
if TYPE_CHECKING: # pragma: no cover — annotation only, avoids circular import.
from soup_cli.utils.advise_history import HistoryEntry
from soup_cli.utils.paths import is_under_cwd
@ -59,6 +62,17 @@ _MIN_ROWS_FOR_TRAINING = 50
# routed into a GPU-intensive RL run).
_MIN_ROWS_FOR_GRPO = 500
# History-bias thresholds (v0.71.5 #163). A choice is "encouraged" when the
# project has >= _HISTORY_MIN_PRECEDENTS accepted verdicts whose mean measured
# outcome is >= _HISTORY_POSITIVE_OUTCOME; "discouraged" when the mean is
# < _HISTORY_NEGATIVE_OUTCOME over the same minimum count. Encouraged choices
# get a small confidence nudge and can flip a marginal tie; discouraged choices
# are suppressed (their effective row-floor is raised).
_HISTORY_MIN_PRECEDENTS = 3
_HISTORY_POSITIVE_OUTCOME = 0.3
_HISTORY_NEGATIVE_OUTCOME = 0.0
_HISTORY_CONFIDENCE_NUDGE = 0.05
# Probe defaults — tiny, no GPU required for the heuristic stubs.
_PROBE_HOLDOUT_DEFAULT = 100
_PROBE_LORA_STEPS_DEFAULT = 100
@ -460,7 +474,7 @@ def _confidence_from_signals(*, row_count: int, diversity: float) -> float:
return max(0.2, min(0.95, 0.4 + 0.4 * size_score + 0.2 * diversity))
def build_verdict(
def _base_verdict(
profile: DatasetProfile,
task_category: str,
*,
@ -584,6 +598,184 @@ def build_verdict(
)
# ---------------------------------------------------------------------------
# History bias (v0.71.5 #163) — tune the rubric from past project outcomes
# ---------------------------------------------------------------------------
def _summarise_history_outcomes(
history: Optional[Sequence["HistoryEntry"]],
*,
project: Optional[str] = None,
) -> Dict[str, Tuple[float, int]]:
"""Aggregate accepted-verdict outcomes per choice from prior history.
Duck-typed (reads ``.choice`` / ``.accepted`` / ``.outcome`` / ``.project``
attributes) so :mod:`advise` never imports :mod:`advise_history` at runtime
— that import direction is owned by ``advise_history`` and reversing it
would create a cycle.
Filtering:
- only entries whose ``accepted is True`` (a rejected verdict carries no
endorsement signal),
- only entries with a finite ``outcome`` in ``[-1, 1]`` (bool / None /
out-of-range skipped — defensive even though ``HistoryEntry`` validates
on construction),
- only entries matching ``project`` when supplied (per-project scoping:
one project's SFT wins must not bias another project's verdict).
Returns ``{choice: (mean_outcome, count)}`` for every choice with >= 1
qualifying entry.
"""
if history is None:
return {}
if isinstance(history, (str, bytes)) or not isinstance(history, Sequence):
raise TypeError("history must be a non-string Sequence or None")
acc: Dict[str, Tuple[float, int]] = {}
for entry in history:
choice = getattr(entry, "choice", None)
if choice not in CHOICES:
continue
if getattr(entry, "accepted", None) is not True:
continue
if project is not None and getattr(entry, "project", None) != project:
continue
outcome = getattr(entry, "outcome", None)
if outcome is None or isinstance(outcome, bool):
continue
if not isinstance(outcome, (int, float)):
continue
f_out = float(outcome)
if not math.isfinite(f_out) or not (-1.0 <= f_out <= 1.0):
continue
total, count = acc.get(choice, (0.0, 0))
acc[choice] = (total + f_out, count + 1)
return {ch: (total / count, count) for ch, (total, count) in acc.items()}
def _is_encouraged(bias: Mapping[str, Tuple[float, int]], choice: str) -> bool:
mean_count = bias.get(choice)
if mean_count is None:
return False
mean, count = mean_count
return count >= _HISTORY_MIN_PRECEDENTS and mean >= _HISTORY_POSITIVE_OUTCOME
def _is_discouraged(bias: Mapping[str, Tuple[float, int]], choice: str) -> bool:
mean_count = bias.get(choice)
if mean_count is None:
return False
mean, count = mean_count
return count >= _HISTORY_MIN_PRECEDENTS and mean < _HISTORY_NEGATIVE_OUTCOME
def _precedent_count(bias: Mapping[str, Tuple[float, int]], choice: str) -> int:
mean_count = bias.get(choice)
return mean_count[1] if mean_count is not None else 0
def build_verdict(
profile: DatasetProfile,
task_category: str,
*,
goal: Optional[str] = None,
roi: Optional[ROIEstimate] = None,
history: Optional[Sequence["HistoryEntry"]] = None,
project: Optional[str] = None,
) -> Verdict:
"""Combine profile + task into a recommendation, optionally history-biased.
Without ``history`` this is byte-identical to the v0.54.0 rubric
(regression-guarded). When ``history`` is supplied the per-project
outcome record nudges marginal decisions (v0.71.5 #163):
- A choice with >= 3 accepted verdicts averaging >= +0.3 outcome is
"encouraged": its confidence is nudged up, and it can flip a marginal
RAG-vs-SFT tie toward SFT.
- A choice with >= 3 accepted verdicts averaging < 0.0 outcome is
"discouraged": it is suppressed (e.g. a project that keeps regressing on
GRPO falls back to SFT-on-traces).
The confidence FLOOR is unchanged; only the tie-break + per-choice
suppression shift. Per-project scoping is enforced in
:func:`_summarise_history_outcomes`.
"""
base = _base_verdict(profile, task_category, goal=goal, roi=roi)
if history is None:
return base
bias = _summarise_history_outcomes(history, project=project)
if not bias:
return base
return _apply_history_bias(base, bias)
def _apply_history_bias(
base: Verdict,
bias: Mapping[str, Tuple[float, int]],
) -> Verdict:
"""Adjust a base verdict using per-project history outcomes."""
roi = base.estimated_roi
task_category = base.task_category
# Marginal RAG → SFT flip: strong SFT track record, no comparable RAG
# track record. RAG is the marginal call (it fired on a heuristic
# variance threshold), so prior SFT success is decisive.
if (
base.choice == "RAG"
and _is_encouraged(bias, "SFT")
and not _is_encouraged(bias, "RAG")
):
n_sft = _precedent_count(bias, "SFT")
return Verdict(
choice="SFT",
confidence=min(0.95, base.confidence + _HISTORY_CONFIDENCE_NUDGE),
reason=(
f"Base rubric leaned RAG, but {n_sft} prior SFT verdicts in "
"this project averaged a positive outcome (precedent) — "
"routing to SFT over RAG."
),
reverse_when=(
"the answer space is small and stable and RAG's freshness "
"outweighs the historical SFT lift — re-run after measuring."
),
task_category=task_category,
estimated_roi=roi,
)
# Discouraged GRPO: a project that keeps regressing on RL falls back to
# SFT-on-traces (raises GRPO's effective floor for this project).
if base.choice == "GRPO" and _is_discouraged(bias, "GRPO"):
n_grpo = _precedent_count(bias, "GRPO")
return Verdict(
choice="SFT",
confidence=base.confidence,
reason=(
f"Base rubric leaned GRPO, but {n_grpo} prior GRPO verdicts in "
"this project averaged a negative outcome (precedent) — "
"falling back to SFT on the reasoning traces."
),
reverse_when=(
"a reliable programmatic reward is now available and the "
"earlier GRPO regressions were reward-shaping bugs, not a "
"fundamental mismatch."
),
task_category=task_category,
estimated_roi=roi,
)
# No flip — nudge confidence when the chosen path has a positive track
# record (DPO keeps its own +0.1; we only ever raise, never lower).
if _is_encouraged(bias, base.choice):
return Verdict(
choice=base.choice,
confidence=min(0.95, base.confidence + _HISTORY_CONFIDENCE_NUDGE),
reason=base.reason,
reverse_when=base.reverse_when,
task_category=task_category,
estimated_roi=roi,
)
return base
# ---------------------------------------------------------------------------
# Probe runner (Part B) — heuristic stubs; real model loading is opt-in
# ---------------------------------------------------------------------------

View File

@ -100,6 +100,16 @@ def _project_name() -> str:
return name[:128]
def current_project_name() -> str:
"""Public accessor for the current project label (v0.71.5 #163).
The CLI passes this to ``build_verdict(..., project=...)`` so the history
bias is scoped to the project the verdict was recorded under. Returns the
same string :func:`record_verdict` stamps into the ``project`` field.
"""
return _project_name()
def record_verdict(
verdict: Verdict,
*,

View File

@ -36,6 +36,7 @@ __all__ = [
"DynamicCurriculumPolicy",
"BucketStats",
"compute_bucket_weights",
"percentile_bucket",
"validate_distributed_curriculum",
]
@ -179,6 +180,48 @@ def _softmax(values: Sequence[float], temperature: float) -> List[float]:
return [e / total for e in exps]
def percentile_bucket(
value: float,
window: Sequence[float],
num_buckets: int,
) -> int:
"""Bucket ``value`` by its percentile rank within a rolling ``window``.
v0.71.5 #149 — replaces step-mod round-robin with a difficulty-signal
bucketing for ``curriculum_metric in {loss, perplexity}``. A value at or
above every window member lands in the top (hardest) bucket; a value
below every member lands in bucket 0. Because the bucket is a function of
the value's rank (not the step), a consistently-high-loss sample is
routed to the same bucket on every recompute.
Args:
value: The current sample's difficulty signal (e.g. loss).
window: Recent difficulty signals (rolling reference distribution).
An empty / ``None`` window returns bucket 0 (warm-up — the caller
should fall back to round-robin until the window fills).
num_buckets: Number of difficulty buckets.
Returns:
Bucket id in ``[0, num_buckets - 1]``.
"""
nb = _reject_bool_int("num_buckets", num_buckets)
if nb < _MIN_BUCKETS or nb > _MAX_BUCKETS:
raise ValueError(
f"num_buckets must be in [{_MIN_BUCKETS}, {_MAX_BUCKETS}], got {nb}"
)
fv = _reject_bool_float("value", value)
if nb == 1:
return 0
if not window:
return 0
le = sum(
1 for w in window if _reject_bool_float("window value", w) <= fv
)
rank = le / len(window)
bucket = int(rank * nb)
return min(nb - 1, max(0, bucket))
def compute_bucket_weights(
stats: Mapping[int, Mapping[str, float]],
policy: DynamicCurriculumPolicy,

View File

@ -23,20 +23,22 @@ quant-check thresholds; operators can tune via --threshold.
from __future__ import annotations
import ipaddress
import json
import math
import os
from dataclasses import dataclass
from typing import Iterable, Mapping, Optional, Tuple
from urllib.parse import urlparse
from typing import Iterable, Mapping, Tuple
from soup_cli.utils.paths import is_under_cwd
# Webhook helpers were lifted into the shared ``utils/webhooks`` module in
# v0.71.5 #207 so every production-trace command can offer --slack-url /
# --discord-url. Re-exported here for back-compat (callers + tests that import
# ``validate_webhook_url`` / ``post_webhook`` from ``drift_alarm``).
from soup_cli.utils.webhooks import post_webhook, validate_webhook_url
_MAX_REFERENCE_ROWS = 1_000_000
_MAX_WEBHOOK_URL_LEN = 4096
_MAX_TEXT_LEN = 1_000_000 # 1 MB / row
_LOOPBACK_HOSTS = frozenset({"localhost", "127.0.0.1", "::1"})
# Default smoothing constant for the KL kernel (Laplace-style add-epsilon
# to defend against `log(0)` when a token in `p` is absent from `q`).
_EPS = 1e-9
@ -84,73 +86,6 @@ def validate_threshold(value: object) -> float:
return f_val
def _is_private_or_link_local(host: str) -> bool:
"""Return True iff ``host`` resolves to a non-loopback private/reserved IP.
Explicit parentheses on the final clause (code-review MEDIUM fix
v0.63.0): Python binds `and` tighter than `or`, but the SSRF gate is
safety-critical and a future edit should not need to re-derive the
precedence rules to verify the logic.
"""
try:
ip = ipaddress.ip_address(host)
except ValueError:
return False
return (
ip.is_private
or ip.is_link_local
or (ip.is_loopback is False and (ip.is_reserved or ip.is_multicast))
)
def validate_webhook_url(url: object) -> str:
"""SSRF-hardened webhook URL validator.
Mirrors v0.29.0 `HF_ENDPOINT` / v0.30.0 OTLP / v0.51.0 `validate_hub_endpoint`
policy:
- scheme allowlist {http, https}
- null-byte / control-char rejection
- ``0.0.0.0`` rejected
- plain HTTP only permitted for loopback hosts
- private / link-local / cloud-metadata IPs rejected
"""
if isinstance(url, bool):
raise TypeError("webhook URL must be str, not bool")
if not isinstance(url, str):
raise TypeError(f"webhook URL must be str, got {type(url).__name__}")
if not url:
raise ValueError("webhook URL must be non-empty")
if "\x00" in url:
raise ValueError("webhook URL must not contain null bytes")
if any(ord(c) < 0x20 for c in url):
raise ValueError("webhook URL must not contain control characters")
if len(url) > _MAX_WEBHOOK_URL_LEN:
raise ValueError(f"webhook URL must be <= {_MAX_WEBHOOK_URL_LEN} chars")
stripped = url.rstrip("/")
parsed = urlparse(stripped)
if parsed.scheme not in ("http", "https"):
raise ValueError(
f"webhook URL must use http/https scheme, got {parsed.scheme!r}"
)
if not parsed.netloc:
raise ValueError("webhook URL is missing a host")
host = parsed.hostname or ""
if host == "0.0.0.0":
raise ValueError(
"webhook URL 0.0.0.0 is ambiguous; use 127.0.0.1 or localhost"
)
if parsed.scheme == "http" and host not in _LOOPBACK_HOSTS:
if _is_private_or_link_local(host):
raise ValueError(
"webhook URL plain HTTP is only allowed for loopback; "
"private/link-local hosts require HTTPS"
)
raise ValueError(
"webhook URL for remote hosts must use HTTPS"
)
return stripped
def compute_token_distribution(rows: Iterable[object]) -> Mapping[str, float]:
"""Compute a normalised whitespace-token frequency distribution.
@ -305,39 +240,6 @@ def run_drift_check(
)
def post_webhook(
*,
url: Optional[str],
payload: Mapping[str, object],
timeout_seconds: float = 5.0,
) -> bool:
"""POST ``payload`` as JSON to ``url``. Returns True on 2xx, False otherwise.
Never raises — webhook delivery must NOT crash the drift-check run.
Lazy-imports ``httpx`` so the runtime cost is paid only when an alarm
actually fires.
"""
if url is None:
return False
try:
validated = validate_webhook_url(url)
except (TypeError, ValueError):
return False
try:
import httpx # type: ignore[import-untyped]
except ImportError:
return False
try:
response = httpx.post(
validated,
json=dict(payload),
timeout=timeout_seconds,
)
return 200 <= response.status_code < 300
except Exception: # noqa: BLE001 — webhook must never crash drift check
return False
__all__ = [
"DriftReport",
"compute_token_distribution",

View File

@ -113,8 +113,20 @@ def attach_curriculum_callback(
getattr(tcfg, "curriculum_dynamic_temperature", 1.0) or 1.0
),
)
# v0.71.5 #149 — thread curriculum_metric so the callback can bucket by
# loss / perplexity percentile (round-robin fallback for `length`). Any
# value that is not one of the three valid metrics (e.g. a missing field
# or a test MagicMock) falls back to `length` so the callback always
# constructs.
curriculum_metric = getattr(tcfg, "curriculum_metric", "length")
if curriculum_metric not in ("length", "perplexity", "loss"):
curriculum_metric = "length"
try:
callback = DynamicCurriculumCallback(policy=policy, output_dir=output_dir)
callback = DynamicCurriculumCallback(
policy=policy,
output_dir=output_dir,
curriculum_metric=curriculum_metric,
)
except (TypeError, ValueError) as exc:
logger.debug("attach_curriculum_callback rejected: %s", exc)
return False

View File

@ -27,7 +27,7 @@ import json
import math
import os
from dataclasses import dataclass
from typing import Sequence
from typing import Any, List, Optional, Sequence, Union
from soup_cli.utils.paths import is_under_cwd
@ -35,6 +35,7 @@ from soup_cli.utils.paths import is_under_cwd
_MAX_SCAN_ROWS = 100_000
_MAX_ROW_CHARS = 1_000_000 # 1 MB / row
_MAX_PREFIX_LEN = 100_000 # hard cap on returned prefix length
_MAX_TOKENS_PER_ROW = 50_000 # token-aware mode (v0.71.5 #205) per-row cap
# Tunable: a frequency below this is meaningless (we want a *near-universal*
# prefix). Operator can pick anything in [0, 1] via --min-frequency.
@ -163,11 +164,112 @@ def detect_common_prefix(
return best_prefix
def detect_common_prefix_tokens(
token_rows: Sequence[Sequence[int]],
*,
min_frequency: float,
) -> List[int]:
"""Token-level analogue of :func:`detect_common_prefix` (v0.71.5 #205).
Returns the longest token-id prefix shared by >= ``min_frequency`` of
rows. Operating on token IDs (not characters) guarantees the prefix
always ends on a token boundary — a multi-byte UTF-8 sequence can never
be split mid-code-point.
"""
threshold = validate_min_frequency(min_frequency)
if isinstance(token_rows, (str, bytes)) or not hasattr(token_rows, "__iter__"):
raise TypeError(
f"token_rows must be an iterable of int sequences, "
f"got {type(token_rows).__name__}"
)
materialised: List[List[int]] = []
for idx, row in enumerate(token_rows):
if isinstance(row, (str, bytes)) or not hasattr(row, "__iter__"):
raise TypeError(
f"token_rows[{idx}] must be a sequence of ints, "
f"got {type(row).__name__}"
)
materialised.append(list(row)[:_MAX_TOKENS_PER_ROW])
if len(materialised) >= _MAX_SCAN_ROWS:
break
if not materialised:
return []
if len(materialised) == 1:
return list(materialised[0]) if threshold >= 1.0 else []
need = max(1, int(math.ceil(threshold * len(materialised))))
best_prefix: List[int] = []
sample_templates = materialised[: min(32, len(materialised))]
for template in sample_templates:
lo, hi = 0, len(template)
best_len = 0
while lo <= hi:
mid = (lo + hi) // 2
if mid == 0:
lo = mid + 1
continue
pfx = template[:mid]
count = sum(1 for r in materialised if r[:mid] == pfx)
if count >= need:
best_len = mid
lo = mid + 1
else:
hi = mid - 1
if best_len > len(best_prefix):
best_prefix = template[:best_len]
return best_prefix
def _resolve_tokenizer(tokenizer: Union[str, Any]) -> Any:
"""Return a tokenizer object from a name (lazy AutoTokenizer) or object.
A pre-built tokenizer-like object (duck-typed ``encode`` / ``decode``)
is returned as-is — this is the injectable test seam + lets advanced
callers pass an already-loaded tokenizer. A string is treated as an HF
model id / local path and lazy-loaded via ``transformers.AutoTokenizer``
(so importing this module never pulls transformers).
"""
if hasattr(tokenizer, "encode") and hasattr(tokenizer, "decode"):
return tokenizer
if not isinstance(tokenizer, str):
raise TypeError(
"tokenizer must be a model id / path string or a tokenizer object"
)
if not tokenizer:
raise ValueError("tokenizer name must be non-empty")
try:
from transformers import AutoTokenizer # noqa: PLC0415
except ImportError as exc:
raise ValueError(
"tokenizer-aware prune-prompt needs transformers — "
"install with: pip install 'soup-cli[train]'"
) from exc
try:
return AutoTokenizer.from_pretrained(tokenizer)
except Exception as exc: # noqa: BLE001 — surface a friendly message.
raise ValueError(
f"could not load tokenizer {tokenizer!r}: {type(exc).__name__}: {exc}"
) from exc
def _encode(tok: Any, text: str) -> List[int]:
"""Encode ``text`` to token IDs (no special tokens), capped per-row."""
try:
ids = tok.encode(text, add_special_tokens=False)
except TypeError:
# Tokenizers / fakes without the kwarg.
ids = tok.encode(text)
return list(ids)[:_MAX_TOKENS_PER_ROW]
def prune_traces(
input_path: str,
*,
output_path: str,
min_frequency: float = _DEFAULT_MIN_FREQUENCY,
tokenizer: Optional[Union[str, Any]] = None,
) -> PrunePromptReport:
"""Read a JSONL of {prompt, output} rows, strip shared prefix, write.
@ -176,6 +278,12 @@ def prune_traces(
``prompt`` field (other fields untouched). When no prefix clears the
threshold, the output is byte-identical to the input plus a
``rows_pruned=0`` report.
v0.71.5 #205: pass ``tokenizer`` (an HF model id / local path string, or
a pre-built tokenizer object) to detect + strip the prefix on token
boundaries instead of characters — the prefix can then never end
mid-UTF-8-code-point. Default (``None``) keeps the whitespace-character
behaviour.
"""
threshold = validate_min_frequency(min_frequency)
@ -198,6 +306,10 @@ def prune_traces(
if not os.path.isfile(input_path):
raise FileNotFoundError(input_path)
# Resolve the tokenizer up front (fails fast on a bad name even on an
# empty input file) — None keeps the legacy character path.
tok = _resolve_tokenizer(tokenizer) if tokenizer is not None else None
# First pass: collect prompts (capped).
prompts: list[str] = []
rows_total = 0
@ -233,9 +345,24 @@ def prune_traces(
min_frequency=threshold,
)
prefix = detect_common_prefix(prompts, min_frequency=threshold)
if tok is None:
return _prune_char_level(
input_path, output_path, prompts, rows_total, threshold
)
return _prune_token_level(
tok, input_path, output_path, prompts, rows_total, threshold
)
# Second pass: write output with prefix stripped where applicable.
def _prune_char_level(
input_path: str,
output_path: str,
prompts: List[str],
rows_total: int,
threshold: float,
) -> PrunePromptReport:
"""Character-level prefix strip (the v0.63.0 default behaviour)."""
prefix = detect_common_prefix(prompts, min_frequency=threshold)
rows_pruned = 0
with open(input_path, encoding="utf-8") as fh_in, \
open(output_path, "w", encoding="utf-8") as fh_out:
@ -253,7 +380,6 @@ def prune_traces(
row["prompt"] = row["prompt"][len(prefix):]
rows_pruned += 1
fh_out.write(json.dumps(row, ensure_ascii=False) + "\n")
return PrunePromptReport(
prefix=prefix,
prefix_chars=len(prefix),
@ -263,9 +389,58 @@ def prune_traces(
)
def _prune_token_level(
tok: Any,
input_path: str,
output_path: str,
prompts: List[str],
rows_total: int,
threshold: float,
) -> PrunePromptReport:
"""Token-aware prefix strip (v0.71.5 #205).
The detected prefix is a list of token IDs; stripping a row decodes the
REMAINING token IDs so the boundary is always a clean token break.
"""
token_rows = [_encode(tok, p) for p in prompts]
prefix_ids = detect_common_prefix_tokens(token_rows, min_frequency=threshold)
prefix_text = tok.decode(prefix_ids) if prefix_ids else ""
plen = len(prefix_ids)
rows_pruned = 0
with open(input_path, encoding="utf-8") as fh_in, \
open(output_path, "w", encoding="utf-8") as fh_out:
for line in fh_in:
line = line.strip()
if not line:
continue
try:
row = json.loads(line)
except json.JSONDecodeError:
continue
if not isinstance(row, dict):
continue
prompt = row.get("prompt")
if prefix_ids and isinstance(prompt, str):
ids = _encode(tok, prompt)
if ids[:plen] == prefix_ids:
row["prompt"] = tok.decode(ids[plen:])
rows_pruned += 1
fh_out.write(json.dumps(row, ensure_ascii=False) + "\n")
return PrunePromptReport(
prefix=prefix_text,
prefix_chars=len(prefix_text),
rows_total=rows_total,
rows_pruned=rows_pruned,
min_frequency=threshold,
)
__all__ = [
"PrunePromptReport",
"detect_common_prefix",
"detect_common_prefix_tokens",
"prune_traces",
"validate_min_frequency",
]

View File

@ -0,0 +1,152 @@
"""Shared Slack / Discord webhook helpers (v0.71.5 #207).
Lifts the SSRF-hardened ``validate_webhook_url`` + best-effort ``post_webhook``
out of ``utils/drift_alarm.py`` (v0.63.0 Part E) so every production-trace
command can offer ``--slack-url`` / ``--discord-url`` without re-implementing
the SSRF gate. ``drift_alarm`` now re-exports these for back-compat.
SSRF policy — full parity with v0.29.0 ``HF_ENDPOINT`` / v0.30.0 OTLP /
v0.51.0 ``validate_hub_endpoint`` / v0.63.0 drift-alarm:
- scheme allowlist {http, https}
- null-byte / control-char rejection
- ``0.0.0.0`` rejected
- plain HTTP only permitted for loopback hosts
- private / link-local / reserved / multicast IPs rejected
``post_webhook`` NEVER raises — webhook delivery must not crash the command
that triggered it. ``httpx`` is lazy-imported so the runtime cost is paid only
when an alarm actually fires.
"""
from __future__ import annotations
import ipaddress
from typing import List, Mapping, Optional, Tuple
from urllib.parse import urlparse
_MAX_WEBHOOK_URL_LEN = 4096
_LOOPBACK_HOSTS = frozenset({"localhost", "127.0.0.1", "::1"})
def _is_private_or_link_local(host: str) -> bool:
"""Return True iff ``host`` resolves to a non-loopback private/reserved IP.
Explicit parentheses on the final clause (mirrors v0.63.0 drift-alarm
code-review MEDIUM fix): Python binds ``and`` tighter than ``or``, but
the SSRF gate is safety-critical and a future edit should not need to
re-derive the precedence rules to verify the logic.
"""
try:
ip = ipaddress.ip_address(host)
except ValueError:
return False
return (
ip.is_private
or ip.is_link_local
or (ip.is_loopback is False and (ip.is_reserved or ip.is_multicast))
)
def validate_webhook_url(url: object) -> str:
"""SSRF-hardened webhook URL validator (returns the canonical URL)."""
if isinstance(url, bool):
raise TypeError("webhook URL must be str, not bool")
if not isinstance(url, str):
raise TypeError(f"webhook URL must be str, got {type(url).__name__}")
if not url:
raise ValueError("webhook URL must be non-empty")
if "\x00" in url:
raise ValueError("webhook URL must not contain null bytes")
if any(ord(c) < 0x20 for c in url):
raise ValueError("webhook URL must not contain control characters")
if len(url) > _MAX_WEBHOOK_URL_LEN:
raise ValueError(f"webhook URL must be <= {_MAX_WEBHOOK_URL_LEN} chars")
stripped = url.rstrip("/")
parsed = urlparse(stripped)
if parsed.scheme not in ("http", "https"):
raise ValueError(
f"webhook URL must use http/https scheme, got {parsed.scheme!r}"
)
if not parsed.netloc:
raise ValueError("webhook URL is missing a host")
host = parsed.hostname or ""
if host == "0.0.0.0":
raise ValueError(
"webhook URL 0.0.0.0 is ambiguous; use 127.0.0.1 or localhost"
)
if parsed.scheme == "http" and host not in _LOOPBACK_HOSTS:
if _is_private_or_link_local(host):
raise ValueError(
"webhook URL plain HTTP is only allowed for loopback; "
"private/link-local hosts require HTTPS"
)
raise ValueError("webhook URL for remote hosts must use HTTPS")
return stripped
def post_webhook(
*,
url: Optional[str],
payload: Mapping[str, object],
timeout_seconds: float = 5.0,
) -> bool:
"""POST ``payload`` as JSON to ``url``. Returns True on 2xx, False otherwise.
Never raises — webhook delivery must NOT crash the calling command.
"""
if url is None:
return False
try:
validated = validate_webhook_url(url)
except (TypeError, ValueError):
return False
try:
import httpx # type: ignore[import-untyped]
except ImportError:
return False
try:
response = httpx.post(
validated,
json=dict(payload),
timeout=timeout_seconds,
)
return 200 <= response.status_code < 300
except Exception: # noqa: BLE001 — webhook must never crash the command
return False
def send_webhooks(
payload: Mapping[str, object],
*,
slack_url: Optional[str] = None,
discord_url: Optional[str] = None,
timeout_seconds: float = 5.0,
) -> List[Tuple[str, bool]]:
"""POST ``payload`` to each provided webhook; return per-target delivery.
Returns a list of ``(label, delivered)`` for every non-``None`` URL,
in ``slack`` then ``discord`` order. ``None`` URLs are skipped (not
attempted). Never raises (delegates to the never-raising
:func:`post_webhook`).
"""
results: List[Tuple[str, bool]] = []
for label, url in (("slack", slack_url), ("discord", discord_url)):
if url:
results.append(
(
label,
post_webhook(
url=url,
payload=payload,
timeout_seconds=timeout_seconds,
),
)
)
return results
__all__ = [
"post_webhook",
"send_webhooks",
"validate_webhook_url",
]

1222
tests/test_v0715.py Normal file

File diff suppressed because it is too large Load Diff