diff --git a/CHANGELOG.md b/CHANGELOG.md
index 68d6507..9d04713 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -12,6 +12,51 @@ reproducing 70+ versions of notes.
## [Unreleased]
+## [0.71.5] - 2026-06-02
+
+### Added
+- **`soup eval against` now reads eval metrics** — `ExperimentTracker.get_metric_series`
+ falls back to the `eval_results` table when the metric is not a per-step
+ training column (`loss` / `lr` / `grad_norm` / `speed` / `gpu_mem`). So
+ `soup eval against --candidate --metric task_accuracy` returns a
+ real score series (benchmark scores live in `eval_results`, not `metrics`)
+ instead of "Empty series". Per-step columns still read from `metrics` — no
+ regression for existing callers.
+- **`soup advise` learns from past project outcomes** — `soup advise` now reads
+ this project's accepted-verdict history (`~/.soup/advise_history.jsonl`) and
+ biases the rubric: 3+ successful SFT precedents flip a marginal RAG call to
+ SFT; 3+ negative GRPO outcomes suppress GRPO in favour of SFT-on-traces; an
+ encouraged choice gets a small confidence nudge. Scoped per-project (one
+ project's record never biases another). No history → identical to before.
+- **Slack/Discord webhooks on four more commands** — `--slack-url` / `--discord-url`
+ (SSRF-hardened, loopback-only HTTP, RFC1918 rejected, never crashes the
+ command) now ship on `soup ingest`, `soup prune-prompt`, `soup ab` (fires only
+ on a `reject_h0` / `accept_h0` decision, not `continue`), and
+ `soup data active-sample` — not just `soup drift-alarm`. The validator + sender
+ moved to a shared `soup_cli/utils/webhooks.py`.
+- **Tokenizer-aware `soup prune-prompt`** — `--tokenizer ` detects
+ and strips the shared system-prompt prefix on **token** boundaries instead of
+ characters, so a multi-byte UTF-8 prefix can never be split mid-code-point.
+ Default (no `--tokenizer`) keeps the whitespace-character behaviour.
+- **Curriculum bucketing by loss percentile** — `DynamicCurriculumCallback` now
+ buckets samples by the percentile rank of the live loss (or perplexity)
+ signal within a rolling window when `data.curriculum_metric` is `loss` /
+ `perplexity`, so a consistently-hard sample is routed to the same difficulty
+ bucket across recomputes. `length` and warm-up still use round-robin.
+- **`--hub` on `soup data push` and `soup data forge`** — `soup data push
+ --hub modelscope|modelers` uploads a dataset via the matching SDK
+ (`repo_type=dataset`, commit message sanitised); `soup data forge --hub
+ --teacher owner/name` pre-fetches the teacher model from that hub
+ (and warns when the teacher is not a repo id so `--hub` is never silently
+ ignored). HF stays the default.
+
+### Notes
+- Live SaaS *pull* adapters for `soup ingest` (Langfuse / LangSmith / Helicone /
+ OpenPipe / OpenAI SDKs, issue #204) remain deferred: they need credentialed
+ vendor accounts with populated trace data to validate honestly. Tracked as an
+ open, `infra-blocked` (external-account) item. `soup ingest` continues to parse
+ the JSONL export you pull from your dashboard.
+
## [0.71.4] - 2026-06-02
### Added
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
index 15bf211..2770c1b 100644
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -120,7 +120,7 @@ src/soup_cli/
templates/ - 17 built-in soup.yaml templates (YAML + manifest.json) with load_template loader (v0.39.0, +bco v0.40.0)
ui/ - Web UI (FastAPI + HTML/JS SPA)
-tests/ - Test suite (274 files, 12474 tests)
+tests/ - Test suite (275 files, 12581 tests)
examples/ - Real-world config examples and datasets
```
diff --git a/README.md b/README.md
index 76d4bd9..9b4d16f 100644
--- a/README.md
+++ b/README.md
@@ -49,21 +49,22 @@ infrastructure instead of improving models. Soup fixes that.
## What's New
-**v0.71.4 — Adapter lifecycle + loop wiring.** The merge, PR, and continuous-loop surfaces go live:
+**v0.71.5 — Ingest, data & prompt polish.** Sharper production-loop ergonomics:
-- **Canary verdict on merge** — `soup adapters merge … --canary suite.json` scores the merged
- adapter and reports **OK / MINOR / MAJOR**; `--strict-verdict` exits non-zero on a MAJOR
- regression. Works with no model load using a pre-scored canary suite.
-- **Evolutionary merge for real** — `soup adapters merge --strategy cmaes --eval suite --budget 1h`
- now runs the full CMA-ES search (merge → score → optimise) and writes the best blend, instead of
- just printing a plan.
-- **Publish an adapter PR** — `soup adapters pr --base-sha --adapter --push
- owner/repo#42` posts the rendered PR straight to a GitHub PR comment.
-- **Continuous fine-tuning loop, wired up** — `soup loop watch --pre-wired` runs the real
- traces → DPO → eval-gate → canary pipeline; `--pack-cans` snapshots every iteration as a
- shareable Soup Can with Registry lineage (`soup loop replay --extract dir`).
-- **Branches ↔ Registry** — `soup adapters branch --attach-to-registry ` /
- `--from-registry ` links training-env snapshots into the Registry lineage DAG.
+- **Alternative model hubs for data** — `soup data push --hub modelscope|modelers` uploads a
+ local JSONL to ModelScope / Modelers, and `soup data forge --hub … --teacher owner/name`
+ pre-fetches the teacher from that hub.
+- **Tokenizer-aware prompt pruning** — `soup prune-prompt --tokenizer ` finds the
+ shared *token* prefix and decodes only the remainder, so BPE multi-byte sequences never get
+ truncated mid-token the way char-slicing can.
+- **Webhooks everywhere** — `--slack-url` / `--discord-url` now work on `soup ingest`,
+ `soup prune-prompt`, `soup ab`, and `soup data active-sample` (same SSRF-hardened validator as
+ `soup drift-alarm`). The A/B harness only pings when the sequential test actually decides.
+- **Curriculum by difficulty percentile** — dynamic curriculum can bucket by `loss` /
+ `perplexity` percentile instead of length round-robin.
+- **Smarter pre-flight `advise`** — `soup advise` now nudges its confidence using your prior
+ verdicts for the same project, and `soup runs replay` can plot a benchmark-score curve, not
+ just the loss curve.
Full history: [CHANGELOG.md](CHANGELOG.md) · [GitHub Releases](https://github.com/MakazhanAlpamys/Soup/releases).
diff --git a/docs/backends-and-ops.md b/docs/backends-and-ops.md
index 3caf16d..d6d0d98 100644
--- a/docs/backends-and-ops.md
+++ b/docs/backends-and-ops.md
@@ -459,6 +459,11 @@ Every completed run also stores an estimated cost (`$` per run) computed from th
captured GPU device name and duration. `soup runs show` renders `—` for CPU /
MPS / unknown GPUs (no fabricated zeros).
+As of v0.71.5, the metric-series lookup that powers replay (`ExperimentTracker.get_metric_series`)
+transparently falls back to the `eval_results` table when a metric has no per-step
+rows — so you can plot a benchmark-score curve (e.g. `mmlu`, `gsm8k`) the same way
+you plot `loss`, without caring which table holds the series.
+
### Tracker integrations (--tracker mlflow / swanlab / trackio)
```bash
diff --git a/docs/commands.md b/docs/commands.md
index 8d192b5..7d9a915 100644
--- a/docs/commands.md
+++ b/docs/commands.md
@@ -90,10 +90,12 @@ soup data download user/ds --samples 1000 Stream first 1000 samples
soup data register --name my-ds --path d.jsonl --format alpaca Register dataset
soup data unregister --name my-ds Remove from registry
soup data push --input d.jsonl --hf-dataset user/name Upload local JSONL as HF dataset
+soup data push --input d.jsonl --hf-dataset u/n --hub modelscope|modelers Upload to an alternative hub
soup data registry List all registered datasets
soup data demo List bundled demo JSONL fixtures
soup data demo alpaca_demo --output ./d.jsonl Copy a bundled demo JSONL fixture
soup data forge --docs ./docs --task sft --target-rows 1000 Synthetic data pipeline + provenance
+soup data forge --docs ./docs --hub modelscope --teacher owner/name Pre-fetch the teacher from an alternative hub
soup data score --input rows.jsonl Composite quality scorecard (PII + toxicity + lang + edu)
soup data decontaminate --input rows.jsonl --benchmarks mmlu,gsm8k Drop benchmark-overlap rows
soup data toxicity --input rows.jsonl -o tox.jsonl Flag toxic rows (keyword baseline)
@@ -137,7 +139,7 @@ soup can publish r.can --hf-hub user/name Publish .can to HF Hub as dataset
soup runs List training runs
soup runs show Run details + loss graph + cost
soup runs compare Compare two runs
-soup runs replay Replay summary + loss curve from history
+soup runs replay Replay summary + loss curve from history (also plots a benchmark-score curve when the metric lives in eval_results)
soup why [run_id] Explain training anomalies (heuristic)
soup tui Full-screen Textual dashboard (requires [tui] extra)
soup train --config soup.yaml --profile Record torch.profiler trace to