project-nomad/admin/inertia/components/benchmark
chriscrosstalk 39d0acbe92
fix(benchmark): partial runs are not the NOMAD Score (relabel + renormalize) (#1088)
A System-Only or AI-Only run is a partial result, not the NOMAD Score
(which is the full-benchmark composite). Two problems addressed:

1. Scoring bug: AI-only runs were NOT renormalized. _calculateNomadScore
   always added the system weights (0.60) to the denominator even for an
   AI-only run (default-zero system scores), so an excellent AI-only run
   scored ~39.8 -- scaled against the full NOMAD 100 where AI is only 40%
   -- while system-only already renormalized correctly. Fix: pass
   systemScores only when the system benchmarks actually ran, so AI-only
   renormalizes to its own 0-100 (39.8 -> ~99.7). Full and System-only
   scores are unchanged.

2. Presentation: partial runs were shown with the full "NOMAD Score"
   label + big gauge, outweighing the small "Partial" notice. Now partial
   runs are relabelled "System Score" / "AI Score" with a PARTIAL badge,
   a muted (neutral) gauge + number, and a "run a Full Benchmark for your
   NOMAD Score" CTA -- applied to both the persistent score section and
   the Phase 3 ScoreReveal via a shared getScoreDisplay() helper. Adds a
   `muted` prop to CircularGauge.

Browser-validated on NOMAD3: AI-only now shows "AI Score" + PARTIAL,
muted gauge, 99.7 (was 39.8); Full still shows "NOMAD Score" green.
Implements the display-layer fix from the Score v2 red-team's W2; the
AI-ceiling saturation (W1) remains v2 work.

Part of #1082.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 13:17:35 -07:00
..
BenchmarkRunView.tsx feat(benchmark): end-of-run score reveal + NVIDIA GPU-util overlay (#1087) 2026-07-20 12:07:45 -07:00
CoreGrid.tsx feat(benchmark): live telemetry during benchmark runs (#1082) (#1084) 2026-07-19 23:28:07 -07:00
LiveReadout.tsx feat(benchmark): live telemetry during benchmark runs (#1082) (#1084) 2026-07-19 23:28:07 -07:00
ResultsSoFar.tsx feat(benchmark): authoritative in-test sysbench numbers + results strip (#1085) 2026-07-20 11:58:57 -07:00
ScoreReveal.tsx fix(benchmark): partial runs are not the NOMAD Score (relabel + renormalize) (#1088) 2026-07-20 13:17:35 -07:00
Sparkline.tsx feat(benchmark): live telemetry during benchmark runs (#1082) (#1084) 2026-07-19 23:28:07 -07:00
StageRail.tsx feat(benchmark): live telemetry during benchmark runs (#1082) (#1084) 2026-07-19 23:28:07 -07:00