* fix(desktop): evict settled session states nothing on screen references Closing a tile never removed its runtime's entry from $sessionStates, so every tile ever closed parked its full transcript in the map for the life of the process. Each leftover entry taxes every subsequent stream flush — the map is spread-copied per delta and the busy/attention/draft projections walk every entry per publish — so the app got slower the longer it ran, which users read as "I need to clean my sessions/dbs". Publish now evicts a settling state when no tile and not the primary view holds its runtime (transition side effects still fire, so the settle keeps its unread dot), and closing a tile drops an already-settled state on the spot. Busy and needs-input states stay: background turns feed the sidebar dots, and a first publish always lands because a resume can publish a beat before the surface binds the runtime. 16 tiles streaming in a 2x2 grid with a day's worth of closed-tile residue: worst-second 34 -> 58 fps, p99 frame 90 -> 28 ms, longtasks 37 -> 0. * perf(desktop): index lineage aliases per sessions-list reference lineageAliases scanned the whole recents list per call, and it is called per cached session state per status projection per message delta — with a populated sessions DB and a few busy sessions that multiplied out to millions of row checks a second during streaming. Build the alias index once per list reference (the list is replaced wholesale, never mutated) and look aliases up in O(1). * perf(desktop): journal each in-flight turn under its own storage key The v1 journal kept every session's tail in one localStorage key, so each throttled write re-parsed and re-stringified EVERY busy session's snapshot — a grid of concurrent streams turned that into a whole-store JSON round trip dozens of times a second, all on the main thread. Per-session keys make a write O(own tail) no matter how many other sessions are streaming. A v1 store migrates on first touch; expired/overflow crash residue is pruned once per renderer. * perf(desktop): stress the multitab scenario across grid/streaming/DB axes The one-stack multitab run hid every cost this round of fixes removed: it drove hook.publish (store only — no journal, no wiring cache), with an empty recents list and no closed-tile residue. Streaming now routes through hook.update (the real gateway write path), and the scenario grows axes for the workloads users actually hit: --zones splits tiles across visible grid zones, --streaming caps how many sessions are mid-turn (zone leaders first), --sessions seeds a lived-in recents list, --dead models settled sessions no surface references. launch.mjs pins HERMES_DESKTOP_CDP_PORT so a non-default --port survives the app's own dev-CDP flag. |
||
|---|---|---|
| .. | ||
| lib | ||
| scenarios | ||
| README.md | ||
| baseline.json | ||
| run.mjs | ||
| serve.mjs | ||
README.md
Desktop perf harness
One systematized way to measure desktop rendering/interaction performance,
diff it against a committed baseline, and fail on regressions. It replaces the
dozen one-off measure-* / profile-* scripts that each reinvented the CDP
client, arg parsing, stats, and output (and never had a baseline).
Quick start
# Isolated instance (recommended) — no running app or LLM credits needed.
# Its own --user-data-dir + HERMES_HOME means it never collides with `hgui`.
npm run perf -- --spawn
# Or: launch an isolated instance once, attach repeatedly (faster iteration).
npm run perf:serve # leaves an instance on :9222
npm run perf # attaches, runs the CI suite, gates on baseline
# One scenario, with a CPU profile:
npm run perf -- stream --cpuprofile --tokens 800
# Representative PRODUCTION numbers (minified React, not the ~3x-slower dev build):
npm run perf -- cold-start stream keystroke transcript --spawn --prod
# Re-capture the baseline on your reference device, then commit baseline.json:
npm run perf -- cold-start stream keystroke transcript --spawn --prod --update-baseline
Dev vs prod
By default the harness measures the dev renderer (fast to spin up, good for
relative regression checks). Pass --prod (with --spawn) to build a
production renderer with the probe included (VITE_PERF_PROBE=1) and measure
minified React — the representative shipped numbers. The committed baseline is
captured with --prod.
Why isolation matters
The measurement this harness exists to run was historically blocked: a running
hgui holds the Electron single-instance lock, so a second instance quit
immediately. --spawn / perf:serve launch with their own --user-data-dir
(separate lock scope), their own HERMES_HOME (separate backend + sessions),
and their own --remote-debugging-port. Synthetic scenarios drive $messages
directly via window.__PERF_DRIVE__, so no LLM credits are spent.
Scenarios
| scenario | tier | measures | replaces |
|---|---|---|---|
stream |
ci | streaming longtasks, frame p95/p99, mutation cadence | measure-synthetic-stream, profile-synth-stream, profile-long-stream |
stream --real |
backend | same, from a real LLM stream | measure-real-stream, profile-real-stream |
keystroke |
ci | composer keystroke → paint latency | measure-latency, profile-typing, leak-typing |
transcript |
ci | large-transcript mount + paint cost | (new) |
render-churn |
ci | per-component render attribution + store churn while N tabs stream | (new) |
idle-cost |
report | busy-but-silent tiles: idle commit rate, + fps while resizing / typing | (new) |
right-pane |
report | file tree + persistent xterm tabs under chat/terminal output and split dragging | (new) |
cold-start |
cold | launch → CDP → driver → first paint (fresh spawn/run) | (new) |
first-token |
backend | Enter → first assistant token painted (TTFT) | (new) |
submit |
backend | Enter → cleared → user msg painted, scroll jump | measure-submit, measure-jump |
session-switch |
backend | route → first-paint → settle | profile-session-switch |
session-load |
backend | how far a session's transcript moves after first paint | (new) |
profile-switch |
backend | rail click → sidebar settled | measure-profile-switch |
ci + cold scenarios need no backend/credits and are gated against
baseline.json (cold-start requires --spawn since it measures a fresh
launch, and must be run in its own invocation). backend scenarios need a live
backend (and --spawn or a real session/credits) and are report-only.
CPU profiling is a cross-cutting --cpuprofile flag on any scenario (it wraps
the run in Profiler.start/stop and prints a top-self-time table), replacing
every standalone profile-* script.
Adding a scenario
Create scenarios/<name>.mjs exporting { name, tier, description, run(cdp, opts) }
where run returns { metrics, detail } (metrics = flat numbers, lower is
better), then register it in scenarios/index.mjs. If it's ci, add a
baseline.json entry (or run --update-baseline).
Layout
lib/cdp.mjs— the one CDP client + target discovery + typing + CPU-profile wrapper + DOM selectors.lib/stats.mjs— percentiles, histograms, CPU-profile self-time ranking.lib/baseline.mjs— load/compare/update the baseline + regression gate.lib/launch.mjs— attach, or spawn a fully isolated instance.scenarios/— one module per measurement.run.mjs— entrypoint.serve.mjs— standalone isolated launcher.
Not migrated (kept as dev utilities)
eval.mjs, reload.mjs, reload-renderer.mjs, probe-renderer.mjs,
probe-thread.mjs, click-session.mjs, diag-*.mjs are interactive dev
helpers, not benchmarks. They can adopt lib/cdp.mjs in a follow-up.