Models sometimes paste old_string (and new_string) straight from
read_file/search_files output, which prefixes every line with the
'N|' gutter. The file has no such prefixes, so all 9 fuzzy strategies
fail and the model burns turns re-reading and re-patching.
When every non-empty line of old_string carries a uniform,
mostly-consecutive 'digits|' prefix — the same verbatim-paste signature
write_file already rejects via _looks_like_read_file_line_numbered_content
— strip the gutter from both strings and retry the strategy chain once.
Mixed prefixes, non-consecutive numbers, or a single numbered line are
treated as genuine content and left alone. An exact literal match always
wins first, so files that really contain gutter-shaped text are unaffected.
Strategy name is reported as '<strategy>+gutter_stripped' for telemetry.
Sabotage-verified: recovery tests fail without the fix.
Follow-up to #82179 addressing helix4u's review comment
(#82179 issuecomment-5229441571). Three parts:
1. Desktop teardown (salvaged from #77436, @4adwentures): the update
hand-off's releaseBackendLock() sent SIGTERM to the primary backend
BEFORE taskkill /T. If the launcher exits first, Windows can no longer
enumerate its descendants and they survive holding the venv — the
Electron path that creates the orphan #82179 then has to repair.
New stopBackendTreesForUpdate() tree-kills the live root first, with
the behavioral vitest from #77436. The scanner half of #77436 is
deliberately NOT taken (superseded by #82158's full-cmdline scan).
2. Tree-aware orphan classification: _orphaned_desktop_backend_pids()
previously refused the whole holder set when any holder had a live
parent. But the scanner legitimately returns an orphaned serve root
AND its descendants (the venv trampoline's uv-managed interpreter
worker — which carries the same backend argv — plus .hermes-runtime
children). Those have a live parent: the orphan root itself. Now
holders inside an accepted orphan root's tree fold into that root
(only roots are returned; taskkill /T reaps descendants), and
live-parent backends defer to the ancestry check instead of refusing
outright. Anything outside an orphan tree still refuses.
3. Tests for the mixed shapes: root+managed-runtime child,
grandchild depth, non-descendant stray alongside an orphan root
(still refuses), descendant exited mid-classify.
E2E on a real Windows box: spawned a detached backend-shaped orphan
that itself spawned children (3 python descendants); the scanner-shaped
mixed holder set classified to [root], taskkill /T reaped root and all
descendants. The live Desktop backend on the box still classified None
(refusal preserved). The first E2E attempt caught exactly the
trampoline/worker case the mocks missed — the live worker re-execs with
the same backend argv and a live parent — which is what part 2 fixes.
Co-Authored-By: 4adwentures <296413879+4adwentures@users.noreply.github.com>
The GUI-updater handoff race: the Desktop fires SIGTERM + app.quit() and
spawns hermes-setup, but its Python backend (`python.exe -m
hermes_cli.main serve`) can survive the teardown. The Desktop is gone --
nothing will respawn that backend -- yet the venv-holder guard refused on
it and the update dead-ended with "Hermes is still running" while the
user had zero windows open (observed twice on 2026-08-09, 01:59 and
02:17, bootstrap-installer.log).
New `_orphaned_desktop_backend_pids()` classifies remaining holders: a
serve/dashboard backend whose supervising parent is provably dead (PID
gone, or recycled -- parent created after the child) is a straggler safe
to reap. Any live-parent backend, non-backend holder, or unprovable case
keeps the refusal exactly as before. Reaping uses the new
`_stop_process_trees()` (taskkill /T /F), mirroring the Desktop's
forceKillProcessTree and install.ps1's venv sweep so the managed
.hermes-runtime interpreter child dies with its launcher (#70026).
Builds on #81327 (salvaged intact underneath): that fixed the same
parent-only-kill gap in install.ps1's venv sweep; this closes the
remaining dead-end in the `hermes update` guard itself.
E2E on a real Windows box: spawned a detached orphan with a
backend-shaped argv -> classifier returned its PID and the tree reap
killed it; a non-backend orphan and the live Desktop backend (parent
alive) both returned None (refusal preserved).
Both halves of this bug were the same failure mode: allowScripts is keyed by
exact name@version, so an entry stops matching the moment a dependency moves
and npm demotes the blocked script to a warning nobody reads. The breakage
surfaces much later as a missing native artifact on one platform.
Assert the two relationships that make the allowlist meaningful — every
versioned pin resolves to a version the lockfile installs, and every package
the lockfile marks as having an install script carries a decision. A
bare-name key stays exempt from the version check so a standing denial like
unicode-animations survives bumps.
Lives in tests-js because the CI change classifier does not run the Python
suite for a manifest-only diff.
Fixing the allowlist only helps a fresh install. npm will not re-run an
install script for a package already on disk, so every checkout that
installed while get-windows was blocked stays bricked: `hermes update`
pulls the fix, `npm install` skips the script, and the build fails on the
same missing binding.
Run `npm rebuild get-windows` from the staging step when the binding is
absent, and if that still yields nothing, print the two commands that
recover the checkout by hand instead of the previous advice to reinstall
dependencies, which is exactly what the user already tried. Gated to a
win32 host building for win32, since no other host can produce the binding.
Co-authored-by: JoaoMarcos44 <JoaoMarcos44@users.noreply.github.com>
get-windows was added to apps/desktop without a root allowScripts entry, so
npm blocked its node-pre-gyp install script and the win32 prebuilt binding
was never downloaded. Every Windows desktop build then died in
stage-native-deps, including the updater's headless rebuild, leaving Windows
users unable to update the app.
The same manifest had drifted twice more: the CVE sweep in 7537de9e74 moved
Electron to 40.10.6 and left the electron@40.10.2 pin behind, and website's
allowlist still names core-js-pure, which its lockfile no longer resolves,
while fsevents runs an install script with no entry at all.
Co-authored-by: gsy324 <gsy324@users.noreply.github.com>
Co-authored-by: elbukott1 <elbukott1@users.noreply.github.com>
Co-authored-by: Brian Franco <BrianFranco@users.noreply.github.com>
The classic CLI's light-mode detection sends an OSC 11 background-color
query and blind-waits 100ms. Terminal managers that swallow OSC 11
(herdr) made every startup pay the full 100ms for nothing, and any
in-order relay that answers slower than 100ms (SSH bridges, WSL,
loaded tmux servers) delivered the reply AFTER prompt_toolkit owned
the tty — the rgb:.../escape payload leaked into the input line as
gibberish characters.
Fix: send the OSC 11 query followed by a DA1 sentinel (ESC [ c) in one
write — the same fence pattern the Ink TUI's TerminalQuerier uses.
Terminals answer queries in order and effectively all of them answer
DA1, so the DA1 reply proves the terminal has already processed (or
ignored) our OSC 11. Fast terminals and herdr-style multiplexers now
resolve in ~1ms; slow relays get their reply consumed instead of
leaked; a hypothetical DA1-mute terminal falls back at a 1s safety
net, same clean timeout path as before.
Adds real-PTY regression tests covering the herdr-style (DA1-only),
slow-relay (+300ms reply), and fully mute emulator behaviors, each
asserting zero leftover bytes in the tty buffer. Sabotage-verified:
the slow-relay test fails against the old un-fenced code with the
exact leak payload in LEFTOVER.
On slow terminals (VPS, containers under load), the OSC 11 background
color response can arrive after TCSAFLUSH completes — leaking into
prompt_toolkit's input buffer and silently consuming the first 1–3
characters of every response.
Add a 50ms post-flush drain window that reads and discards any late
bytes via select() + os.read() before prompt_toolkit grabs the tty.
Fixes#40250
Adds the failed-result guard the salvage review called for:
_deliver_queued_first_response now takes deliver_media and the queued
follow-up call site passes deliver_media=not _delivery_result.get('failed').
A failed turn still delivers its normalized failure text (pinned by
test_run_agent_sends_normalized_failure_before_queued_followup), but its
attachments are no longer uploaded as if the turn succeeded — mirroring
the completed-turn path's 'not agent_result.get(failed)' guard.
Regression test added.
Ensure queued follow-up resends keep MEDIA-backed attachments by replaying the
first response through the gateway's text-plus-media delivery flow instead of a
plain adapter text send.
Closes the residual the contributor's own triage comment flagged: % was
excluded from the special-char class to protect the CJK LIKE fallback,
but a non-CJK query never reaches that fallback (is_cjk gates it), so
'50%' still hit MATCH raw and silently returned zero results. Strip %
whenever the sanitized query contains no CJK; the CJK path keeps its
pre-existing contract. Regression tests for both directions.
_sanitize_fts5_query's strip step only removed +{}():"^ . Every other
character FTS5's grammar rejects outside a quoted phrase reached MATCH
raw and raised, and — as the step's own comment says about the colon it
was fixed for — the execute site swallows that into zero results. Session
search silently found nothing for ordinary queries:
it's fts5: syntax error near "'"
gateway/run.py fts5: syntax error near "/"
user@host fts5: syntax error near "@"
a,b fts5: syntax error near ","
why? fts5: syntax error near "?"
e=mc2 fts5: syntax error near "="
Complete the class and assemble it with re.escape, because written as a
regex literal the backslash was eaten as an escape and never made it in
(C:\path\file still raised after the first pass).
Measured against a real FTS5 table over 651 realistic queries:
373 unparsable before, 77 after. The remainder is leading/trailing "." and
"-", which #43889 already covers.
% is deliberately left in: the CJK path falls back to a LIKE search that
needs it literal and escapes wildcards itself, so stripping it widened
those queries onto unrelated rows (test_cjk_like_escapes_wildcards).
Extends the picker fix to the read that actually uses the key: switch_model's
user-provider credential resolution (the ${VAR} api_key expansion and the
key_env fallback at the resolve-credentials step) still read os.environ raw
and passed the result to resolve_runtime_provider as explicit_api_key — so
under multiplex_profiles the actual switch, not just the picker listing,
could adopt another profile's key. Same _scoped_key_env helper, same
fail-closed semantics; identical behavior when multiplexing is off.
Adds end-to-end switch_model tests pinning that an installed scope wins over
the process environment for both read shapes.
854007d1c routed the remaining main-agent fallback key reads through
agent.secret_scope so the multiplexed gateway's per-profile scope applies.
list_authenticated_providers - which gateway/slash_commands.py calls
directly for /model - still resolved custom-endpoint and fallback-entry
credentials with raw os.environ.get(key_env), so under multiplex_profiles
one profile's picker reads whatever key the process environment happens to
hold, i.e. another profile's.
no multiplexing : profileA-key (unchanged)
scope installed : profileB-key (was profileA-key)
Route both reads through a _scoped_key_env() helper over
secret_scope.get_secret(). get_secret is identical to os.getenv when
multiplexing is off, so single-profile deployments are byte-for-byte
unchanged; a fail-closed UnscopedSecretError is treated as "no credential
visible for this profile", which is how the picker already handles a
missing key.
Scope: only the two key_env credential reads. The other environment reads
in that function are provider-presence probes (AWS creds, LM_BASE_URL),
a separate concern.
Per-turn .env adoption could rewrite agent.api_key while leaving
_credential_pool_entry_id on a previously rotated fallback. The next 429
then marked the healthy fallback exhausted via credential_id precedence
(#79156).
- Sync pool entry id after a successful env credential refresh
- First look does not stomp a pool-rotated key with the env primary
- mark_exhausted_and_rotate prefers api_key_hint when it disagrees with
credential_id
Fixes#79156
_detect_venv_python_processes() returned cmdline_raw[:120]. Gateways
autostarted via the managed-runtime interpreter carry a >120-char exe path
(.hermes-runtime\python\generation-...\cpython-3.11-...), so the truncated
cmdline ended inside the exe path, before '-m hermes_cli.main gateway run'.
The Desktop preflight's pausable-gateway exemption
(_scan_venv_blockers._is_pausable_gateway) therefore never matched, the
gateway was reported as a blocker, and every Desktop update aborted with
'Update didn't finish' even with all windows closed — the updater's own
gateway pause never got a chance to run.
Fix: return the full cmdline from the detector and truncate only at
display time (_format_venv_python_holders_message and the scan's JSON
cmdline field, after redaction).
Reproduced live on Windows 11: scan reported blocked=true for
'...cpython-3.1' (truncated); after the fix the same gateway pair scans
clear with pausable_gateways=2.
An alias that family-matches multiple catalog models (/model opus) used to
silently pick one via _model_sort_key heuristics. The heuristics have
guessed wrong repeatedly — dated snapshots like claude-opus-4-20250514
parsed as version 20,250,514 and outranked claude-opus-4-8; suffix
tiebreaks landed on the cheapest tier — and every wrong guess silently
switches the user to a model they did not ask for.
resolve_alias now raises AmbiguousAliasError whenever more than one model
matches the alias family; switch_model catches it at all three call sites
(explicit-provider path, current-provider path, authenticated-provider
fallback) and returns a failure result listing the candidates
(best-guess-first ordering, capped at 10) with instructions to pick an
exact name. A single match still resolves automatically, and DIRECT_ALIASES
exact mappings are unaffected.
The date-stamp split from #67571 is kept, demoted from selection logic to
display ordering of the candidate list.
Supersedes the auto-pick approach of #67571; credit to @Sahaun and @GottZ
for the date-stamp parser analysis that this builds on.
_model_sort_key treated YYYYMMDD snapshot stamps (e.g.
claude-opus-4-20250514) as version components, so 20250514 > 8
and resolve_alias("opus", "anthropic") returned the wrong model.
Fix: split components ≥ 19_000_101 (smallest plausible date stamp)
out of the version tuple, keeping them as a trailing tiebreaker so
bare IDs sort before their dated snapshots and newer snapshots
before older ones. Shorter numeric components (mistral-large-2411,
gpt-4-0613) keep their current behavior. No models.dev dependency
in the sort path.
_resource_attributes() in otlp_exporter.py built its own hardcoded
resource dict (service.name/instance.id/telemetry.scope only) instead
of reusing gateway_health_export.py's _runtime_resource_attributes(),
which already applies the resource_attributes allowlist from config.
Result: operator-configured attributes like deployment.environment.name
reached metrics and diagnostic logs but never spans.
Span resource building now delegates to the same
_runtime_resource_attributes() helper metrics/logs already use,
removing the duplicate implementation instead of patching it in place.
Generic thinking fields (reasoning / reasoning_content + the
reasoning_details text charge) are replayed for at most the NEWEST
assistant turn on every transport: Anthropic strips all-but-newest at
convert time, Bedrock Converse never replays thinking, and strict
chat-completions providers reject or one-space-pad the field. The tail
budget walks charged them on every message anyway, spending 19-24% of
the budget (per the issue's 1,025-message measurement) on bytes that
provably never reach the wire — so the tail cut landed early and each
compaction discarded more real transcript than configured.
_estimate_msg_budget_tokens now partitions the replay keys:
* _ALWAYS_REPLAYED_BUDGET_KEYS (codex_reasoning_items,
codex_message_items) — charged unconditionally. These ride the wire
on every retained turn (#55572), and codex_reasoning_items now also
carries native server-side compaction checkpoints (#81747).
* _NEWEST_TURN_ONLY_BUDGET_KEYS (reasoning, reasoning_content) + the
reasoning_details text charge — charged only for the newest assistant
turn via charge_stale_thinking, resolved by the three budget walks
(tail cut, raw-budget re-walk, proactive-prune boundary).
Default stays the conservative full charge for callers without
turn-position context. A partition invariant test pins that any future
_REPLAY_BUDGET_KEYS entry must be classified into exactly one class.
Direction credit: #73669 (@x7peeps) and #73730 (@webtecnica) both
attacked this; the keep_open reviews asked for provider/API-mode-aware
accounting that keeps Codex carriers charged — this implements that
shape.
get_messages() only deserializes content and tool_calls; the structured
reasoning columns (reasoning_details, codex_reasoning_items,
codex_message_items) come back as the raw TEXT they were stored as.
Feeding those rows straight back into a write, which is exactly what
the POST /api/sessions/{id}/fork handler does by piping get_messages()
into replace_messages(), hit an unguarded json.dumps() and stored the
already-serialized string encoded a second time. On replay of the fork,
json.loads() then yields the inner string instead of a list, and every
consumer's isinstance(..., list) gate silently drops it: preserved
Anthropic thinking blocks, Codex encrypted-reasoning/message-item
replay, and OpenRouter multi-turn reasoning context are all lost after
a fork, with one more encoding layer added per fork.
The /branch copy loop had the same defect from the other side: it
forwarded reasoning but none of the structured columns, and both TUI
branch writers persisted role/content alone, dropping reasoning and
reasoning_content along with them.
Route the six dumps sites in append_message and _insert_message_rows
through a shared guard that keeps already-serialized strings as-is;
structured values from the live runtime are dumped exactly as before.
Forward the reasoning fields in all three branch writers, matching the
set gateway/slash_commands.py already forwards on its own /branch path.
Field renders its visible label as <label htmlFor>, but that only
associates with labelable elements (input, select, ...), not the
div[role="group"] DeliverCheckboxes renders. The checkbox group had
no accessible name for assistive tech.
Field now stamps an id on its label (`${htmlFor}-label`) and
DeliverCheckboxes references it via aria-labelledby, so the group
picks up the same visible label text instead of duplicating it.
Addresses review feedback on PR #73886.
The band is a few lines of conversation floating over another app, not a
transcript. Tool blocks, file-diff panels, and background notices ("Self-
improvement review: patched ...") pushed the actual answer out of the
capped band and read as junk pinned over the window below.
The HUD-mode note tells the model that an unqualified "this" means the app
behind the strip. It says nothing about the app that was behind it a minute
ago, and the user drags the strip from app to app mid-thought: parked over
Spotify, "pause that and play X here" is one request spanning two apps, and
only the second half has a window under it.
Those earlier windows are already in context as read_window_below results, so
the note only has to say they still count. Without it the latest window reads
as the only one and half the request is silently dropped.
No new tool names, so the existing gating tests cover it unchanged.
Document the actual transport split: legacy editable drafts by default, optional rich drafts, persistent rich final sends, and in-place rich final edits for edit-based streams.
Anything longer reads as the band waiting for something. Also drops the note
claiming a sub-second hold disappears into the fade — that was written when
losing window focus never started the hold at all, so what looked like too
short a stage was no stage.
A hairline shadow under the bar's bottom edge. Without it the HUD reads as
pasted onto the other app rather than floating above it; kept tight and faint
so it never becomes a glow around a card.
The composer's completion list — `/`, `@`, `:`, and the help hint — hangs off
the top of the bar. That is right everywhere the composer has a window above
it, and wrong in the one place the bar is parked against the screen's top edge:
the list rendered off screen, so `/` looked like it did nothing at all.
Flip it below the bar in that orientation and cap it to the room it actually
has, since the app's own cap assumes a full window.
The list hangs over the band, and both were half-lit over a third thing — the
app underneath. No pair of opacities reads well in that stack.
So the list goes fully opaque and the band falls back behind it: dimmer,
fractionally smaller, slightly out of focus, scaled from the bar's edge so it
reads as depth rather than as the panel shrinking.
Apply the same metadata invariant to sealed split chunks, keep expect_edits on live previews, and exercise the real Telegram adapter path with rich messages enabled but rich drafts disabled.\n\nCredits PR #78525 by @Slobaka for the reproduced rich_messages/rich_drafts combination.
A finalized native draft is the first persistent send and will not be edited again. Omitting expect_edits lets Telegram use sendRichMessage for the persistent final instead of degrading tables through MarkdownV2.\n\nAdapted from PR #46536.
Clicking away to another app is the commonest way the HUD gets let go of, and
it fires no focusout — the composer stays document.activeElement while the
window is inactive. Chrome stops matching `:focus` on an unfocused window all
the same, so the band lost its focus state with no hold running and snapped
shut instead of stepping down to the glanceable stage.
The sheet also goes heavier than the text in front of it. There is no blur to
separate the band from what it lies over, so it is the only thing keeping
half-opacity text off someone else's UI.