The version details gain an Artifact row (embedded runtime with its
payload tag, or external) and a Runtime row (the tree the backend of
this session actually spawned from). Users copy this when they report
issues, so the two-axis state is visible without a terminal.
The rules come from one source, package.json engines, instead of
copies in the build script. The payload then embeds the EXACT host
versions the gates approved: the node dist is downloaded at the host
node version (and must be an official nodejs.org release), the staged
uv is the host binary, and npm ships inside the node dist. The
installer moves to Node 26 so source installs and embedded installs
run the same node major.
--no-install and --no-package are retired and rejected loudly. A
skipped step is a different artifact, and a different artifact is
not a reproduction. CI drops --no-install; its own npm ci remains
only as a cache warmer with retry protection. The payload stages
(uv python install, pip --target site-packages, node dist) already
ran unconditionally inside the desktop build.
nix run .#build-desktop-app-bundle -- --tag=vX.Y.Z --no-install
The derivation is pure: it wraps the pinned node/uv/git around
scripts/build-bundled-desktop.mjs. The wrapped script runs impurely
on the source tree (payload downloads, electron-builder, codesign).
This replaces the ad-hoc `nix shell nixpkgs#nodejs_22 ...` incantation
and pins the build toolchain to the repo flake.lock.
The dev-tree guard fires in any non-managed checkout — including the
checkout the test runner itself sits in. The update-flow suites mark
their root managed (the guard has its own test file), and the guard
skips an unclassifiable PROJECT_ROOT instead of crashing on test
sentinels.
The old external desktop app installed a git checkout at
$HERMES_HOME/hermes-agent. An embedded app never uses it. Doctor
suggests deletion only for a demonstrably untouched tree (clean
status, on main, no stashes; any probe failure counts as local
work). Doctor deletes nothing itself.
The update flow stashes local changes and moves the checkout to the
update branch. At a managed install root that is the purpose. In any
other checkout it yanks a working tree off its feature branch. The
guard asks first, refuses without a terminal, and honors --yes.
hermes_cli/runtime_tree.py replaces install_manifest.py. A tree with
.git is a git checkout and `hermes update` owns it. A tree without
.git is sealed, and the distribution field of the build stamp names
the steward that replaces it (desktop-app, docker, nix). The refusal
message comes from a per-steward table.
.hermes-install.json dies: staging stops writing it into the payload,
the CLI never reads it, and the update channel lives in config.yaml
(update.channel; main is the default and keeps the current behavior).
Eject gates on Sealed(desktop-app) and is a full handoff: it tells
the user that Setup replaces the desktop app. --channel on a git
checkout writes config instead of a manifest.
Two decisions land together because the code cannot compile between
them:
- Staging has no per-item skip. A stage failure throws and the build
fails. The payload manifest shrinks to a complete-payload sentinel
(schemaVersion 3, tag, commit). External builds write an
external:true stub.
- Backend selection is a constant of the artifact. resolvePayload
requires every runtime item directory; when it resolves, the app
spawns the embedded backend without a look at any checkout. A
payload with no runnable interpreter is a damaged artifact and
raises an error instead of a silent checkout fallback.
decideResidentRuntime, the adoption-era checkout examination, and the
installMode parameter of shouldUseAppUpdater are deleted. The app
self-update gate is now: embedded stamp AND packaged. The update
channel moves to config.yaml (update.channel); Electron mirrors it
with a narrow parser for the version pill. The resident vocabulary is
renamed to embedded; thin builds are now called external.
Docker already stamps distribution:docker in CI, and .dockerignore
already excludes .git from the image. Only the nix stamp lacked the
field. nix/desktop.nix already had it.
The distribution field names who replaces a gitless tree. The desktop
payload writes desktop-app. The CLI reads this value to give the
correct update instruction.
The sandbox (bubblewrap) and the Wayland E2E stack exist on Linux
only. Gate them so the macOS devshell evaluates. Add actionlint for
workflow validation on both platforms.
The renderer speaks in releases on the stable channel and in commits
on the main channel. The statusbar pill shows "(update)" and names the
release tag in its tooltip; a commits-behind count reads as an
alarming +N on a channel where a release is one step. The updates
overlay names the release ("Hermes v0.21.0 is ready to install")
instead of the no-changelog copy, because a release feed carries no
commit rows by design.
The new VersionDetails panel shows version, branch, commit, source,
and distribution from the build stamp on the About page and in the
updates overlay, so support screenshots carry full provenance. The
statusbar tooltip stacks the same details in one panel; TooltipContent
changes from per-line marker chips to a single column panel, and a new
TooltipDetails helper renders muted secondary rows.
gen-share-codes.ts only picks up lint fixes (import order, blank
lines).
A complete payload makes the launch "resident": the backend spawns
directly from the payload in resources, with no materialized checkout
and no bootstrap. The payload CPython resolves imports through its own
hermes-bundle.pth, so the spawn needs no PYTHONPATH and survives
renames, Gatekeeper translocation, and read-only mounts. Writable
state (pycache, lazy installs) goes under HERMES_HOME.
decideResidentRuntime keeps existing users safe: a checkout whose
manifest says source-managed wins, and a pre-manifest checkout with no
desktop marker (the CLI-first cohort) wins too. Desktop-managed trees
go resident; an eject reverses the preference on the next launch.
Bundled installs update through electron-updater and the GitHub
Releases feed instead of git. checkUpdates() maps feed failures to the
same structured error shape as the git path, so an offline check never
surfaces as a raw IPC rejection. The download progress listener comes
off the singleton after each attempt, so a retry cannot stack ghost
listeners. Source installs on the stable channel compare against the
newest release tag, not commits behind main.
stage-agent-payloads.mjs assembles the resources-resident runtime that
ships inside the bundled installer: the repo tree at the release tag
(no .git, with the prebuilt TUI and dashboard JS), a static uv, a
uv-managed CPython, the full site-packages tree from uv.lock, and a
node dist. A hermes-bundle.pth with relative paths makes the payload
interpreter resolve repo/ and site-packages/ wherever the app bundle
sits — no venv, no PYTHONPATH, no absolute paths.
Each CI runner stages natively for its own (os, arch), so there are no
cross-platform wheel-tag tables. Banner probes verify that every staged
binary was built FOR the target: uv prints its build triple, python
reports platform.machine(), node reports process.arch. A wrong-arch
payload fails the build instead of shipping. Packages with no
win_arm64 wheel build from sdist on the arm64 Windows runner; user
machines never compile.
The script stays dormant unless HERMES_DESKTOP_BUNDLED=1, and writes a
thin stub manifest otherwise, so dev builds are unchanged.
scripts/build-bundled-desktop.mjs runs the same sequence locally on
any platform. The desktop-bundled-release workflow builds each
(os, arch) target on a tag push, signs through Azure OIDC (Windows)
and the Apple secrets (macOS) when they exist, and attaches the
artifacts plus the latest*.yml feed files to the GitHub release.
.hermes-install.json marks a checkout as source-managed or
desktop-bundled and records its update channel. A missing file means
source mode on the main channel, so no existing install changes.
`hermes update` reads the manifest:
- On a bundled install it refuses and points at the in-app updater.
- On the stable channel it fast-forwards the checkout to the newest
final release tag (vX.Y.Z, three-digit major cap so legacy CalVer
tags never match) instead of origin/main. The ZIP fallback resolves
the tag through the GitHub API because that path runs when git file
I/O is broken.
- `update.channel` in config.yaml overrides the channel for source
installs. "auto" defers to the manifest.
`hermes update --eject` is the exit from desktop management. On a
bundled install it downloads Hermes Setup and launches it pinned to
the exact commit the bundle was built from; the installer creates a
normal source checkout at ~/.hermes/hermes-agent. Hermes Setup accepts
the new `--pin-commit <sha>` argument for this flow. On a
source-managed install, --eject with --channel only switches the
channel. The "ejected" manageStyle is the permanent opt-out that stops
future auto-adoption.
All packagers (Docker, Nix, desktop) write the same install-stamp.json
with scripts/write_install_stamp.py or with equivalent inline data. The
new hermes_cli/version_info.py reads the stamp first, falls back to
live git for source installs, and reports "unknown" when neither
exists. It caches the result per process.
The stamp replaces three separate provenance paths:
- the HERMES_REVISION env var from the Nix wrapper,
- the .hermes_build_sha file from the Docker build arg,
- live git probes in banner.py and dump.py.
hermes_cli/build_info.py and the desktop's write-build-stamp.mjs are
deleted with them. The desktop build calls the shared Python script.
The banner, `hermes --version`, `hermes dump`, the TUI session panel,
and the desktop About panel now show the same derived version: the
release version, plus "+N" when the build is N commits past the
release tag, or "+?" for a dirty tree with no countable tag. The
release_date field is gone from every surface.
The dirty probe uses `git status --porcelain -uno`: it runs on the
startup-banner path, and an untracked-file scan costs real time on
large checkouts.
release.py now tags each release vX.Y.Z from the package version. The
old CalVer date tags stay readable as history. get_last_tag() prefers
the newest SemVer tag and falls back to the legacy CalVer tags for the
first SemVer release.
__release_rev_count__ records the commit count of the release-bump
commit. Immutable Nix builds carry no git history, so they derive the
"+N commits since release" display from this number and the flake's
revCount.
read_file on a workspace FIFO/socket blocked until the exec timeout —
the existing device guard is name-based (/dev/*, /proc/*) and cannot
see an arbitrary special file. Add _special_file_kind(): one os.stat
on the resolved path, refusing FIFO/socket/char/block devices with a
plain note ('no read was attempted') instead of hanging. Host-visible
filesystems only; regular files, dirs, and missing paths unchanged.
Also adds evals/readtool/: an A/B harness that runs the real AIAgent
against hostile-file fixtures (huge lockfile, one-line bundle, FIFO,
NFD filenames, lying extensions) and measures accuracy, turns, tool
calls, and tokens. Measured for this guard (3 reps, file-only arm):
qwen3.8-max fifo task tokens 122k -> 26k (-79%), turns 9.3 -> 5.0;
opus-4.8 tokens 40k -> 23k; accuracy held 1.00 both arms.
User-visible surface from the #82616 session-continuity campaign:
- sessions.md: 'Repair Stranded Gateway Sessions' (evidence rules,
dry-run-first, why adoption is never automatic) and 'Continuity After
Crashes and Restarts' (atomic identity, self-heal, recency resolution,
reset-boundary fence)
- cli-commands.md: repair-routing row in the hermes sessions table
Docs build verified (en + zh-Hans).
After a partial update (stash restore overwriting providers/base.py with
an older version), the NousProfile singleton was instantiated from a
ProviderProfile class that predates the supports_prompt_cache_key field
(added in f4fb23f3d). Accessing profile.supports_prompt_cache_key raised
AttributeError, crashing every API call with:
'NousProfile' object has no attribute 'supports_prompt_cache_key'
Use getattr(profile, 'supports_prompt_cache_key', False) so a stale
profile degrades to 'no prompt cache key' instead of crashing.
hermes update runs in the PRE-pull Python process. After git pull
updates the source files on disk, sys.modules still holds the OLD
hermes_cli.config and hermes_cli.config_migrations. Function-level
imports return the cached module, so DEFAULT_CONFIG["_config_version"]
is the OLD value and check_config_version() reports (33, 33) —
"up to date" — even though the freshly-pulled code has v34 with a
migration to run.
The personality reset migration (#81946) was silently skipped this
way: display.personality: kawaii stayed active after updates that
should have reset it. Every user who updated from a pre-v34 codebase
to a post-v34 codebase was affected.
Fix: _run_config_check_fresh and _run_migrate_config_fresh call
importlib.reload() on hermes_cli.config_defaults, hermes_cli.config,
and hermes_cli.config_migrations before calling check_config_version
and migrate_config. This forces the modules to be re-read from the
updated source files on disk.
Optional skills merged to main after a user's install was cut were
invisible to 'hermes skills install official/...' until they ran
'hermes update' — the OptionalSkillSource only scanned the local
optional-skills/ checkout.
Now, when an official/<category>/<skill> identifier is not found
locally, OptionalSkillSource resolves it against the live default
branch of NousResearch/hermes-agent: one Trees API call enumerates
optional-skills/*/SKILL.md dirs (cached on disk via the shared index
cache, 1h TTL), then the full skill directory is downloaded byte-exact
(including root-level install scripts, LICENSE, tests/ — files the
generic GitHubSource.fetch path drops). search() and inspect() also
surface remote-only skills so discovery works pre-update too.
Local checkout always wins when present; offline degrades to the old
local-only behavior; traversal and ambiguous bare names are refused;
provenance stays official/builtin.
Review follow-ups on the composite salvage (whole-bug-class sweep):
- session.branch and _persist_branch_seed copied parent history without
display_kind/display_metadata, so a tagged timeline marker (personality
pivot, model switch, auto-continue) re-entered the branched session as a
bare role=user row after a restart — re-planting the phantom-ordinal
class this PR fixes. Both projection dicts now carry the tags; regression
asserts added to both branch tests (mutation-checked: fail without the
fix).
- ui-tui renderer learns display_kind=personality_switch (was falling
through to an opaque user bubble; desktop got the case in commit 1).
- programmatic-integration docs: document the two new 4004 refusals
(boolean ordinal, bare confirm_truncate).
- hermes_state comment: archived rows are searchable only with
include_inactive=True, not by default search — align comment with the
actual FTS filter.
- strip stray trailing blank line in test_tui_gateway_server.py
Two hardening guards extracted from #82766 by @StanleyStetson:
- bool is an int subclass, so a JSON `true` in truncate_before_user_ordinal
coerced via int() to ordinal 1 and aimed a CONFIRMED rewind at the second
user turn — the same silent-loss class as #82756. Reject with 4004.
- confirm_truncate with no truncation target is leaked client rewind state
on an ordinary submit; fail fast with 4004 instead of silently ignoring
the flag, so the corrupted client state is surfaced.
Part of the composite fix for #82756.
Guarding the *aim* of a rewind still leaves every other way of aiming it
wrong terminal. All three reported incidents (#70516, #80763, #82756) ended
at the same write — `replace_messages()` in the `prompt.submit` truncation
path — and all three were unrecoverable for the same reason: the rows are
DELETEd, which also evicts them from the FTS index, so there is no `active=0`
archive and nothing to restore from.
The codebase already draws this distinction and already has the safe half of
it. `archive_and_compact` is documented as "the durability-preserving
alternative to replace_messages"; `rewind_to_message` — the `/undo` path —
soft-deletes to `active=0, compacted=0` and keeps the rows "on disk for audit
/ forensic inspection". The desktop rewind is the same user-facing operation
as `/undo` and was the one taking the destructive branch.
`replace_messages(..., archive_dropped=True)` flips the DELETE to a
content-preserving `UPDATE messages SET active = 0`, reusing the existing
transaction and the existing `active=0, compacted=0` marking so the dropped
turns stay readable via `get_messages(..., include_inactive=True)` and stay
out of session search (`compacted=0` = "the user took it back", vs
compaction's `compacted=1` = "summarized away, still discoverable").
The live transcript is byte-identical either way — only the durability of the
dropped turns changes. The parameter defaults to False, so the fork handler,
the ACP adapter and `gateway/session.py` keep their current semantics
untouched; a test pins that.
`active_only=True` stays on the call: #80216 still applies, and archiving must
not disturb rows an earlier compaction deliberately archived.
Test doubles for `replace_messages` in the gateway suite are widened to the
real signature — they are stand-ins for SessionDB, and a double that does not
accept what production passes silently converts this write into a 5008.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`truncate_before_user_ordinal` is an index into the list of *real* user
turns. The gateway builds that list with `role == "user" and not
display_kind`, and `test_prompt_submit_truncate_ordinal_skips_display_kind_rows`
already pins why: "Without the filter, a trailing marker shifts the ordinal
so the wrong message is targeted for truncation."
`_apply_personality_to_session` broke that invariant at the producer. Its
pivot marker rides as `role=user` — deliberately, so strict
OpenAI-compatible providers accept it mid-conversation (the same reason
`_append_model_switch_marker` does) — but unlike the model-switch marker it
carried no `display_kind`. The gateway therefore counted it as a real user
turn while no client ever renders it as one.
After a personality change the two sides address different lists: every
later rewind/edit/regenerate resolves one slot too early, and
`replace_messages()` hard-DELETEs the extra span. That is the reported
signature — an in-range, valid ordinal, `confirm_truncate: true`, and a cut
that moved backwards with no user rewind action.
Tag the pivot like the model-switch marker, and teach the desktop to
project the kind as a timeline row so a persisted marker is never rendered
— or counted — as a user turn on the client side either. Both ends must
exclude it; excluding it on only one end just inverts the drift.
The regression test drives the real injection point rather than a
hand-written marker dict. Without the fix it fails with "the pivot shifted
the ordinal: the cut landed at 3 instead of 5", losing a turn the user
never asked to drop.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The result-error path in _deliver_result is not inside an except block,
so sys.exc_info() always returns (None, None, None) — the condition was
always False. Simplify to a plain logger.error call with accurate comment.
hermes-cron-tick.service starts without TELEGRAM_HOME_CHANNEL/DISCORD_HOME_CHANNEL
in the unit env; the per-run load_hermes_dotenv reload lived only on the agent
path (after the no_agent short-circuit returns), so every deliver=telegram/all
script job failed with 'no delivery target resolved'. Load the dotenv at the top
of the no_agent branch; override=False keeps the gateway's in-process tick
behavior unchanged.
_normalize_bundle_path rejected absolute paths, .. traversal, and a bare
drive-letter prefix, but permitted a colon inside a later path component.
On NTFS a bundle member named scripts/helper.py:payload writes a hidden
Alternate Data Stream into the visible file scripts/helper.py. The skill
scanner walks with rglob('*'), which does not enumerate streams, so both
operator review and the guard scanner miss the executable bytes.
Reject a colon in any component (the whole class, not just the trailing
one). This subsumes the previous bare drive-letter check, which is folded
into the single colon guard. '/' is the only legal separator once
normalized, so no portable bundle path needs a colon.
Adds an OS-independent quarantine_bundle regression plus a direct
normalizer unit test covering leading/mid/trailing-component colons,
bare/qualified drive letters, and the empty stream name.
Reported-by: JoaoMarcos44 <87440198+JoaoMarcos44@users.noreply.github.com>
Trim verbose comments in conversation_loop.py and run_agent.py to 2 lines
each. Fix the same bug class in the compression summary path at
chat_completion_helpers.py: remove _thinking_prefill from the explicit
pop tuple and move the generic underscore-key sweep to after
_drop_thinking_only_and_merge_users, so the drop pass can recognize
prefill stubs there too.
- warn (not debug) on final text-turn flush failure: a failure here
reopens the exact #81641 data-loss window with _persist_session as
the only remaining retry, unlike the verify siblings which retry
in-loop; include session id for triage
- trim the flush-site comment to sibling proportion, pointing to the
test module for the full incident narrative
- test: assert _persist_session presence before indexing, so a wiring
change fails with a clean assertion instead of ValueError from max()
A pure-text assistant turn (finish_reason=stop) had no durable write of
its own. Its answer reached the user through the streaming / interim
display path, which is display-only and never touches state.db, and the
first durable write was finalize_turn's _persist_session — after the
loop exits and behind post-turn work that can include micro-compaction's
aux-LLM call.
Anything that ended the process or tore the session down inside that
window lost a reply the user had already been shown. On a remote
(non-loopback) backend the window is easy to hit: WS 1006 closures drive
ws_orphan_reap teardown, and affected sessions ended up with user rows
and zero assistant rows in state.db.
The neighbouring exits of the same loop already close this gap:
* the tool-call exit flushes the assistant(tool_calls) block before
handing control to _execute_tool_calls (#49045)
* the verify-on-stop and pre_verify exits flush final_msg before
appending their nudge (#65919 §7)
Apply that same idiom to the ordinary text exit rather than adding a new
persistence mechanism. The intrinsic _DB_PERSISTED_MARKER dedup makes the
later _persist_session a no-op for this row, so no duplicate rows and no
extra write — the same write, just earlier.
Unlike the tool-call exit, a failed flush must not abort the turn: no
side effect runs after this point and the answer is already produced, so
the failure is logged and _persist_session remains the retry.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The live comment poller inferred completion from the job list. An empty
job list looks the same as a finished run: GitHub has not spawned the
jobs yet, so nothing is pending, and the poller posted a final
"all good!" comment and exited.
The run status is now the authoritative signal. collect_run_jobs()
returns whether the CI run and every watched sibling run report
status=completed, and the loop exits only when no job is pending AND
all runs are complete. While a run is still queued or in progress with
no visible jobs, the comment shows "waiting for jobs to start" instead
of a final banner.
Six test files still selected an OS branch with a faked host. Each one now
carries the marker for the host that owns the branch, or derives the
expectation from the real host:
- test_clipboard: macos_only on the has_clipboard_image dispatch. The fake
picked the branch, but _macos_has_image needs osascript.
- test_claw: windows_only on the tasklist/powershell scan, with return_value
in place of a side_effect list that pinned the call count.
- test_linux_desktop_entry: the parametrize over "darwin"/"win32" becomes one
marked test per host. A fake left POSIX paths and a POSIX XDG layout.
- test_graphical_browser_detection: linux_only on the display-server arm. The
$BROWSER check runs before the platform branch, so its test stays unmarked.
- test_auth_nous_provider: the fixture pinned linux so the macOS certifi
fallback could not change the result. The assertion now reads the host, so
the macOS lane covers the fallback too.
- test_tts_macos_output and test_voice_mode: the afplay policy exists because
CoreAudio init raises a TCC prompt, which no Linux runner reproduces.
tests/conftest.py refuses collection when one test carries two OS markers.
Each marker skips on all but one host, so two of them make a test that runs
nowhere while every lane reports green. tests/test_os_marker_gating.py pins
that behavior.
The docstring on TestConfirmDestructiveSlash said the Windows job runs it.
The class has no marker, so -m windows_only deselects it.