Commit Graph

2696 Commits

Author SHA1 Message Date
Jeffrey Quesnelle daf67f2e59
Merge branch 'main' into feat/hermes-relay-tool-metrics 2026-08-04 12:43:38 -04:00
kshitij 1709f82c15 chore: remove dead tomli dependency declaration
requires-python is >=3.11 so tomllib is always in stdlib; the
tomli fallback branch in _lint_toml_inproc was unreachable. Removes
the dependency from pyproject.toml + uv.lock and deletes the dead
try/except ImportError fallback in the code.
2026-08-04 14:34:24 +05:30
kshitijk4poor 8c19e29259 refactor(file-ops): fold simplify-pass findings
- write_file: encode content once, share bytes between bytes_written and
  the sha256 verification (drops a second full-content encode per write)
- patch_parser: replace the except-TypeError retry around
  write_file(pre_content=...) with signature-based feature detection so a
  TypeError raised inside a capable implementation propagates instead of
  triggering a duplicate write; tests for both duck-typing contracts
- tests: real-ops V4A BOM round-trip + _file_has_bom disk-probe guard
  (the teknium1-review regression previously only covered by a fake)
- comment: document dirs_created's long-standing "parent ensured" meaning
2026-08-04 14:34:24 +05:30
kshitijk4poor fcae5ad49a fix(file-ops): surrogatepass in bytes_written encode (review finding)
Content that flowed through a surrogateescape decode (backend output via
patch_replace) can carry lone surrogates; a strict encode raises
UnicodeEncodeError where the old wc -c path could not. Mirrors the
existing sha256 verification encode.
2026-08-04 14:34:24 +05:30
阿泥豆 eb78ab235f fix(file-ops): decouple BOM detection from pre_content, add V4A backward compat
Bug 1 (UTF-8 BOM loss on V4A UPDATE):
_file_has_bom() trusted pre_content for BOM detection, but the most
common pre_content provider — read_file_raw() — deliberately strips
BOMs so the agent never sees U+FEFF glyphs.  Passing BOM-stripped
content through pre_content caused a false-negative: the method
returned False and write_file() silently removed the marker on rewrite.

Fix: _file_has_bom() now always probes the first 3 bytes on disk
(head -c 3), ignoring pre_content for BOM purposes.  pre_content is
still used by two other consumers — line-ending detection and lint/LSP
delta computation — neither of which is affected by BOM stripping.

Bug 2 (backward compatibility):
_apply_update() called write_file(path, content, pre_content=...) as a
keyword argument.  Duck-typed file_ops implementations that only
implement the two-argument write_file(path, content) contract would
raise TypeError.

Fix: wrap the call in try/except TypeError, falling back to the
two-argument form when the keyword is not accepted.

Also declare tomli in pyproject.toml (pre-existing conditional import
for pre-3.11 Python, caught by the pre-commit dep scan after staging
file_operations.py).

Tests:
Add TestV4ABomRoundTrip with two cases:
  - UPDATE on BOM-bearing file preserves the marker
  - UPDATE on plain file does not inject a BOM

Addresses teknium1 review on PR #55661.
2026-08-04 14:34:24 +05:30
阿泥豆 cb3e8e9fb1 perf(file-ops): eliminate redundant subprocess calls in write_file and V4A patch path
write_file currently spawns up to 6 subprocesses per call:
  1. mkdir -p (separate call before atomic write)
  2. cat (to read pre-content for lint/BOM/line-ending detection)
  3. _atomic_write (mktemp + write + mv — the essential one)
  4. wc -c (to measure bytes written)
  5. _check_lint_delta (post-write lint — also essential)
  6. LSP snapshot (also essential)

This PR removes three of them without changing any observable behavior:

1. Fold mkdir -p into _atomic_write shell script (−1 subprocess/write)
   The atomic write script already runs a single shell; adding mkdir -p
   to it costs zero extra processes.

2. Add optional pre_content parameter to write_file (−1 subprocess/patch)
   patch_replace and V4A _apply_update already read the file for fuzzy
   matching. Passing that content as pre_content skips the redundant cat
   inside write_file. Fully backward-compatible: callers that don't pass
   pre_content still read from disk as before.

3. Replace wc -c with len(content.encode('utf-8')) (−1 subprocess/write)
   We already have the content in memory; encoding it to get the byte count
   is equivalent to wc -c for UTF-8 text.

4. Remove redundant _check_lint loop in apply_v4a_operations (−N subprocesses/V4A)
   write_file already runs _check_lint_delta internally. The old code ran a
   bare _check_lint(f) loop over all modified files — a re-read + re-lint
   without post_content context. Now lint results propagate from write_file
   via a four-tuple return, zeroing out the extra subprocesses.

Net effect:
  - write_file: 6 → 3 subprocesses per call (new files)
  - patch_replace: 6 → 5 subprocesses per call (pre_content skips cat)
  - V4A multi-file patches: saves 1 subprocess per modified file
  - A typical 4-file V4A patch drops from ~28 to ~16 subprocess calls
2026-08-04 14:34:24 +05:30
EndeavorYen 952d86b797 fix(file-sync): serialize concurrent sync cycles 2026-08-03 22:53:32 +05:30
kshitij 13172fca02 refactor(tool-search): drop dead fallback ladder in _available_source_summary
Simplify-pass finding: _listing_group_label already falls back to 'other' for empty source names, and _classify_source guarantees source_name=='' only when source=='other' — both  legs were dead by construction. Aligns the summary path's grouping with the listing path.
2026-08-03 19:11:30 +05:30
Carbon 48e12a06fa perf(tools): shrink lazy tool catalog overhead 2026-08-03 19:11:30 +05:30
Hafiz Ahmad Ashfaq 0de8c32f52 fix(dashboard): cache plugins hub payload and avoid auth probes 2026-08-03 18:44:38 +05:30
Jakub Wolniewicz ffb54305c4 perf(session-search): project fields before enrichment 2026-08-03 17:50:58 +05:30
PRATHAMESH75 fe6330de03 fix(stt): thread confidence thresholds into faster-whisper's own gate (#74178)
build_local_transcribe_kwargs read stt.local.no_speech_prob_threshold /
stt.local.logprob_threshold only for Hermes' post-filter
(_is_hallucinated_segment). faster-whisper's model.transcribe() never
received them, so its internal defaults (no_speech_threshold=0.6,
log_prob_threshold=-1.0) always applied and silently dropped
low-confidence segments before they reached the post-filter — making
those config knobs dead for the first gate.

Non-English speech decodes at a lower avg_logprob, so the English-tuned
defaults discard whole utterances (empty transcript despite correct
capture and language detection). Map the same config values through to
model.transcribe() so both gates stay in sync and the knobs work.
Defaults are unchanged, so behavior is identical unless a user tunes them.

Fixes #74178
2026-08-03 14:30:05 +05:30
kshitij ebf967ff2c polish(mcp): simplify-pass folds on the lazy-startup salvage
Five review findings folded:
- schema cache writes via utils.atomic_json_write (fsync; was bare
  tmp+replace), file moved to cache/mcp_schema_cache.json with 0o600
  (sibling precedent: registry discovery cache)
- phantom-tool reconciliation: after a lazy server's first-use connect,
  cached tools the live server no longer offers are deregistered (were
  permanent registry ghosts burning circuit-breaker strikes on every
  'Unknown tool' round-trip); stale fingerprint logged
- cache-load path now runs _scan_mcp_description like the eager path
  (cache file is user-writable JSON; defense-in-depth)
- write-through skips the disk rewrite when the entry is unchanged
  (a flapping stdio server was rewriting byte-identical JSON per
  revival)
- _lazy_server_fingerprints no longer write-only dead state (consumed
  by the reconciliation logging)

444 mcp tests green (440 pre-fold + 4 new guards); phantom-dereg and
write-skip mutation-checked.
2026-08-03 14:24:37 +05:30
kshitij 1d5ecad568 feat(mcp): lazy server startup from schema cache (design from #56832)
Wires the fingerprint-keyed schema cache (previous commit, @Vansh5632's
design from #56832) into the startup path, re-derived onto main's
current connect machinery:

- register_mcp_servers: servers with mcp_servers.<name>.lazy=true whose
  config fingerprint matches a valid cache entry register tools from
  cache WITHOUT spawning; miss/stale falls back to eager connect.
- First tool use routes through _ensure_lazy_server_connected, which
  composes with the connect cooldown (#50394) and _server_connecting
  dedup rather than duplicating the connect path.
- resource/prompt utility handlers (list_resources/get_prompt) also
  connect-on-first-use — closes the gap flagged in the original
  sweeper review.
- Write-through: a live connect refreshes the cache entry.

Config gate is per-server, default OFF, matching the
idle_timeout_seconds key pattern. 24 lazy/cache tests + 440 mcp-wide
green; mutation-checked (cache-read disabled -> registration test
fails; connect bypassed -> 3 first-use tests fail).
2026-08-03 14:24:37 +05:30
Vansh Gilhotra 135a29452a feat(mcp): add fingerprint-keyed on-disk MCP tool-schema cache
Stores per-server tool manifests in ~/.hermes/mcp_schema_cache.json so
tools can be registered into the agent snapshot without spawning the
stdio child at startup. Entries are keyed by server name plus a
fingerprint of the connection-defining config (command/args/url/
transport/tool filters), so any config change invalidates the entry.

Extracted from #56832.
2026-08-03 14:24:37 +05:30
liuhao1024 f07f47fe7d fix(lazy-deps): skip the install ladder on package-manager installs
Salvage of #48637 (Fixes #48628). On a NixOS-style install the venv's
site-packages lives in the read-only store, so ensure()'s
uv -> pip -> ensurepip ladder spends ~15s bootstrapping ensurepip only
to fail against a target it can never write. Fail fast with an
actionable message pointing at the system package manager.

Retargeted onto current main (the PR's base predates the durable-target
subsystem by ~8.1K commits) with two corrections to the original:

- Gate on _lazy_install_target() is None. The container deployment sets
  HERMES_MANAGED=true AND HERMES_LAZY_INSTALL_TARGET (a writable
  volume); the original guard would have blocked installs that path
  legitimately satisfies, breaking the NixOS-container mode.
- Reason string starts with 'unsupported ' because
  refresh_active_features classifies FeatureUnavailable by that prefix;
  the original wording made 'hermes update' report a hard failure
  instead of a skip.

Placed after _unsupported_feature_reason so a platform-specific reason
(more actionable) wins, and so ensure() agrees with
refresh_active_features, which pre-checks that same function.
2026-08-03 14:09:46 +05:30
f1aggo_macair 911d380296 fix(tools): allocate snapshot temp paths with mktemp instead of $BASHPID
Extracted from #54314 (@flag0x369), re-derived onto current main: macOS
ships bash 3.2 as /bin/bash, which lacks $BASHPID entirely — the
variable expands to empty string, collapsing every concurrent writer's
'unique' temp path onto the same file (torn snapshot writes under
concurrency). mktemp allocates per-writer unique paths portably.
Live-verified: /bin/bash -c 'echo $BASHPID' prints empty on this box.
2026-08-03 13:47:29 +05:30
f1aggo_macair 0125281609 fix(tools): allow Unicode letters in workdir validation
The workdir allowlist regex was ASCII-only, so perfectly normal
non-ASCII workdirs (Chinese Obsidian vault paths, accented dirnames)
were rejected with 'disallowed character'. Replace the regex with a
per-character check that accepts Unicode letters/digits (str.isalnum)
plus the same safe ASCII punctuation set, while still rejecting shell
metacharacters, control characters (newlines/tabs), and NUL.

Salvaged from PR #54314.

Co-authored-by: kshitij <82637225+kshitijk4poor@users.noreply.github.com>
2026-08-03 13:47:29 +05:30
Xue-1997 72e8e2983a fix(session-search): strip ANSI from recalled messages
Recalled session messages can carry raw ANSI escape sequences (e.g.
archived terminal output), which then re-enter the model's context.
Strip them in _shape_message before content is truncated/returned,
reusing tools.ansi_strip.strip_ansi.

Re-applied onto current main (the original hunk predates the
max_content_len truncation in _shape_message; stripping happens on the
raw content before truncation so escape bytes never count against the
budget). Extracted from #40276.
2026-08-03 13:47:03 +05:30
Teknium d127fb2197 fix(approval): stop treating newlines inside quoted arguments as command starts
A raw newline in the _CMDPOS start-position class made ANY multi-line
quoted argument look like a command boundary, so hermes send message
bodies, multi-line git commit -m messages, and heredoc text that merely
mentioned dangerous command names tripped the unconditional hardline
blocklist and could not run at all.

Mask newlines inside single/double quotes (detection-only, mirroring the
quote tracking in _iter_shell_command_starts) before building detection
variants. Real threats keep blocking: unquoted newlines stay command
separators, command substitutions inside quotes still anchor, and
_mark_command_starts still re-inserts newlines at genuine quote-aware
command starts. Masking runs on the RAW command before normalization,
which strips escapes and would otherwise corrupt quote state.

Regression tests cover both directions: multi-line quoted data passes
(hermes send, git commit -m, heredocs); bare/chained/substituted
shutdown-class and rm-floor commands still block.
2026-08-02 23:11:56 -07:00
Teknium a01f979b6e perf(tools): restore benchmark-sensitive phrasing in compacted description
Round-2 A/B (gpt-4o-mini, 6 reps) showed two passages could not survive
paraphrase: the DO-NOT-USE list needs the arrow-list shape with the
'no reasoning needed' qualifier (prose form regressed mechanical-work
routing 6/6->1/6), and the self-report rule needs the concrete
'claiming uploaded successfully may be wrong' framing (without it,
side-effect verification regressed 6/6->2/6). With both restored:
30/42 vs 30/42 on gpt-4o-mini and intent-parity on claude-haiku-4.5.
Final size: 1,900 chars (from 3,963).
2026-08-02 22:44:58 -07:00
Teknium 4be0d56023 perf(tools): compact delegate_task description by deduping against param schema
The top-level delegate_task description repeated content the model already
receives through parameter descriptions: the concurrency limit (tasks param),
the full nesting clause (role param), context-passing guidance (goal/context
params), and background semantics (background param). Every API call paid for
the duplication (~4,000 chars).

The description now carries only what exists nowhere else in the schema:
use/don't-use routing (execute_code, cronjob), the no-poll rule, the
non-durability warning, the self-report verification contract with concrete
verbs, the language-passing example, the leaf blocked-tool list, and model
inheritance. 3,963 -> 1,704 chars (~570 tokens saved per API call), and the
top-level text is now static (dynamic limits flow only through the two param
descriptions, which are already rebuilt per get_definitions() call).

A/B benchmark across 4 models (gpt-4o, gpt-4o-mini, claude-haiku-4.5,
llama-3.3-70b) showed the naive compaction in PR #72813 regressed weaker
models on exactly the passages it cut (side-effect verification 8/8->0/8 on
gpt-4o-mini; language passing 3/3->0/3 on haiku-4.5). This version keeps
those benchmark-sensitive hooks verbatim.

Tests pin the contracts at keyword level (not prose-literal) plus a size
ceiling, and verify dynamic limits still reach the model via the tasks/role
param descriptions.

Refs #72737, supersedes the delegate_task half of PR #72813.
2026-08-02 22:44:58 -07:00
Mahdi Hedhli 75901a295d perf(tts): pipeline sync per-sentence synthesis with playback
The universal sync fallback in stream_tts_to_speaker ran strictly serially
per sentence — synthesize, play, and only then start synthesizing the next
sentence — so every sentence boundary added a full synthesis-time of dead
air. Chunked streamers (elevenlabs/openai/gemini/xai) already avoid this;
every other provider (edge, piper, plugin providers) paid it on each reply
in voice mode and the wake-word loop.

_SyncSentencePipeline overlaps the two: one single-threaded synthesis
worker (sentences stay FIFO; providers never see concurrent calls from
this loop — same effective concurrency as before) feeds one playback
worker through a small bounded queue, so sentence n+1 synthesizes while
sentence n plays. Lookahead is bounded (backpressure + at most a couple of
temp files), stop_event short-circuits both stages, synthesis failures are
isolated per sentence, temp files are always unlinked, and the finally
block flushes the pipeline BEFORE tts_done_event fires so continuous voice
mode never reopens the mic over its own voice. synthesize/play are
resolved late so existing monkeypatch-based tests work unchanged.

Measured with a real local model provider (OmniVoice plugin, Apple
Silicon), same 3-sentence reply, playback simulated at the produced clips'
true durations, best-of-2 interleaved runs under identical load:

                     serial   pipelined
  time to first word  10.8s        4.4s
  mid-reply dead air  11.2s        1.8s   (second gap: 0.03s)
  full reply wall     33.2s       17.0s

Tests: 4 new (timestamp-proven overlap, order + per-sentence failure
isolation, stop skips queued playback, temp-file hygiene); the existing
sync-fallback and display-callback tests pass unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 10:25:15 +05:30
Eugeniusz Gilewski 9421c5afdf perf(agent): precompile response and skill-scan regexes (#33208)
strip_think_blocks passed the same response-scrubbing strings through
re's pattern dispatcher on every response. Skills Guard repeated the
same work for 121 patterns against every scanned line.

Compile the existing expressions once and reuse Pattern.sub/search. Keep
each generic tool-call tag in its own paired expression so mismatched
openers retain their payload while existing stray-closer cleanup remains
unchanged.

Part of #33208
Salvaged from #32713 by @ErnestHysa.

Co-authored-by: ErnestHysa <takis312@hotmail.com>
2026-08-03 09:56:36 +05:30
Teknium aac74be2f1 fix(approval): classify CLI/TUI approval timeouts separately from explicit denials
When an approval prompt expired without a response, every CLI-side path
collapsed the timeout into the same 'deny' choice as an explicit user
refusal, so the agent was told the user denied the action when the user
simply never answered. The gateway wait already distinguished the two
('timed out without user response... Silence is not consent.'); this
brings the CLI/TUI/ACP surfaces to parity.

- prompt_dangerous_approval(): input()-path expiry now returns a distinct
  'timeout' choice (still fail-closed).
- cli.py _approval_callback + hermes_cli/callbacks.py approval_callback:
  deadline expiry returns 'timeout' instead of 'deny'.
- check_all_command_guards / _run_approval_gate CLI tails: 'timeout' maps
  to outcome='timeout' with a 'timed out without user response... Silence
  is not consent.' BLOCKED message (matching the gateway wording);
  explicit deny keeps outcome='denied' and gains user_consent=False for
  shape parity.
- computer_use: 'timeout' verdict threads through the CLI adapter and
  yields a 'prompt timed out — the user did not respond' error instead of
  'denied by user'.
- ACP permissions bridge: FutureTimeout returns 'timeout' (other failures
  still 'deny'); elicitation maps 'timeout' to 'cancel' like the gateway's
  unresolved outcome; codex wire mapping documents deny/timeout→decline.
- write_approval already treats unknown choices as 'stage, not drop', so
  a timeout now stages the memory write instead of silently refusing it.

Every timeout path remains fail-closed — the action never runs; only the
classification reported to the agent changes.
2026-08-02 20:21:59 -07:00
Alex Fournier 14c8bd646c Merge updated model metrics into tool metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-02 20:11:20 -07:00
Teknium f01c193be4 refactor(schema): trim terminal and execute_code schema prose ~40%
Every tool schema ships on every API call. The terminal schema was
5,641 chars (~1,410 tokens) and execute_code 2,842 (~710) — the two
largest core tools, padded with repeated war stories and triple-stated
rules. Schema token audit across 88 tools: ~33k tokens total.

This trims prose while preserving every hard rule (each still stated
exactly once):
- terminal description 2,324 -> 1,233 chars: tool-redirect lines
  collapsed to one sentence; background/notify guidance deduplicated
  (was stated in desc + 2 params); PTY/pager rules merged.
- background/notify_on_complete/watch_patterns params 692/508/1,114 ->
  ~330/250/490 chars: kept the mutual-exclusion contracts, the
  rate-limit consequence, and the bounded-vs-long-lived distinction;
  dropped narrative repetition.
- execute_code description tightened (helper docs inlined to one line
  each; when-to-use kept).

Net: terminal schema 5,641 -> 3,386 chars, execute_code 2,842 -> 2,522
— ~700 tokens saved on EVERY request with the terminal+code toolsets.
One test updated (pinned a removed phrase; now pins the rule's new
phrasing).
2026-08-02 16:02:26 -07:00
Teknium 80631c4aea feat(terminal): recoverable truncation — spill full output + report pre-truncation size
Truncated terminal output was information LOSS: the middle was gone
and the only recovery was re-running the command (data: 1,394
truncation markers in a 250k-call window, with re-runs and grep
retries chained behind the big ones).

Truncation is now deferred retrieval (opencode/goose/qwen-code
pattern, codex's original_token_count idea):

- tools/environments/base.py: _BoundedOutputCollector gains an
  optional spill tee — when foreground output overflows the capture
  window, the FULL stream is teed to
  ~/.hermes/cache/terminal-output/out-*.log (lazy file creation with
  backlog backfill, 5MB hard cap, 7-day opportunistic cleanup,
  disk errors never break execution). All three _wait_for_process
  returns attach {output_total_chars, full_output_path} via a shared
  finalizer.
- tools/terminal_tool.py: redacts the spill with the same
  redact_terminal_output pass as the visible output (no secret
  persists unmasked), then surfaces output_total_chars,
  full_output_path, and a truncation_note pointing at
  search_files/read_file instead of a re-run.

Non-truncated results are byte-identical; internal unbounded
consumers (file-ops cat reads, RPC reads) are untouched (spill only
arms with bounded_capture=true).
2026-08-02 15:52:14 -07:00
Teknium 1c6d1a23c0 feat(patch): list match locations in ambiguous old_string errors
'Found N matches for old_string' (190+ occurrences in a 250k-window)
previously reported only the count, forcing a re-read of the file to
find the occurrences before retrying. The error now appends up to 5
'L<line>: <snippet>' rows (80-char cap per snippet, overflow noted),
so the model can disambiguate in ONE follow-up — add neighboring
context or choose replace_all — without the intermediate read.

Applies to both patch modes (replace + V4A) since they share
fuzzy_find_and_replace.
2026-08-02 15:51:43 -07:00
Teknium 794d6c434e fix(search): restore zero-match probes on the rg engine after auto-multiline early-return
Integration regression between two merged PRs: #77102's rg-path early
return (added to skip the grep-era line-oriented warning) also skipped
the #77011 zero-match steering probes (case-insensitive, hidden-file,
literal-vs-regex), silencing them on the primary engine. Caught by
#77001's rebased CI run (its branch carried both features together for
the first time).

_search_content now runs the zero-match probe block for BOTH engines
and only exempts the rg path from the legacy line-oriented \n warning
(rg auto-enables --multiline; the grep fallback keeps the explanation).
2026-08-02 15:34:22 -07:00
Teknium eb62143006 feat(execute_code): recovery hints for known sandbox failure classes
The top execute_code failure shapes in production (state.db mining)
are sandbox-contract confusions, not logic bugs: importing tools that
aren't in the sandbox from hermes_tools (23x in one window, incl.
importing the built-in helpers json_parse/shell_quote/retry), importing
third-party packages absent from the sandbox interpreter (matplotlib
6x), and indexing tool-result dicts as strings. The stderr traceback
alone sends models into re-diagnosis loops.

Failed scripts (exit != 0) now carry one actionable 'hint' field:
- unavailable hermes_tools import -> lists the tools that ARE
  importable in this session + points to normal tool calls otherwise;
- built-in helper import -> 'no import needed, call it directly';
- ModuleNotFoundError -> 'sandbox has stdlib only; use terminal() with
  the project venv for third-party packages';
- string-indexing errors -> 'tool functions return dicts, do not
  json.loads them'.

Bounded 4KB stderr scan, first match wins, never raises; successful
scripts and unknown failures are untouched.
2026-08-02 15:13:24 -07:00
Teknium 1cefabc8af feat(search): auto-enable multiline mode for newline patterns
A regex \n (or a raw newline) in a search_files content pattern cannot
match in rg's default line-oriented mode. It previously either
hard-errored ('the literal "\n" is not allowed in a regex' — 17
occurrences in the production window) or, after the newline-warning
patch, returned 0 matches with an explanation — either way the model's
cross-line search intent required a manual workaround.

The rg engine now detects the pattern shape (_pattern_has_regex_newline,
already used by the warning path: odd-backslash \n escape or raw
newline; escaped \\n literals excluded) and enables -U/--multiline up
front, noting the mode switch in the result warning. Plain patterns are
untouched; the old line-oriented explanation is retained for the grep
fallback engine, which has no multiline mode.
2026-08-02 15:13:04 -07:00
Teknium 2a3a7e6f53 feat(skills): dedup repeat skill_view calls with an unchanged-content stub
skill_view re-sent full skill content on every call: ~286k tokens of
verbatim repeat views in a 400k-msg production window (one session
loaded the same skill 9 times), and a single repeat view of a large
skill costs ~25k tokens.

Mirrors read_file's proven unchanged-stub pattern: a per-task cache
keyed on (resolved name, file_path) with an mtime+size fingerprint of
the served file. On a repeat view of an UNCHANGED file, return a short
stub pointing at the earlier result. This does NOT violate the
skills-are-loaded-fully rule — the stub only ever replaces content
that is already fully present earlier in the same conversation, and:

- any on-disk change (patch, external edit) invalidates the entry;
- context compression clears the cache (wired next to
  reset_file_dedup in conversation_compression.py) so post-compression
  re-views return full content;
- setup-needed views are never deduped (readiness can change without
  the file changing);
- no task_id -> no dedup; caches are task-isolated; 200-entry cap.

Live E2E: repeat view of hermes-agent-dev 99,739 chars -> 374-char
stub.
2026-08-02 15:12:43 -07:00
Teknium 2c8a932f80 feat(file): verify write_file content on disk and say so (verified: true)
write_file confirmed only SIZE (wc -c) after writing — never content.
Models compensated by re-reading files immediately after writing them
(154 verify-reads in a 400k-msg production window), and a corrupted
write (truncated pipe, backend FS oddity) could silently pass.

The write path now compares the on-disk sha256 against the intended
content (one shell call). Three outcomes:
- match -> result carries verified: true; the schema tells the model
  an explicit contract: do NOT re-read to check the write landed.
- mismatch -> hard error ('The write did not persist correctly'),
  mirroring patch_replace's existing post-write verification.
- backend can't hash (no sha256sum) -> flag omitted, write unaffected.

Hashes the shim-adjusted content (after CRLF/BOM preservation) so
Windows-line-ending and BOM round-trips verify correctly; surrogatepass
encoding matches the rest of the codebase's hashing of model text.
2026-08-02 15:12:23 -07:00
Teknium 5d675a2ca7 feat(patch): whitespace-visualized diagnosis on residual no-match errors
When old_string survives all 9 fuzzy strategies without a match but
the closest candidate line matches after stripping whitespace, the
failure is whitespace-shaped (tabs vs spaces, indent depth). The
did-you-mean hint now appends a two-line diagnosis with leading
whitespace made visible:

  Whitespace difference detected (→ = tab, · = space):
    file has: →def start(self):
    you sent: ····def start(self):
  Use the exact whitespace shown in 'file has'.

Pattern ported from crush's diagnoseMismatch (agent-codebase survey) —
it converts the residual dead-end error into a one-turn fix. Only the
leading run is visualized (interior spacing stays readable); content-
shaped misses and raw-exact candidates are unchanged.
2026-08-02 15:12:02 -07:00
null-runner 820f63e842 [verified] fix(desktop): preserve background delegates across session switches 2026-08-02 15:11:37 -07:00
Teknium 6f5d6b1f5b feat(terminal): auto-save parser-limit-blocked payloads as runnable scripts
Follow-up to the recovery-recipe commit on this branch, per review:
instead of only TELLING the model to re-author the payload via
write_file (2 turns), materialize the blocked command to
~/.hermes/cache/blocked-scripts/blocked-*.sh and point the recovery
at it directly: 'saved to <path> - review it, then run
terminal(command="bash <path>")' (1 turn).

Safety posture is unchanged or better:
- Nothing is executed here; the file is only written.
- The bash <path> follow-up goes through the normal execution
  pipeline, including the referenced-script content guard, which
  inspects script files named in commands - the payload is MORE
  visible to policy than it was inline.
- Genuine hardline blocks (destructive ops) never save anything
  (test-asserted).
- Save failures fall back to the previous manual write_file recipe.
- 7-day opportunistic cleanup of saved payloads.
2026-08-02 15:11:21 -07:00
Teknium b1711c6f2e fix(terminal): blocked-command errors carry a concrete recovery recipe
Two block classes from production mining (250k-call window) that
models answered with blind rephrase-retries:

1. Parser-limit / malformed-payload hardline blocks (198x): these fire
   on oversized inline payloads (heredocs, giant one-liners), not on a
   forbidden operation - but the message read like a permanent ban.
   The block now appends: 'RECOVERY: ... write the script to a file
   with write_file, then run bash /path/script.sh - do not retry
   inline.' Genuine hardline blocks (destructive filesystem
   operations) are unchanged.

2. Backgrounding-wrapper blocks (200x): the guidance now spells out
   the exact corrected call shape - 're-send WITHOUT the wrapper as
   terminal(command="<cmd>", background=true,
   notify_on_complete=true)' - instead of describing the feature
   abstractly.
2026-08-02 15:11:21 -07:00
Teknium 0b149ca030 fix(process): wait timeout result reads as status, not failure
process(action='wait') hitting its window returned status='timeout'
with a terse note — models read it as an error and re-issued identical
waits (process is the #1 exact-duplicate tool call in production: 511
dupes in a 400k-msg window; wait is 57% of all process actions).

The timeout result now carries:
- process_running: true — machine-readable 'this is a status, not a
  failure'
- an explicit note: 'Wait window of Ns elapsed — the process is still
  running. This is not an error. Uptime: Ms.' plus the right next step:
  when notify_on_complete is set, 'you will be notified on exit — do
  more work instead of waiting again'; otherwise a pointer to
  notify_on_complete for next time.
- the clamp note (requested > max) now composes with the status note
  instead of replacing it.

Exited/interrupted results are unchanged.
2026-08-02 15:11:02 -07:00
Teknium e7aa06c3a6 feat(search): hidden/gitignored probe on zero-match results
Third zero-match steering tier: when a content search finds nothing in
visible files, probe once with rg --hidden --no-ignore --count-matches.
If the pattern exists only in dotdirs or gitignored files, the result
says so and tells the model to search the hidden path explicitly.

Found live by the benchmark battery: a task with a match inside
.hidden/ returned a bare 0 and the model missed the file entirely in
2/3 baseline runs.
2026-08-02 15:10:42 -07:00
Teknium 5797b50288 feat(search): zero-match probes and multi-path recovery
Two dead-turn classes from production mining (state.db, recent window):

1. 13.9% of 19.6k content searches return 0 matches with no steering.
   Now a 0-match content search runs one cheap rg -i --count-matches
   probe (plus an rg -F probe when the pattern has regex metachars) and
   attaches what it found: '0 exact matches, but N case-insensitive
   matches — casing may be wrong' / 'N literal matches — metacharacters
   need escaping'. True zero-match results stay clean (no noise).

2. 122 'Path not found' failures came from models passing several paths
   in ONE path string ('dir1 dir2 dir3', comma lists). Instead of
   failing wholesale, split the string, search every path that exists,
   merge results, and report skipped parts in a warning. Single-path
   misses keep the existing Similar-paths hint; all-missing multi-path
   strings still error.

Both probes are bounded (count-only rg, 30s timeout, max 2 invocations)
and wrapped so a probe failure can never break the search result.
2026-08-02 15:10:42 -07:00
Teknium a18a2f170c feat(terminal): echo cwd in result when a command changes the working directory
Production mining (state.db, 400k-msg window): 60.2% of 104k terminal
calls carry a defensive 'cd X && ' prefix (~925k tokens of pure prefix)
and 2,462 failed calls led with cd — the model cannot see cwd state, so
it re-asserts it on every call and runs pwd/ls diagnostics after
directory changes.

The result dict now includes a 'cwd' field whenever the session cwd
after the command differs from the cwd it started in (cd, pushd,
chained cd). Stable-cwd commands are unchanged (no field, no noise).
Per-command workdir overrides stay transient by contract and never
echo. Schema note added so models learn to trust session cwd instead
of prefixing. Pattern borrowed from crush's <cwd> injection.

realpath comparison avoids false echoes through symlinks; the echo is
wrapped defensively so a backend without .cwd can never break the
result path.
2026-08-02 15:10:13 -07:00
Teknium 99d6f55e38 feat(patch): detect already-applied edits and return success no-op
The #1 patch failure class in production (state.db mining, 250k-window)
is a re-send of an edit that already landed: 'old_string and new_string
are identical' (299 occurrences) plus a share of hunk-not-found errors
where the new text is already in the file. These errored, sending
models into re-read/re-patch loops.

New tools/fuzzy_match.is_already_applied(content, old, new) — a
conservative check requiring (1) non-trivial new_string (>=8 chars),
(2) EXACT presence of new_string, (3) old_string gone (unless
identical). Wired into three sites:

- patch_replace (replace mode): returns success + no_change: true +
  an explicit note instead of the identical-strings / no-match error.
- V4A validation phase: an already-applied hunk validates as a no-op
  so multi-hunk patches no longer fail wholesale when one hunk landed
  in a prior call.
- V4A apply phase: mirrors the same skip so the two phases agree.

Genuine no-matches (new text absent) and half-applied renames (old
text still present) keep their error behavior — covered by tests.
2026-08-02 15:09:53 -07:00
Teknium af27e60603 feat(file): raise read_file default limit from 500 to 2000 lines
Production mining (state.db, 28.5k read_file calls in the recent
window) shows 74.3% of reads truncated — nearly all by the 500-line
default, not the char budget (22,443 line-limit vs 13 char-budget
truncations). That churn produced 12,229 redundant re-reads, and after
a truncated result the most common next move was fleeing to terminal
cat/sed (4,608 times) — the pagination contract was not trusted.

Median truncated file is 2,422 total lines, so a 2000-line default
makes 44% of today's truncated first-reads complete in one call while
the unchanged ~100K-char budget still caps worst-case result size
(same ceiling as before: 500 lines x 2000-char line cap = the same
100K). Schema max was already 2000.

Touchpoints: DEFAULT_READ_LIMIT + both read_file signatures
(file_operations.py), read_file_tool + schema text (file_tools.py),
execute_code sandbox stub docs (code_execution_tool.py), 3 tests
pinning the old default.
2026-08-02 15:09:33 -07:00
Teknium 677473273e feat(terminal): output-pattern failure hints for common error classes
When a command exits non-zero, scan the first 4KB of output for
well-known failure shapes and attach one short, actionable recovery
hint to the tool result ('hint' field):

- gh 'Unknown JSON field' (9.2k occurrences in a 250k-call window)
- git merge conflicts (1.2k) — stop verbatim retries
- command not found (1.0k), incl. python->python3 and pip->pip3
- ModuleNotFoundError (739) — venv activation guidance
- 'already exists' (633), gh rate limits (133), permission denied
- exit-code-only tier: 124 timeout, 126 not-executable, 137 SIGKILL

Hints are suppressed when the existing exit_code_meaning tier already
explains the code (grep=1 etc). Pattern order = production frequency
from state.db mining; first match wins; pure function, no I/O.
2026-08-02 15:08:35 -07:00
xrazai e57a8f5cb9 fix(gateway): preserve delegates during session reaping 2026-08-02 14:02:08 -07:00
kshitij fb6446fc9e fix(cron): scope cron approval context per session
Replace the process-global HERMES_CRON_SESSION env var with a per-session
ContextVar so a cron tick in the gateway process cannot leak into unrelated
live gateway/API/TUI turns. The cron scheduler now sets the ContextVar
inside the job's try/finally scope and resets it on cleanup. Gateway, API
server, ACP adapter, and TUI gateway all pass cron_session='' to explicitly
mark their sessions as non-cron, masking any stale process env.

Co-authored-by: hinablue <hinablue@gmail.com>
Closes #37968
2026-08-03 00:25:20 +05:30
AlexFucuson9 9eb8e20c68 fix: use lazy logging with %s formatting in logger calls
Replace f-string interpolation in logger calls with lazy %-style
formatting across 10 files (38 instances). This follows Python logging
best practices — the message is only formatted if the log level is
enabled, avoiding unnecessary string concatenation overhead.

Files changed:
- trajectory_compressor.py (6)
- mini_swe_runner.py (2)
- agent/tool_executor.py (1)
- agent/model_metadata.py (1)
- agent/agent_runtime_helpers.py (3)
- agent/chat_completion_helpers.py (3)
- agent/conversation_loop.py (8)
- tools/skills_hub.py (2)
- tools/environments/docker.py (10)
- gateway/kanban_watchers.py (2)
2026-08-02 23:17:40 +05:30
AMATH 9d7c2d450c fix: _check_disk_usage_warning runs full rglob on every terminal call 2026-08-02 22:47:50 +05:30
kshitij 9b50a99b39 refactor(tools): existence stat outside the global tracker lock
Simplify-pass advisory on the #25387 salvage: _read_tracker_lock is one
global lock guarding every task's read/search bookkeeping (15 sites). A
hung stat on a dead network mount inside it would stall all of them.
Check-then-recheck matches the dedup mtime pattern 30 lines below: read
the entry under the lock, stat outside, reacquire only to evict.
2026-08-02 22:45:28 +05:30