Commit Graph

22221 Commits

Author SHA1 Message Date
kshitij 1d3242ac72 chore: map universeszym@mail.ustc.edu.cn to Johnny-xuan 2026-08-03 11:02:47 +05:30
kshitijk4poor 5633764733 test: wait for ledger-wrapped sends in drain assertions
Two consumer tests assumed a send completes synchronously within the
handler turn: the continuation-drain test polled on handler-call count
then immediately asserted on adapter.sent, and the split-brain heal
test drained with bare zero-delay yields. With ledger calls hopping to
worker threads around each send, the reply can land microseconds after
those checks. Poll for the actual sends with a bounded 2s window —
same invariants, scheduling-robust.
2026-08-03 11:00:49 +05:30
kshitijk4poor 5b36d64583 test: raise blocking-probe timeouts for loaded CI runners
CI slices failed the offload tests with 0.5s witness timeouts: on a
loaded shared runner the event loop thread can take >0.5s to get
scheduled even when NOT blocked, making the probe report a false
positive. A genuinely blocked loop can never set the progress event at
any timeout (the witness coroutine can't run at all), so 5s only
absorbs scheduler flake without weakening the invariant. Mutation
re-verified: reverting the offload still fails all 4 tests.
2026-08-03 11:00:49 +05:30
kshitijk4poor b7e3cc37be test: move the redelivery event-loop test to the class that has its helpers
The sweep-path test parametrizes over _runner/_adapter, which live on
TestGatewayRedeliverySweep; main later added
TestUnconnectedPlatformKeepsItsBudget at the cherry-pick anchor point and
the test landed in that class, where the helpers don't exist
(AttributeError x2). Placement-only move.
2026-08-03 11:00:49 +05:30
ibaldr89 498800a22e fix(gateway): offload delivery ledger I/O 2026-08-03 11:00:49 +05:30
kshitij 07dd2f8fc9 chore: AUTHOR_MAP kshitij@k4poor.dev -> kshitijk4poor 2026-08-03 10:48:49 +05:30
kshitij cc04825c5d fix: content-based diff for model-switch marker merge (#76870)
The original PR #77274 used positional slicing (current_history[len(history):])
to detect the model-switch-only mutation.  But _append_model_switch_marker
strips prior markers in-place before appending the new one, so when a prior
marker existed (every switch after the first in a session), the net length
delta is zero and the slice produces an empty list — the merge path is dead
code for the common case.

Replace with a content-based diff: strip markers from both the turn-start
snapshot and the current history, then check that the non-marker content is
identical.  This correctly handles the strip-and-replace behavior.

Also guard against auto-compression making result["messages"] shorter than
the turn-start history — use the full result as the base when that happens.

Added test covering both no-prior-marker and prior-marker cases.
2026-08-03 10:48:49 +05:30
JonthanaHanh f9ed58e6ac fix(tui-gateway): merge agent output on model-switch history_version mismatch (#76870)
When a model switch occurs mid-turn, `_append_model_switch_marker()`
appends a marker to session history and increments `history_version`.
The turn completion guard then sees `current_version != history_version`
and discards all agent output — producing empty assistant messages in
the session DB.

Detect when the only history mutation during the turn was one or more
model-switch markers.  In that case, merge the agent's new messages
into the current history (which now contains the marker) instead of
discarding them.  Genuine desyncs (undo/compress/retry) still surface
the warning as before.

Fixes #76870
2026-08-03 10:48:49 +05:30
Mahdi Hedhli 75901a295d perf(tts): pipeline sync per-sentence synthesis with playback
The universal sync fallback in stream_tts_to_speaker ran strictly serially
per sentence — synthesize, play, and only then start synthesizing the next
sentence — so every sentence boundary added a full synthesis-time of dead
air. Chunked streamers (elevenlabs/openai/gemini/xai) already avoid this;
every other provider (edge, piper, plugin providers) paid it on each reply
in voice mode and the wake-word loop.

_SyncSentencePipeline overlaps the two: one single-threaded synthesis
worker (sentences stay FIFO; providers never see concurrent calls from
this loop — same effective concurrency as before) feeds one playback
worker through a small bounded queue, so sentence n+1 synthesizes while
sentence n plays. Lookahead is bounded (backpressure + at most a couple of
temp files), stop_event short-circuits both stages, synthesis failures are
isolated per sentence, temp files are always unlinked, and the finally
block flushes the pipeline BEFORE tts_done_event fires so continuous voice
mode never reopens the mic over its own voice. synthesize/play are
resolved late so existing monkeypatch-based tests work unchanged.

Measured with a real local model provider (OmniVoice plugin, Apple
Silicon), same 3-sentence reply, playback simulated at the produced clips'
true durations, best-of-2 interleaved runs under identical load:

                     serial   pipelined
  time to first word  10.8s        4.4s
  mid-reply dead air  11.2s        1.8s   (second gap: 0.03s)
  full reply wall     33.2s       17.0s

Tests: 4 new (timestamp-proven overlap, order + per-sentence failure
isolation, stop skips queued playback, temp-file hygiene); the existing
sync-fallback and display-callback tests pass unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 10:25:15 +05:30
Teknium dd600d1ace fix(discord): suppress link embeds in tool preview markdown links
Wrap the masked-link destination in angle brackets so Discord does not
unfurl an OG-preview embed under every tool progress bubble. quote()
percent-encodes any <> inside the URL itself, so the wrapper cannot be
broken out of.
2026-08-02 21:48:06 -07:00
Coffee☕️ e599f100ec fix(discord): avoid truncated URL link targets 2026-08-02 21:48:06 -07:00
Coffee☕️ 7af104b37b refactor(discord): keep link formatting adapter-local 2026-08-02 21:48:06 -07:00
Coffee☕️ 56941eb329 refactor(gateway): share markdown link formatting 2026-08-02 21:48:06 -07:00
Coffee☕️ 911d8dfbf4 refactor(discord): simplify tool preview links 2026-08-02 21:48:06 -07:00
Coffee☕️ df9e039d2d fix(discord): preserve links in truncated tool previews 2026-08-02 21:48:06 -07:00
kshitij 037825c1f2 fix: check NUL bytes before size limit in _read_referenced_script
On Linux, /usr/bin/python3 is >1MB, so the size check fired before
the NUL check could run — the binary was returned as unsafe=True
(blocked) instead of (None, False) (skip). Reorder: read the bounded
chunk first, check for NUL bytes (binary → skip), then check size
(oversized text → fail closed).
2026-08-03 10:11:39 +05:30
CriptoGus c98ed22e42 fix(cron): stop lifecycle guard false-positives and crashes on .py/binary scripts
The gateway lifecycle guard (cron/lifecycle_guard.py) applied shell-style
tokenization and script-reference resolution to non-shell content, with two
regressions:

#77131 - every .py cron script using pathlib division was hard-blocked:
  Path.home() / ".hermes" / ".env" tokenizes the bare "/" operator as an
  executable path, which resolves to the filesystem root; the regular-file
  check then fails closed as unsafe. Since Python runs under the
  interpreter, never through a POSIX shell, the shell-script reference walk
  is a false-positive generator on Python sources. check_gateway_lifecycle
  now skips the walk for *.py scripts (the direct command regex still scans
  the full text), and _iter_referenced_shell_scripts skips pure-separator
  tokens.

#76762 - terminal commands invoking a binary by absolute path (e.g.
  /usr/bin/python3) crashed the guard with ValueError: embedded null byte:
  the walk read the binary's bytes, decoded them as text, and re-tokenized
  machine code; the recursion then hit Path.resolve() on a NUL-bearing
  path while only OSError was caught. _read_referenced_script now skips
  NUL-containing files (binaries are not referenced shell scripts) and
  resolve() tolerates ValueError.

Shell scripts (.sh/.bash/.zsh) keep the full deep scan; literal lifecycle
commands in .py scripts are still blocked by the direct regex. New tests
cover all four behaviors.
2026-08-03 10:11:39 +05:30
kshitij 21040c4ab6
Merge pull request #77339 from kshitijk4poor/chore-dblank-email
chore: add contributor email mapping for danielblankhh
2026-08-03 10:04:45 +05:30
kshitijk4poor 45d3416d82 chore: add contributor email mapping for danielblankhh 2026-08-03 10:04:35 +05:30
Anthony Tran 947fdeab3b fix(whatsapp): guard bridge reconnect against hangs and unhandled rejections
startSocket() awaits useMultiFileAuthState() and fetchLatestBaileysVersion()
before it creates a socket or registers event handlers, and the close handler
re-entered it via a bare setTimeout(startSocket, ...). That leaves two
unrecoverable failure modes on a reconnect:

- a rejection is an unhandled promise rejection (fatal on modern Node)
- a hang leaves the bridge permanently disconnected with nothing left to
  retry, while its HTTP server keeps answering 503 to the gateway

The second mode was observed in the field: fetchLatestBaileysVersion() is a
plain fetch to raw.githubusercontent.com with no AbortSignal, and after a
stream:error 503 disconnect the bridge logged 'Reconnecting in 3s...' once
and then sat silent and disconnected for 27+ hours until manually restarted.

Fix, as two pure helpers in bridge_helpers.js (keeping bridge.js side-effect
free to test):

- createReconnectScheduler(): every (re)connect entry point now catches a
  failed startSocket() and reschedules it instead of dying or going silent
- createVersionResolver(): bounds the version fetch with a 15s timeout and
  falls back to the last known-good version (or the Baileys default before
  first success) instead of pending forever
2026-08-03 10:03:57 +05:30
kshitij 648c01c693 fix(matrix): pickle key from resolved device ID, skip migration after reset, cache enc info
Follow-up fixes for the combined Matrix crypto salvage (#71073,
#71543, #71547):

1. Construct _pickle_key from client.device_id (resolved from whoami)
   instead of self._device_id (the configured value). Without this, when
   #71543 makes the token's real device win over a stale
   MATRIX_DEVICE_ID, the pickle key is built from the stale value and
   the Olm account is stored under a key that can never be looked up
   again — perpetuating the same decryption failure the PRs aim to fix.

2. Skip _migrate_legacy_crypto_pickle when the store was just deleted
   by _reset_crypto_store_if_device_changed — there is no account to
   migrate. Also check the migration return value and log a warning on
   failure instead of proceeding to olm.load() which fails with a
   cryptic BAD_ACCOUNT_KEY.

3. Add a local dict cache (_enc_info_cache) to _CryptoStateStore so the
   homeserver fallback in get_encryption_info() does not make a network
   round-trip on every is_encrypted() call. MemoryStateStore does not
   implement set_encryption_info, so the existing cache-back is a no-op.

4. Log homeserver encryption-info query failures at DEBUG level instead
   of silently returning None (which would cause OlmMachine to treat an
   encrypted room as unencrypted).

5. Update _CryptoStateStore docstring to mention the homeserver fallback.
2026-08-03 10:03:35 +05:30
ckaznocha 80d5a57b94 fix(matrix): commit the migrated account only after the session sweep
The account was written under the new pickle key before sessions were
re-pickled. The account is effectively the migration's commit marker —
once it reads under the current key, the fast path short-circuits every
later startup — so a sweep that errored or was interrupted left the
remaining legacy-key sessions stranded permanently with no retry.

Sweep first, commit the account last, and return False on sweep failure
so the migration is retried on the next start.

Also corrects the unreadable-row log: it claimed rows were being dropped
while no DELETE was ever issued. Such rows are left in place (already
unusable; deleting crypto material on a guess is not worth it) and the
message now says so.

Adds session-sweep coverage, which was previously absent: rows rewritten
under the current key, rows already current left alone, unreadable rows
left in place, and a failed sweep that leaves the account uncommitted.
The existing migration test now fakes the olm C-extension so the suite
no longer requires libolm.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 10:03:35 +05:30
ckaznocha 3a43422065 fix(matrix): migrate crypto store when the Olm pickle key changes
The Olm account pickle key is derived from the account ID plus the
configured device ID (acct:device_id). If the crypto store's account was
created before MATRIX_DEVICE_ID was set — e.g. the very first password
login, where the device ID is only known after connecting — it gets
pickled under "<acct>:default". Setting MATRIX_DEVICE_ID afterwards (a
reasonable thing to do once you know the device ID you want to pin)
changes the derived pickle key, and every subsequent unpickle attempt
fails with BAD_ACCOUNT_KEY. In optional-E2EE mode that failure is
swallowed and encryption silently stays disabled instead of surfacing an
actionable error.

_migrate_legacy_crypto_pickle() detects the BAD_ACCOUNT_KEY failure,
tries the known legacy pickle keys, and re-pickles the account (plus
every stored olm/megolm session — sessions share the same pickle key, so
migrating only the account would leave them unreadable on the next
decrypt and silently break key sharing with peers) under the current
key. It only reports failure when no known key can unpickle the
account, with a log message pointing at what changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-03 10:03:35 +05:30
ckaznocha 907bdc7856 fix(matrix): let the token's own device win over a stale MATRIX_DEVICE_ID
connect() resolved client.device_id as `self._device_id or resolved_device_id`,
so a configured MATRIX_DEVICE_ID masked the live whoami() device. With
persisted A, configured A and a rotated token reporting B, the reset
compared A to A and never fired — exactly the token-rotation case this PR
claims to handle.

An access token is bound to one device and the homeserver only accepts key
uploads for that device, so a configured value naming a different one
cannot work. The live whoami() device now wins on conflict and logs an
error naming both. The configured value is still preferred when whoami()
reports no device.

Adds a connect()-level regression for persisted A + configured A +
whoami B, and corrects test_connect_uses_configured_device_id_over_whoami,
whose stated premise this inverts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 10:03:35 +05:30
ckaznocha a4b686c29e fix(matrix): reset Olm crypto store when the access token's device ID changes
The crypto store is keyed by Matrix user ID, not device ID, so swapping in
a new access token (which mints a new device_id) silently inherits the
previous device's Olm account. That account's identity keys can never be
published under the new device ID, and the pickle key embeds the old
device ID anyway — the result is stale-key mismatches and cross-signing
signatures the homeserver refuses to replace, degrading E2EE in ways that
are hard to diagnose (peers silently withhold room keys).

_reset_crypto_store_if_device_changed() compares the store's persisted
device ID against the live one at connect time and wipes the store on
mismatch, so a fresh Olm account is generated for the new device instead
of reusing stale key material.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-03 10:03:35 +05:30
webtecnica 799102a4d8 fix(matrix): fallback to homeserver query for room encryption detection (#71067)
_CryptoStateStore.get_encryption_info only consulted mautrix's in-memory
MemoryStateStore, which has no record of m.room.encryption for rooms the
bot joined in the past (the raw-sync path never feeds those state events
through set_encryption_info). On a fresh crypto store this returns None
for all previously-joined rooms, so OlmMachine reports them as unencrypted,
never tracks peer devices, and silently drops all inbound messages.

Fix: pass the mautrix Client into _CryptoStateStore so get_encryption_info
can fall back to a live GET /_matrix/client/v3/rooms/{room_id}/state/
m.room.encryption query when the in-memory store returns None. The result
is cached back via set_encryption_info so subsequent lookups (and
OlmMachine device tracking) hit the fast path.
2026-08-03 10:03:35 +05:30
LFDM Core ff3c9848d4 fix(compression): route overhead-aware tokens in post-tool compress + add recovery-path tests
Post-tool compression path passed context_compressor.last_prompt_tokens (0 in the no-usage
fallback) to _compress_context instead of the overhead-aware _real_tokens computed just above
— same tool-blind bug as the overflow handlers (upstream PR #77169 review, teknium1). Also adds
production-path regression tests asserting the 413, context-overflow (two wordings), and
Anthropic long-context recovery handlers pass estimate_request_tokens_rough(..., tools=...)
(sentinel-patched) to _compress_context.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 10:01:45 +05:30
LFDM Core 4ffc8449a3 fix(compression): overflow handlers pass overhead-aware token size to LCM recovery (issue 441)
Root fix (Option A) for design-session "Context compression exhausted" crashes. The three
compression-retry handlers after an API overflow/413/long-context error (conversation_loop.py
~4229/4488/4747) passed the tool-BLIND messages-only estimate (approx_tokens) to _compress_context,
so hermes-lcm's forced-overflow recovery armed on the message count and missed overflows driven by
tool-schema/system overhead. Now they pass estimate_request_tokens_rough(api_messages, tools=...)
— the same overhead-aware estimator already used at :4580 — so recovery arms on the TRUE request
size; LCM's _overflow_recovery_assembly_cap self-subtracts the overhead so the full request fits.

Empirically validated on real failed session 6dddf1a67b76 (LCM engine, floor=24000/cap=248000):
observed 256,359 >= 248,000 -> arms; recovery 231,313->208,559 msg-tokens -> full request
233,605 < 272,000 FITS. Prior messages-only path did NOT arm (231K < 248K) and crashed.

Durable copy: ~/.hermes/local-patches/optionA-overflow-overhead-aware.patch (survives hermes update
reset). Upstream PR pending. classify_api_error call at :3667 intentionally unchanged (not recovery).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 10:01:45 +05:30
kshitij e8c1882766 refactor(ui-tui): single source for the floating-panel kind set
Review follow-up: $isStatusRuleOccluded and FloatingOverlays each
enumerated the same six overlay kinds — adding a 7th floating panel
required updating both or the timer gate silently missed it. Extracted
hasFloatingPanel as the shared predicate (completions stays local to
FloatingOverlays; it deliberately never occludes the status rule).
Full ui-tui suite green (1487 tests).
2026-08-03 10:00:24 +05:30
briandevans ee8765c930 fix(ui-tui): gate status-rule timers on real occlusion, not $isBlocked
`$isBlocked` answers "is text input suspended", not "is the status rule
covered". appLayout uses it only to hide the input rows (appLayout.tsx:384);
`StatusRulePane` renders outside that guard, at :365 for `at="top"` and :449
for `at="bottom"`. So the previous revision paused FaceTicker /
SessionDuration / IdleSince under prompts that leave the rule fully on
screen — approval, billing, subscription, confirm, clarify, sudo and secret
all render through PromptZone in NORMAL FLOW above ComposerPane
(appOverlays.tsx:58-162, appLayout.tsx:553-568). They push the rule down;
they do not cover it. Freezing a visible clock is a worse bug than the churn
being removed.

Replace it with `$isStatusRuleOccluded`, a narrow derived store over
overlay + ui state covering only what actually paints over the rule:

- `widget` — the modal widget slot renders at viewport level
  (ActiveWidgetSlot, sdk/host.tsx:209) so it can anchor the full-screen
  absolute `Overlay` against the whole terminal.
- the FloatingOverlays set (modelPicker, pager, petPicker, sessions,
  skillsHub, pluginsHub) — but only when `ui.statusBar === 'top'`. That
  panel is `position="absolute" bottom="100%"` inside ComposerPane's
  relative Box (appOverlays.tsx:387), so it grows UPWARD over the top rule
  and can never reach the bottom one.

Deliberately excluded: the PromptZone flow states above; `agents` and
`journey`, which unmount the entire ComposerPane subtree (appLayout.tsx:553)
so React's effect cleanup already clears the intervals; `ambient`, an
in-flow dock; and composer completions, which share the floating grid but
are a render prop that changes per keystroke — re-arming a 1s interval on
every character would restart the countdown each time and starve the tick.
`statusBar: 'off'` needs no branch: StatusRulePane returns null for both
slots, so the timers never mount.

Tests: the store-level cases are re-split into occluding and non-occluding
sets, and an AppLayout-level `describe` mounts the real layout so the rule
sits in its true position — asserting that under approval and sudo the rule
is still rendered AND its clock advances (1m 0s to 1m 30s), that a floating
model picker suppresses the clocks with the rule at the top, and that the
same picker leaves them armed with the rule at the bottom.
2026-08-03 10:00:24 +05:30
briandevans 873af72f8c fix(ui-tui): pause status-chrome timers while a blocking overlay is open
`StatusRulePane` renders `StatusRule` outside the `!isBlocked` guard in
appLayout.tsx, so the status rule stays mounted underneath approval,
model-picker, pager, sessions and every other blocking overlay. Its three
timer-driven components keep firing the whole time: `FaceTicker` (glyph,
1s clock, verb rotation), `SessionDuration` (1s) and `IdleSince` (1s).
Every tick re-renders a rule nobody can see, and in an Ink TUI that churn
reads to the user as the dialog flickering.

Gate all three components' interval creation on the existing `$isBlocked`
computed store, so nothing is armed while an overlay covers the rule.

The pause alone would leave the elapsed read-outs frozen at the moment the
overlay opened, so each effect re-seeds `now` from the wall clock when it
re-arms. `SessionDuration` and `IdleSince` already did this; `FaceTicker`
gains the same re-sync. Closing a five-minute overlay now resumes at the
true elapsed time instead of the pre-overlay value.

No new store is introduced — `$isBlocked` already exists and already ORs
the current OverlayState field set.
2026-08-03 10:00:24 +05:30
kshitij 372494fdf9 style(desktop): eslint --fix import order in sessions-section test
Review follow-up on the #75714 salvage: perfectionist/sort-imports
would fail lint CI; autofixed.
2026-08-03 10:00:20 +05:30
StanleyStetson 416f91d271 fix(test): export VirtualSessionListProps and assert sessions array identity change 2026-08-03 10:00:20 +05:30
StanleyStetson 09ed1ac950 fix(desktop): memoize sidebar flatRows and row renderers to prevent scroll jitter (fixes #73629)
Fix direction inspired by PR #73674 by @drbronson with added Vitest component unit tests.
2026-08-03 10:00:20 +05:30
kshitij 0cbedc58f0 docs: note atexit re-finalization interaction in shutdown() docstring
The PR's skip-live-sessions optimization is partially defeated by
server._shutdown_sessions() registered via atexit (server.py:1172),
which runs on SystemExit after shutdown() returns for the SIGTERM and
stdin_closed paths. The orphan path (os._exit(0)) bypasses atexit.

This is a pre-existing issue — the old finalize-first order had the
same atexit interaction. The comment documents the gap and suggests a
follow-up: gate _shutdown_sessions on not session.get('running').
2026-08-03 09:59:58 +05:30
briandevans eb4f514b2d fix(tui_gateway): retain live-turn sessions unfinalized when the drain deadline expires
The drain reserves a slice of the shutdown budget so flush_all_sessions
still runs when in-flight turns outlast the window. But that flush was
unconditional: a session whose turn was still running got its one-shot
_finalize_session spent mid-turn, and the executor.shutdown(wait=False,
cancel_futures=True) immediately after does not join the turn. The
session was then permanently un-finalizable and its active-session lease
had been released out from under live work — the same persistence and
lifecycle race the drain exists to close, just relocated past the
deadline instead of removed.

Give _turn_futures a session association (Future -> sid, the same key
space as server._sessions) at both submit sites, and on deadline expiry
exclude the sids whose futures are still running from the flush. Those
sessions are retained unfinalized and therefore recoverable; sessions
with no live turn finalize exactly as before. The done-callback now pops
under the lock, since a bare dict.pop is not the drop-in set.discard was.

wait semantics, the reserve math and the bounded per-tick sleep are
unchanged, so this adds no shutdown latency. All three shutdown callers
(orphan, sigterm, and the tight stdin_closed wait=2.0 path) funnel
through this one function and are covered.
2026-08-03 09:59:58 +05:30
briandevans e239184521 fix(tui_gateway): bound the shutdown drain sleep by the time left to it
The drain loop slept a flat 0.05s per tick, so it could overshoot its deadline
by up to one tick and spend part of the reserve withheld for
flush_all_sessions(). For a small `wait` the reserve is itself half the budget,
so a single overshoot can consume all of it: at wait=0.34 the drain budget is
0.17s but the loop requested 4 x 0.05 = 0.20s of sleep.

Clamp each tick to the remaining time. The new test asserts on the summed
*requested* sleep rather than wall-clock, which is deterministic: every sleep is
bounded by the strictly-decreasing remainder, so the total can never exceed the
drain budget regardless of how the scheduler interleaves.
2026-08-03 09:59:58 +05:30
briandevans 1fba2dbe85 fix(tui_gateway): drain in-flight turns before finalizing sessions on compute-host shutdown
ComputeHost.shutdown() called flush_all_sessions() before its own in-flight
turn drain loop. server._finalize_session latches on session["_finalized"]
and every later call returns immediately, so that one flush was spent while
turns were still producing output: the unflushed tail was never persisted,
commit_memory_session wrote long-term memory from a truncated transcript, the
session's DB row was marked ended while it was live, on_session_end fired with
completed=False/interrupted=True against a running session, and the
active-session lease was released out from under a turn. The drain loop exists
precisely so that mid-turn work survives a teardown; finalizing first defeated
it. Reachable from all three teardown paths: the parent/orphan guard (which
os._exit(0)s immediately after), the SIGTERM/SIGINT handler, and stdin close.

Drain first, then flush. A slice of the caller's budget (_FLUSH_RESERVE_SECS,
never more than half of it so a short explicit wait still gets a real drain) is
withheld from the drain so the flush still runs when turns outlast the window:
HostSupervisor SIGKILLs the host _SHUTDOWN_TIMEOUT_SECS after SIGTERM — 10.0s,
the same value as shutdown()'s default wait — so a drain allowed to consume the
whole budget would leave the durability write racing that kill. `wait` itself is
unchanged, so total shutdown latency and the SIGTERM->SIGKILL margin are
unchanged.
2026-08-03 09:59:58 +05:30
HexLab98 28d994f26c test(gateway): cover restart after-turn deferral (#77184) 2026-08-03 09:57:58 +05:30
HexLab98 db3f7e4eb9 fix(gateway): defer in-band restart until active turns finish (#77184)
request_restart was calling stop() immediately, so the requesting turn stayed
in the drain wait set and got force-killed at restart_drain_timeout. Wait for
active work to reach zero first, then stop against an idle gateway.
2026-08-03 09:57:58 +05:30
kshitij f105db2136
Merge pull request #77335 from kshitijk4poor/chore/author-map-atran28
chore: add ATran28 to AUTHOR_MAP
2026-08-03 09:57:46 +05:30
kshitij 307ef0061c chore: add ATran28 to AUTHOR_MAP
at828@proton.me -> ATran28 (GH id=1445620)
Needed for PR #77270 whatsapp bridge reconnect fix.
2026-08-03 09:57:30 +05:30
kshitij 622c220ebb refactor(agent): rebuild hoisted patterns from shared tag-name tuples
Simplify-pass follow-up on the #69653 salvage: the original code built
these patterns from a name loop; hand-expanding them into 10 literals
lost that single source. Tag-name tuples restore it (adding a 6th
reasoning tag is now a one-place change), and the gnarly named-function
pattern regained a pointer to its step-1c rationale. Byte-equivalence
of every rebuilt pattern verified programmatically (alternation-order
neutrality probed: the \b and > anchors make order irrelevant).
2026-08-03 09:56:36 +05:30
Eugeniusz Gilewski 9421c5afdf perf(agent): precompile response and skill-scan regexes (#33208)
strip_think_blocks passed the same response-scrubbing strings through
re's pattern dispatcher on every response. Skills Guard repeated the
same work for 121 patterns against every scanned line.

Compile the existing expressions once and reuse Pattern.sub/search. Keep
each generic tool-call tag in its own paired expression so mismatched
openers retain their payload while existing stray-closer cleanup remains
unchanged.

Part of #33208
Salvaged from #32713 by @ErnestHysa.

Co-authored-by: ErnestHysa <takis312@hotmail.com>
2026-08-03 09:56:36 +05:30
joaomarcos 3b53a3c560 chore: drop stale diagnostic report from PR #77021
teknium1 flagged ISSUE_76870_RELATORIO_CAUSA_RAIZ.md as containing
stale metadata (references an unrelated local branch) and asked to
drop the standalone report, keeping only the focused server.py fix
and regression test.
2026-08-03 09:56:24 +05:30
joaomarcos a93969dbd1 fix(tui): snapshot history after pending model switch applies (#76870)
Deferred model switches append a marker and bump history_version at
turn start; the dispatcher was snapshotting history before that
mutation, so the version-mismatch guard rejected the turn's own
result as a stale/concurrent write. Move the snapshot to after
_apply_pending_model_switch/_sync_agent_model_with_config, under
history_lock, so the turn's own preparatory mutation is included in
its baseline while the anti-stale guard still catches real external
writes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-03 09:56:24 +05:30
kshitij 8f61e61991
chore: add contributor email mapping for LFDMcore (core@lfdm.co) (#77331) 2026-08-03 09:55:44 +05:30
kshitij 1228bfb270
Merge pull request #77329 from kshitijk4poor/chore/attrib-ckaznocha
chore: AUTHOR_MAP for ckaznocha
2026-08-03 09:54:46 +05:30
kshitij b6cdddd3df chore: AUTHOR_MAP for ckaznocha (matrix crypto PRs) 2026-08-03 09:54:28 +05:30
Gabriel Lesperance 68391930d0 perf(tui): bound scroll rendering and preserve anchors 2026-08-03 09:53:51 +05:30