- warn (not debug) on final text-turn flush failure: a failure here
reopens the exact #81641 data-loss window with _persist_session as
the only remaining retry, unlike the verify siblings which retry
in-loop; include session id for triage
- trim the flush-site comment to sibling proportion, pointing to the
test module for the full incident narrative
- test: assert _persist_session presence before indexing, so a wiring
change fails with a clean assertion instead of ValueError from max()
A pure-text assistant turn (finish_reason=stop) had no durable write of
its own. Its answer reached the user through the streaming / interim
display path, which is display-only and never touches state.db, and the
first durable write was finalize_turn's _persist_session — after the
loop exits and behind post-turn work that can include micro-compaction's
aux-LLM call.
Anything that ended the process or tore the session down inside that
window lost a reply the user had already been shown. On a remote
(non-loopback) backend the window is easy to hit: WS 1006 closures drive
ws_orphan_reap teardown, and affected sessions ended up with user rows
and zero assistant rows in state.db.
The neighbouring exits of the same loop already close this gap:
* the tool-call exit flushes the assistant(tool_calls) block before
handing control to _execute_tool_calls (#49045)
* the verify-on-stop and pre_verify exits flush final_msg before
appending their nudge (#65919 §7)
Apply that same idiom to the ordinary text exit rather than adding a new
persistence mechanism. The intrinsic _DB_PERSISTED_MARKER dedup makes the
later _persist_session a no-op for this row, so no duplicate rows and no
extra write — the same write, just earlier.
Unlike the tool-call exit, a failed flush must not abort the turn: no
side effect runs after this point and the answer is already produced, so
the failure is logged and _persist_session remains the retry.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>