A single assistant turn routinely requests several independent reads — a
dialectic turn typically emits `search_memory` and `search_messages` together —
but the loop awaited them one at a time, so the turn cost their sum when it only
needed to cost the slowest. Measured on a production `minimal` dialectic: 1.2s
in `search_memory`, then a further 1.9s in `search_messages`, both pure reads.
Measured across 42 production dialectic requests (103 iterations, 77 of them
multi-tool, 3.34 tool calls per iteration on average), running each turn's calls
concurrently would cut 11.3% of total wall-clock: median 9.3% per request, p90
29.7%, best case 51.0%. Only 4 of the 42 requests gain nothing.
This is safe because tool handlers own their sessions: each opens a short-lived
one via `tracked_db()`, and only the mutating handlers (`create_observations`,
`update_peer_card`, `delete_observations`) take `ctx.db_lock`, so concurrent
reads do not contend on shared state.
The per-call telemetry ContextVars move inside the task. `asyncio` copies the
context per task, so `set_current_tool_call_seq` and `set_last_tool_metadata`
now bind to their own call instead of being written and read across one shared
context — which the previous code could only keep straight by never overlapping.
`gather` preserves argument order, so `tool_results` and `all_tool_calls` stay in
the order the model asked for them.
Fan-out is capped at MAX_CONCURRENT_TOOL_CALLS (4). Production has produced 18
tool calls in one iteration, and firing all of them at once would mean that many
simultaneous embedding + pgvector queries on a single instance. The cap leaves
the measured common case fully parallel while bounding the tail.