honcho/tests
Tomas Sereikis 5535415e89 perf(llm): run one turn's tool calls concurrently instead of serially
A single assistant turn routinely requests several independent reads — a
dialectic turn typically emits `search_memory` and `search_messages` together —
but the loop awaited them one at a time, so the turn cost their sum when it only
needed to cost the slowest. Measured on a production `minimal` dialectic: 1.2s
in `search_memory`, then a further 1.9s in `search_messages`, both pure reads.

Measured across 42 production dialectic requests (103 iterations, 77 of them
multi-tool, 3.34 tool calls per iteration on average), running each turn's calls
concurrently would cut 11.3% of total wall-clock: median 9.3% per request, p90
29.7%, best case 51.0%. Only 4 of the 42 requests gain nothing.

This is safe because tool handlers own their sessions: each opens a short-lived
one via `tracked_db()`, and only the mutating handlers (`create_observations`,
`update_peer_card`, `delete_observations`) take `ctx.db_lock`, so concurrent
reads do not contend on shared state.

The per-call telemetry ContextVars move inside the task. `asyncio` copies the
context per task, so `set_current_tool_call_seq` and `set_last_tool_metadata`
now bind to their own call instead of being written and read across one shared
context — which the previous code could only keep straight by never overlapping.
`gather` preserves argument order, so `tool_results` and `all_tool_calls` stay in
the order the model asked for them.

Fan-out is capped at MAX_CONCURRENT_TOOL_CALLS (4). Production has produced 18
tool calls in one iteration, and firing all of them at once would mean that many
simultaneous embedding + pgvector queries on a single instance. The cap leaves
the measured common case fully parallel while bounding the tail.
2026-08-29 13:26:19 +03:00
..
alembic feat: make session_name nullable for documents and update related SDKs (#347) 2026-01-26 13:33:11 -05:00
bench telemetry: zero-initialize bounded-label metrics so an absent series means a broken scrape (#927) 2026-08-20 10:24:49 -04:00
crud fix: stop top_k=0 from reaching Turbopuffer on message search (#1084) 2026-08-26 17:07:41 -04:00
deriver fix(embedding): truncate in batch embed and return results breakdown (#1019) 2026-08-20 11:42:38 -04:00
dialectic Scopes Phase 2b: `scope` option on chat, representation, context, and search (#897) 2026-08-12 18:34:55 -04:00
dreamer Scopes Phase 2a: scope-kind peers, guardrails, and scopes CRUD routes (#884) 2026-08-12 16:21:40 -04:00
integration feat: defer embedding messages (#704) 2026-06-11 10:31:04 -04:00
live_llm fix(openai): fix content normalization in openai backend history adapter (#1064) 2026-08-26 12:14:32 -04:00
llm perf(llm): run one turn's tool calls concurrently instead of serially 2026-08-29 13:26:19 +03:00
reconciler telemetry: zero-initialize bounded-label metrics so an absent series means a broken scrape (#927) 2026-08-20 10:24:49 -04:00
routes feat: Add workspace-level chat (#931) 2026-08-24 15:54:23 -04:00
scripts feat: add new cloudevents for api routes (#637) 2026-05-20 18:25:30 -04:00
sdk Scopes SDK Changes (#1030) 2026-08-19 10:48:29 -04:00
sdk_typescript add read db (#773) 2026-06-10 13:28:36 -04:00
startup add read db (#773) 2026-06-10 13:28:36 -04:00
telemetry review: fix db_connections_open leak on the detach path; comment/doc polish 2026-08-24 14:37:20 -04:00
unified feat: Add workspace-level chat (#931) 2026-08-24 15:54:23 -04:00
utils fix: stop top_k=0 from reaching Turbopuffer on message search (#1084) 2026-08-26 17:07:41 -04:00
vector_store fix: stop top_k=0 from reaching Turbopuffer on message search (#1084) 2026-08-26 17:07:41 -04:00
webhooks Tighten Transaction Scopes (#525) 2026-04-08 11:14:50 -04:00
__init__.py Refactor clients.py to add modern features and more flexible configuration (#459) 2026-04-20 02:46:37 -04:00
conftest.py fix(scopes): scope observer sessions in SQL instead of a fetched name list (#1065) 2026-08-25 13:24:31 -04:00
test_advanced_filters.py fix(filter): make ne on jsonb metadata keys null-safe (#1036) 2026-08-24 09:33:08 -04:00
test_cache_key_namespace.py perf(cache): hash-tag the cache namespace so an instance uses one shard (#1058) 2026-08-25 12:34:42 -04:00
test_cache_redaction.py Scopes Phase 2a: scope-kind peers, guardrails, and scopes CRUD routes (#884) 2026-08-12 16:21:40 -04:00
test_config.py fix(llm): support per-request provider timeouts (#832) 2026-08-04 12:43:00 -04:00
test_datetime_parsing.py Make embeddings configurable (#678) 2026-05-14 15:03:35 -04:00
test_db_resilience.py fix(deriver): increase deriver polling backoff (#1015) 2026-08-14 16:38:20 -04:00
test_dependencies.py add read db (#773) 2026-06-10 13:28:36 -04:00
test_dialectic_prompts.py fix(dialectic): revamp workspace and pair chat system prompts (#1066) 2026-08-25 12:30:06 -04:00
test_generate_jwt_script.py feat: add generate_jwt.py script for creating scoped JWTs (#757) 2026-06-09 13:49:55 -04:00
test_models_vector_dim.py feat: add new cloudevents for api routes (#637) 2026-05-20 18:25:30 -04:00
test_schema_validations.py Align API contract with DB contract for IDs (#684) 2026-05-14 16:37:39 -04:00
test_search.py fix(crud): preserve joined_at for active session peers (#1059) 2026-08-25 10:51:59 -04:00
test_security.py feat: Add workspace-level chat (#931) 2026-08-24 15:54:23 -04:00
test_session_allowlist.py fix(filter): make the filter DSL reject bad input instead of 500ing, and fix negation over unset fields (#947) 2026-08-13 10:39:19 -04:00
test_workspace_chat.py fix(dialectic): revamp workspace and pair chat system prompts (#1066) 2026-08-25 12:30:06 -04:00