* fix(llm): forward provider_params.timeout to the OpenAI-compatible embedding client #832 and #903 added a configurable request timeout for the LLM registry and the Gemini embedding client respectively, but the OpenAI-compatible embedding client (src/embedding_client.py) was never wired up. It constructed AsyncOpenAI with no timeout at all, so a stalled socket against a slow or contended OpenAI-compatible backend (e.g. a self-hosted embedding model under load) wedges the deriver worker's event loop indefinitely — the exact failure #785/#903 describe, just via a code path #903 didn't cover. EmbeddingModelConfig now carries provider_params through from resolve_embedding_model_config, mirroring how resolve_model_config already does it for ModelConfig, and the OpenAI branch of _EmbeddingClient.__init__ extracts `timeout` via the existing request_timeout_from_extra_params helper. Unset stays unset — no existing behavior changes. Reproduced and verified against a real self-hosted deployment (local Ollama backend under load): before this fix, a single stuck embedding call blocked all deriver queue processing for 20+ minutes with no error logged, twice in one session. * fix(embedding): use first-class timeout on embedding model config provider_params is the LLM per-request escape hatch; embedding timeouts are client-construction knobs and belong next to max_batch_size. Wire the field for OpenAI and Gemini, omit the OpenAI kwarg when unset so the SDK default stays, and keep Gemini's 10-minute floor when unset. * test(embedding): live coverage for first-class embedding timeout Exercise EmbeddingModelConfig.timeout on one representative OpenAI and Gemini model: configured timeout lands on the SDK client, and a near-zero timeout aborts before the provider answers. --------- Co-authored-by: Aakash Kattelu <aakash@plasticlabs.ai> |
||
|---|---|---|
| .. | ||
| api-reference | ||
| contributing | ||
| documentation | ||
| guides | ||
| migrations | ||
| README.md | ||
| openapi.json | ||
README.md
This subdirectory contains the peer-paradigm documentation for Honcho (Honcho v2.0.0 onwards).