Merge remote-tracking branch 'origin/main' into abigail/dev-2219

This commit is contained in:
ajspig 2026-08-04 15:25:37 -04:00
commit 96100839a7
17 changed files with 735 additions and 10 deletions

View File

@ -0,0 +1,103 @@
---
name: verify
description: Build, launch, and drive a local Honcho stack to verify a change at its runtime surface (the /v3 HTTP API and the deriver queue). Use when verifying a diff or confirming a change works in the running app.
---
# Verifying changes in Honcho
## Prerequisites
Docker Compose is the preferred way to run the stack — `docker-compose.yml` at
the repo root brings up Postgres (pgvector), Redis, the API server, and the
deriver worker together. If Docker Compose isn't available, the stack can also
run directly on the host (see Launch below); you'll need Postgres with pgvector
and Redis reachable, plus a `.env` with connection strings and an LLM provider
key for any flow that hits a model (deriver, dialectic, dreamer).
Working in a worktree? It carries neither `.env` nor `node_modules`. Copy
`.env` from the main checkout (that's where the provider keys live), and run
`bun install` in `sdks/typescript` if you'll run the full test suite —
otherwise the pre-push gate fails on a phantom `Cannot find package 'zod'`.
## Launch
First check whether the stack is already running via Docker Compose:
```bash
docker compose ps # look for api / deriver / database / redis
docker compose up -d # start it if not
```
A running stack is not the same as *your branch's code* running — check the
CREATED column; the images may be weeks old. For verifying a diff, the cheap
path is to reuse the stack's Postgres/Redis containers but run the branch's
API as a host process on a spare port:
```bash
uv run uvicorn src.main:app --port 8901 # branch code, stack's DB/Redis
uv run python -m src.deriver # if the diff touches the worker
```
This avoids both an image rebuild and the port/project-name conflicts a second
compose stack in a worktree would cause. Without Docker at all, run the two
processes the same way against host Postgres + Redis (API default port 8000
via `uv run fastapi dev src/main.py`).
When reading server logs, note that with the main `.env` the telemetry emitter
spams connection warnings at an unreachable endpoint — `grep -v
telemetry.emitter` before looking for the real error.
## Drive it
Prefer driving through the `honcho-cli` skill; fall back to the SDKs
(`sdks/python`, `sdks/typescript`), then raw REST against `/v3`, in that order
when a method isn't available at the higher level. The `honcho-integration`
skill covers how to use the SDKs.
When a change updates endpoints or makes schema changes, verify across all
three surfaces — CLI, SDKs, and REST — since they can drift independently.
For LLM-path changes, the fastest synchronous surface is dialectic at the
`minimal` reasoning level: send a couple of messages, then hit
`/v3/.../peers/{peer_id}/chat` and observe the response.
## Configuration
Configuration is central to both driving the app and running tests: it's how
API keys reach the server and how a config-related change gets exercised at
all. Settings come from environment variables or files, with precedence
env > `.env` > `config.toml` > defaults. To verify a configuration change,
set the relevant option through one of these layers, restart the affected
process, and observe the behavioral difference at the surface — the same
mechanism lets you point provider base URLs, model choices, and timeouts at
credentials and proxies you actually have.
Two concrete levers: deep-nested settings (e.g.
`[dialectic.levels.minimal.model_config.overrides.provider_params]`) are
miserable as env vars — drop a partial `config.toml` in the repo root instead.
And for load-time config behavior, `uv run python -c "import src.config; ..."`
is faster than booting the server.
## Test suites
Three test types matter here. All run in CI, but they're also runnable locally
with whatever API keys and configuration you have — Honcho's config surface is
large, so options like per-agent timeouts and provider base URLs can be pointed
at your own keys/proxies to exercise a change:
```bash
# Unit tests (pytest; spins up its own infra via fixtures)
uv run pytest tests/
# Unified tests — step-based end-to-end flows defined in JSON
# (config hierarchy, multi-turn interactions, LLM-as-judge assertions)
uv run python -m tests.unified.run
uv run python -m tests.unified.run --test-dir tests/unified/test_cases
# Live LLM tests — real provider calls, for testing specific backends.
# Needs provider API keys AND model vars — without LIVE_LLM_*_MODELS the
# tests silently deselect that provider. See tests/live_llm/README.md.
export LLM_ANTHROPIC_API_KEY=...
export LIVE_LLM_ANTHROPIC_45_PLUS_MODELS=claude-sonnet-4-5
uv run pytest tests/live_llm -n 0 --live-llm --no-header -q
```

View File

@ -138,6 +138,7 @@ model = "gpt-5.4-mini"
# api_key_env = "DERIVER_CUSTOM_BACKUP_API_KEY"
# [deriver.model_config.overrides.provider_params]
# verbosity = "low"
# timeout = 3600.0
# Peer card settings
[peer_card]

View File

@ -189,8 +189,29 @@ Each model config supports an `overrides.provider_params` dict for passing arbit
[deriver.model_config.overrides.provider_params]
# These are passed directly to the provider SDK
verbosity = "low"
# Per-request timeout in seconds; useful for queued workers that can wait longer
timeout = 3600.0
```
Because provider params live on each model config, background workers such as
the Deriver and Dreamer can use longer request timeouts while synchronous
chat paths keep tighter defaults.
`timeout` gotchas:
- The value is validated **at config load**: it must coerce to a positive,
finite number of seconds (numbers or numeric strings like `"3600"`), or the
process refuses to start with an error naming the offending config path.
This applies to both the primary model config and its `fallback.overrides`.
- The unit is always **seconds**, regardless of transport. OpenAI and
Anthropic receive it as the SDK's `timeout` kwarg; Gemini has no such
kwarg, so Honcho converts it to milliseconds on `http_options.timeout`.
- When unset, nothing is forwarded and each SDK's default applies — adding
this key is opt-in and changes no existing behavior.
- A too-tight timeout doesn't fail once: the aborted request goes through the
normal retry/fallback chain before the caller sees an error, so the
observed latency is several multiples of the timeout.
#### Transport passthrough keys
Three keys inside `provider_params` are recognized as request-level escape hatches and forwarded to the underlying transport. Where a transport actually validates and merges one of these keys, its value must be a mapping — a non-mapping value raises a configuration error (see the per-transport behavior below; a key a transport ignores is not validated):

View File

@ -1,4 +1,5 @@
import logging
import math
import os
from pathlib import Path
from typing import Annotated, Any, ClassVar, Literal, cast
@ -66,6 +67,37 @@ ThinkingEffortLevel = Literal[
StructuredOutputMode = Literal["json_schema", "json_object"]
PROVIDER_TIMEOUT_ERROR_TEXT = (
"provider_params.timeout must be a positive number of seconds"
)
def coerce_provider_timeout(value: Any) -> float:
"""Coerce a `provider_params.timeout` value to positive, finite seconds.
Canonical implementation shared by config-load validation (here) and
per-request validation (`src.llm.request_builder.request_timeout_from_extra_params`,
which translates the ValueError into a ValidationException). Lives in
config.py because src.exceptions imports src.config, so config validators
cannot raise Honcho exception types.
"""
if isinstance(value, bool):
raise ValueError(PROVIDER_TIMEOUT_ERROR_TEXT)
if isinstance(value, int | float):
timeout = float(value)
elif isinstance(value, str):
try:
timeout = float(value.strip())
except ValueError as exc:
raise ValueError(PROVIDER_TIMEOUT_ERROR_TEXT) from exc
else:
raise ValueError(PROVIDER_TIMEOUT_ERROR_TEXT)
if not math.isfinite(timeout) or timeout <= 0:
raise ValueError(PROVIDER_TIMEOUT_ERROR_TEXT)
return timeout
class ModelOverrideSettings(BaseModel):
"""Advanced module-level transport overrides."""
@ -91,6 +123,14 @@ class ModelOverrideSettings(BaseModel):
),
)
@field_validator("provider_params")
@classmethod
def _validate_provider_timeout(cls, v: dict[str, Any]) -> dict[str, Any]:
"""Reject bad `timeout` values at config load; normalize good ones to float."""
if "timeout" not in v:
return v
return {**v, "timeout": coerce_provider_timeout(v["timeout"])}
class PromptCachePolicy(BaseModel):
"""Per-call prompt-caching configuration.

View File

@ -355,6 +355,9 @@ async def query_documents(
Returns:
Sequence of matching documents
"""
if top_k <= 0:
return []
# Use provided embedding or generate one
if embedding is None:
try:

View File

@ -318,11 +318,12 @@ class RepresentationManager:
total = max_observations
# Calculate how many observations to get from each source
# Calculate how many observations to get from each source.
# Floor of 1 when a semantic query was explicitly requested.
semantic_observations = (
min(
max(
0,
1,
semantic_search_top_k
if semantic_search_top_k is not None
else total // 3,

View File

@ -9,7 +9,10 @@ from anthropic.types import TextBlock, ThinkingBlock, ToolUseBlock
from pydantic import BaseModel, ValidationError
from src.llm.backend import CompletionResult, StreamChunk, ToolCallResult
from src.llm.request_builder import apply_sdk_passthroughs
from src.llm.request_builder import (
apply_sdk_passthroughs,
request_timeout_from_extra_params,
)
from src.llm.structured_output import repair_response_model_json, schema_instruction
@ -74,6 +77,10 @@ class AnthropicBackend:
# from ModelConfig.provider_params. Shallow merge with operator-wins.
apply_sdk_passthroughs(params, extra_params)
timeout = request_timeout_from_extra_params(extra_params)
if timeout is not None:
params["timeout"] = timeout
# The '{' prefill forces a JSON-first response, which suppresses
# tool_use blocks — skip it when tools are available and rely on the
# conditional instruction + repair fallback instead.
@ -157,6 +164,11 @@ class AnthropicBackend:
# Operator escape hatch: forward Anthropic SDK passthrough kwargs
# from ModelConfig.provider_params. Shallow merge with operator-wins.
apply_sdk_passthroughs(params, extra_params)
timeout = request_timeout_from_extra_params(extra_params)
if timeout is not None:
params["timeout"] = timeout
# See complete(): no '{' prefill when tools are available, so
# tool_use blocks stay reachable on the streamed path too.
use_json_prefill = (

View File

@ -4,6 +4,7 @@ from collections.abc import AsyncIterator
from datetime import datetime, timedelta, timezone
from typing import Any, ClassVar, cast
from google.genai import types as genai_types
from pydantic import BaseModel
from src.exceptions import LLMError, ValidationException
@ -14,7 +15,10 @@ from src.llm.caching import (
build_cache_key,
gemini_cache_store,
)
from src.llm.request_builder import coerce_passthrough_mapping
from src.llm.request_builder import (
coerce_passthrough_mapping,
request_timeout_from_extra_params,
)
from src.llm.structured_output import repair_response_model_json, schema_instruction
GEMINI_BLOCKED_FINISH_REASONS = {
@ -289,19 +293,38 @@ class GeminiBackend:
# extra_query has no SDK-level equivalent and is ignored. Shallow
# merge with operator-wins. Operators are responsible for not setting
# unknown fields that google-genai's validation will reject.
http_options: genai_types.HttpOptions | None = None
if extra_params:
operator_extra_body = extra_params.get("extra_body")
if operator_extra_body:
config.update(
coerce_passthrough_mapping("extra_body", operator_extra_body)
)
raw_http_options = config.get("http_options")
if isinstance(raw_http_options, genai_types.HttpOptions):
http_options = raw_http_options
elif isinstance(raw_http_options, dict):
http_options = genai_types.HttpOptions.model_validate(
raw_http_options
)
operator_extra_headers = extra_params.get("extra_headers")
if operator_extra_headers:
http_options = config.setdefault("http_options", {})
existing_headers = http_options.setdefault("headers", {})
if http_options is None:
http_options = genai_types.HttpOptions()
existing_headers = dict(http_options.headers or {})
existing_headers.update(
coerce_passthrough_mapping("extra_headers", operator_extra_headers)
)
http_options.headers = existing_headers
timeout = request_timeout_from_extra_params(extra_params)
if timeout is not None:
if http_options is None:
http_options = genai_types.HttpOptions()
# Gemini has no native timeout kwarg; set the httpx-level value in ms.
http_options.timeout = int(timeout * 1000)
if http_options is not None:
config["http_options"] = http_options
return config
def _normalize_response(

View File

@ -11,7 +11,10 @@ from pydantic import BaseModel, ValidationError
from src.exceptions import ValidationException
from src.llm.backend import CompletionResult, StreamChunk, ToolCallResult
from src.llm.request_builder import apply_sdk_passthroughs
from src.llm.request_builder import (
apply_sdk_passthroughs,
request_timeout_from_extra_params,
)
from src.llm.structured_output import (
StructuredOutputError,
empty_structured_output,
@ -397,6 +400,10 @@ class OpenAIBackend:
# if the operator supplies `extra_body.reasoning`, it replaces any
# value Honcho auto-injected above.
apply_sdk_passthroughs(params, extra_params)
timeout = request_timeout_from_extra_params(extra_params)
if timeout is not None:
params["timeout"] = timeout
return params
def _normalize_response(

View File

@ -11,10 +11,14 @@ from typing import Any, cast
from pydantic import BaseModel
from src.config import ModelConfig, PromptCachePolicy
from src.config import ModelConfig, PromptCachePolicy, coerce_provider_timeout
from src.exceptions import ValidationException
from .backend import CompletionResult, ProviderBackend, StreamChunk
from .backend import (
CompletionResult,
ProviderBackend,
StreamChunk,
)
# Operator escape-hatch keys recognized inside ModelConfig.provider_params.
PASSTHROUGH_KEYS = ("extra_body", "extra_headers", "extra_query")
@ -99,6 +103,45 @@ def build_config_extra_params(config: ModelConfig) -> dict[str, Any]:
return extra_params
def request_timeout_from_extra_params(
extra_params: dict[str, Any] | None,
) -> float | None:
"""Return a validated per-request provider timeout from extra params.
Config-sourced timeouts are already validated and normalized at config
load (`coerce_provider_timeout` in src.config); this guards extra_params
passed programmatically at call time.
"""
if not extra_params or "timeout" not in extra_params:
return None
try:
return coerce_provider_timeout(extra_params["timeout"])
except ValueError as exc:
raise ValidationException(str(exc)) from exc
def _strip_none_params(
params: dict[str, Any],
keys: tuple[str, ...],
) -> dict[str, Any]:
"""Remove specified keys from extra params when their values are None."""
return {k: v for k, v in params.items() if not (k in keys and v is None)}
def _normalize_extra_params(extra_params: dict[str, Any]) -> dict[str, Any]:
"""Normalize and clean shared extra params before they reach backends.
Centralizes per-key coercion and null-stripping so new keys are added
here rather than spawning one-off normalizers.
"""
result = dict(extra_params)
timeout = request_timeout_from_extra_params(result)
if timeout is not None:
result["timeout"] = timeout
return _strip_none_params(result, ("timeout",))
async def execute_completion(
backend: ProviderBackend,
config: ModelConfig,
@ -120,6 +163,7 @@ async def execute_completion(
**build_config_extra_params(config),
**(extra_params or {}),
}
merged_extra_params = _normalize_extra_params(merged_extra_params)
if cache_policy is not None:
merged_extra_params["cache_policy"] = cache_policy
@ -158,6 +202,7 @@ async def execute_stream(
**build_config_extra_params(config),
**(extra_params or {}),
}
merged_extra_params = _normalize_extra_params(merged_extra_params)
if cache_policy is not None:
merged_extra_params["cache_policy"] = cache_policy

View File

@ -629,3 +629,72 @@ class TestRepresentationManagerSave:
assert len(saved.created_documents) == 0
mock_embed.assert_not_awaited()
mock_save.assert_not_awaited()
class TestVectorQueryTopKFloor:
"""Regression for HONCHO-19Q / HONCHO-4Q4.
A top_k of 0 reached Turbopuffer, which rejects it with a 400
('top_k must be between 1 and 10000'). Two independent paths produced it:
the working-representation budget split (``total // 3`` rounds to 0 for
max_conclusions < 3) and the dialectic ``search_memory`` tool, whose
LLM-supplied top_k has an upper clamp but no floor.
"""
@pytest.mark.asyncio
async def test_query_documents_returns_empty_without_querying_on_zero_top_k(self):
"""The choke point every semantic document query routes through."""
from src.crud.document import query_documents
with (
patch(
"src.crud.document.embedding_client.embed", new=AsyncMock()
) as mock_embed,
patch(
"src.crud.document.query_external_vector_document_ids",
new=AsyncMock(),
) as mock_vector,
):
for top_k in (0, -1):
assert (
await query_documents(
None,
"workspace",
"query",
observer="observer",
observed="observed",
top_k=top_k,
)
== []
)
mock_embed.assert_not_awaited()
mock_vector.assert_not_awaited()
@pytest.mark.asyncio
async def test_requested_semantic_search_always_gets_budget(
self,
db_session: AsyncSession,
sample_data: tuple[models.Workspace, models.Peer],
):
"""max_conclusions < 3 must not allocate 0 to an explicitly requested search."""
test_workspace, test_peer = sample_data
manager = RepresentationManager(
test_workspace.name, observer=test_peer.name, observed=test_peer.name
)
for max_observations in (1, 2, 100):
with patch(
"src.crud.query_documents", new=AsyncMock(return_value=[])
) as mock_query:
await manager._get_working_representation_internal( # pyright: ignore[reportPrivateUsage]
db_session,
include_semantic_query="what do they like?",
embedding=[0.1],
max_observations=max_observations,
)
assert mock_query.await_args is not None
top_k = mock_query.await_args.kwargs["top_k"]
assert top_k >= 1, f"max_observations={max_observations} gave top_k={top_k}"
assert top_k <= max_observations

View File

@ -0,0 +1,117 @@
from __future__ import annotations
import time
from typing import Any
import anthropic
import httpx
import openai
import pytest
from src.llm.request_builder import execute_completion
from .conftest import make_backend, require_provider_key, wrap_async_method
from .model_matrix import LiveModelSpec, ProviderName, get_live_model_specs
pytestmark = [pytest.mark.live_llm]
GENEROUS_TIMEOUT_SECONDS = 120
TIGHT_TIMEOUT_SECONDS = 0.01
# Well under the 600s client default; generous enough to absorb SDK retries.
TIGHT_TIMEOUT_WALL_CLOCK_LIMIT_SECONDS = 30
TIMEOUT_EXCEPTIONS: dict[ProviderName, tuple[type[BaseException], ...]] = {
"anthropic": (anthropic.APITimeoutError,),
"openai": (openai.APITimeoutError,),
# google-genai raises httpx or aiohttp timeouts depending on its transport;
# aiohttp surfaces as asyncio.TimeoutError (== builtins.TimeoutError).
"gemini": (httpx.TimeoutException, TimeoutError),
}
PROVIDER_MARKS = {
"anthropic": pytest.mark.requires_anthropic,
"openai": pytest.mark.requires_openai,
"gemini": pytest.mark.requires_gemini,
}
def representative_specs() -> list[Any]:
"""One spec per provider — timeout plumbing is transport-level, not model-level."""
params: list[Any] = []
for provider in ("anthropic", "openai", "gemini"):
specs = get_live_model_specs(provider=provider)
if not specs:
continue
params.append(
pytest.param(specs[0], marks=PROVIDER_MARKS[provider], id=specs[0].id)
)
return params
def assert_timeout_reached_sdk(
model_spec: LiveModelSpec, call_kwargs: dict[str, Any], timeout_seconds: float
) -> None:
if model_spec.provider == "gemini":
http_options = call_kwargs["config"]["http_options"]
assert http_options.timeout == int(timeout_seconds * 1000)
else:
assert call_kwargs["timeout"] == timeout_seconds
def sdk_call_target(backend: Any, model_spec: LiveModelSpec) -> tuple[Any, str]:
if model_spec.provider == "gemini":
return backend._client.aio.models, "generate_content"
if model_spec.provider == "anthropic":
return backend._client.messages, "create"
return backend._client.chat.completions, "create"
@pytest.mark.asyncio
@pytest.mark.parametrize("model_spec", representative_specs())
async def test_live_provider_timeout_reaches_the_wire(
model_spec: LiveModelSpec,
monkeypatch: pytest.MonkeyPatch,
) -> None:
require_provider_key(model_spec)
backend, config = make_backend(
model_spec, provider_params={"timeout": GENEROUS_TIMEOUT_SECONDS}
)
target, attribute = sdk_call_target(backend, model_spec)
calls = wrap_async_method(monkeypatch, target, attribute)
result = await execute_completion(
backend,
config,
messages=[{"role": "user", "content": "Reply with the single word: ok"}],
max_tokens=256,
)
assert isinstance(result.content, str)
assert result.content.strip()
assert len(calls) == 1
assert_timeout_reached_sdk(model_spec, calls[0]["kwargs"], GENEROUS_TIMEOUT_SECONDS)
@pytest.mark.asyncio
@pytest.mark.parametrize("model_spec", representative_specs())
async def test_live_tight_provider_timeout_aborts_request(
model_spec: LiveModelSpec,
) -> None:
require_provider_key(model_spec)
backend, config = make_backend(
model_spec, provider_params={"timeout": TIGHT_TIMEOUT_SECONDS}
)
started = time.monotonic()
with pytest.raises(TIMEOUT_EXCEPTIONS[model_spec.provider]):
await execute_completion(
backend,
config,
messages=[{"role": "user", "content": "Reply with the single word: ok"}],
max_tokens=256,
)
elapsed = time.monotonic() - started
assert (
elapsed < TIGHT_TIMEOUT_WALL_CLOCK_LIMIT_SECONDS
), f"tight timeout took {elapsed:.1f}s — per-request timeout likely not applied"

View File

@ -449,3 +449,82 @@ async def test_anthropic_backend_stream_no_prefill_when_tools_present() -> None:
"If not responding with a tool call, respond with valid JSON"
in call["messages"][0]["content"]
)
@pytest.mark.asyncio
async def test_anthropic_backend_passes_timeout_to_completion_request() -> None:
"""Anthropic completion requests receive per-request provider timeout."""
client = Mock()
client.messages.create = AsyncMock(
return_value=SimpleNamespace(
content=[TextBlock(type="text", text="ok")],
usage=SimpleNamespace(
input_tokens=10,
output_tokens=5,
cache_creation_input_tokens=0,
cache_read_input_tokens=0,
),
stop_reason="end_turn",
)
)
backend = AnthropicBackend(client)
await backend.complete(
model="claude-haiku-4-5",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
extra_params={"timeout": 45},
)
await_args = client.messages.create.await_args
if await_args is None:
raise AssertionError("Expected Anthropic create call")
assert await_args.kwargs["timeout"] == 45.0
@pytest.mark.asyncio
async def test_anthropic_backend_passes_timeout_to_stream_request() -> None:
"""Anthropic stream requests receive per-request provider timeout."""
class FakeStream:
"""Minimal async stream manager for Anthropic streaming tests."""
async def __aenter__(self):
"""Return the stream object used by the backend."""
return self
async def __aexit__(self, *_args: object) -> bool:
"""Do not suppress stream errors."""
return False
def __aiter__(self):
"""Return the async iterator used by the backend."""
return self
async def __anext__(self):
"""End the fake stream immediately."""
raise StopAsyncIteration
async def get_final_message(self):
"""Return the final message required by the backend."""
return SimpleNamespace(
usage=SimpleNamespace(output_tokens=1),
stop_reason="end_turn",
)
client = Mock()
client.messages.stream = Mock(return_value=FakeStream())
backend = AnthropicBackend(client)
chunks = [
chunk
async for chunk in backend.stream(
model="claude-haiku-4-5",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
extra_params={"timeout": "60"},
)
]
assert chunks[-1].is_done is True
assert client.messages.stream.call_args.kwargs["timeout"] == 60.0

View File

@ -98,6 +98,40 @@ async def test_gemini_backend_maps_thinking_effort_to_thinking_level() -> None:
assert call["config"]["thinking_config"] == {"thinking_level": "low"}
@pytest.mark.asyncio
async def test_gemini_backend_maps_timeout_to_http_options() -> None:
"""Gemini requests receive provider timeout through config http_options."""
client = Mock()
client.aio.models.generate_content = AsyncMock(
return_value=SimpleNamespace(
candidates=[
SimpleNamespace(
finish_reason=SimpleNamespace(name="STOP"),
content=SimpleNamespace(parts=[SimpleNamespace(text="ok")]),
)
],
usage_metadata=SimpleNamespace(
prompt_token_count=12,
candidates_token_count=6,
),
parsed=None,
)
)
backend = GeminiBackend(client)
await backend.complete(
model="gemini-2.5-flash",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
extra_params={"timeout": "90"},
)
await_args = client.aio.models.generate_content.await_args
if await_args is None:
raise AssertionError("Expected Gemini generate_content call")
assert await_args.kwargs["config"]["http_options"].timeout == 90_000
@pytest.mark.asyncio
async def test_gemini_backend_rejects_budget_and_effort_together() -> None:
backend = GeminiBackend(Mock())
@ -385,7 +419,7 @@ async def test_gemini_backend_forwards_provider_params_extra_headers() -> None:
if await_args is None:
raise AssertionError("Expected Gemini generate_content call")
call = await_args.kwargs
assert call["config"]["http_options"]["headers"] == {"X-Trace-Id": "abc123"}
assert call["config"]["http_options"].headers == {"X-Trace-Id": "abc123"}
@pytest.mark.asyncio

View File

@ -510,6 +510,7 @@ async def test_openai_backend_converts_anthropic_style_tools() -> None:
assert call["tool_choice"] == "required"
@pytest.mark.asyncio
async def test_openai_backend_translates_canonical_any_tool_choice_to_required() -> (
None
):
@ -571,6 +572,78 @@ def test_openai_convert_tool_choice(canonical: Any, expected: Any) -> None:
assert OpenAIBackend._convert_tool_choice(canonical) == expected # pyright: ignore[reportPrivateUsage]
@pytest.mark.asyncio
async def test_openai_backend_passes_timeout_to_completion_request() -> None:
"""OpenAI completion requests receive per-request provider timeout."""
client = Mock()
client.chat.completions.create = AsyncMock(
return_value=SimpleNamespace(
choices=[
SimpleNamespace(
finish_reason="stop",
message=SimpleNamespace(
content="ok",
tool_calls=[],
reasoning_details=[],
),
)
],
usage=SimpleNamespace(
prompt_tokens=10,
completion_tokens=5,
prompt_tokens_details=None,
),
)
)
backend = OpenAIBackend(client)
await backend.complete(
model="gpt-4.1",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
extra_params={"timeout": 12.5},
)
assert _await_kwargs(client.chat.completions.create)["timeout"] == 12.5
@pytest.mark.asyncio
async def test_openai_backend_passes_timeout_to_structured_parse_request() -> None:
"""OpenAI structured parse requests receive per-request provider timeout."""
client = Mock()
client.chat.completions.parse = AsyncMock(
return_value=SimpleNamespace(
choices=[
SimpleNamespace(
finish_reason="stop",
message=SimpleNamespace(
parsed=_StructuredResponse(answer="ok"),
content='{"answer":"ok"}',
tool_calls=[],
refusal=None,
),
)
],
usage=SimpleNamespace(
prompt_tokens=10,
completion_tokens=5,
prompt_tokens_details=None,
),
)
)
backend = OpenAIBackend(client)
await backend.complete(
model="gpt-4.1",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
response_format=_StructuredResponse,
extra_params={"timeout": "30"},
)
assert _await_kwargs(client.chat.completions.parse)["timeout"] == 30.0
@pytest.mark.parametrize(
"model",
[

View File

@ -1,6 +1,8 @@
import pytest
from pydantic import BaseModel
from src.config import ModelConfig
from src.exceptions import ValidationException
from src.llm.caching import PromptCachePolicy
from src.llm.request_builder import execute_completion
from tests.llm.conftest import FakeBackend
@ -95,3 +97,51 @@ async def test_provider_params_are_merged_into_extra_params(
call = fake_backend.calls[0]
assert call["extra_params"]["top_p"] == 0.9
assert call["extra_params"]["custom_flag"] is True
async def test_provider_timeout_is_normalized_into_extra_params(
fake_backend: FakeBackend,
) -> None:
"""Numeric-string provider timeout values are normalized before backends."""
config = ModelConfig(
model="gpt-4.1-mini",
transport="openai",
provider_params={"timeout": "42.5"},
)
await execute_completion(
fake_backend,
config,
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
)
call = fake_backend.calls[0]
assert call["extra_params"]["timeout"] == 42.5
@pytest.mark.parametrize(
"timeout",
["slow", "", 0, -1, True, float("nan"), float("inf"), "nan", "inf"],
)
async def test_provider_timeout_rejects_invalid_values(
fake_backend: FakeBackend,
timeout: object,
) -> None:
"""Invalid provider timeout values fail before provider SDK calls."""
config = ModelConfig(
model="gpt-4.1-mini",
transport="openai",
provider_params={"timeout": timeout},
)
with pytest.raises(
ValidationException,
match=r"provider_params\.timeout must be a positive number of seconds",
):
await execute_completion(
fake_backend,
config,
messages=[{"role": "user", "content": "Hello"}],
max_tokens=100,
)

View File

@ -81,3 +81,50 @@ def test_representation_batch_target_input_cannot_exceed_max_input_tokens() -> N
MAX_INPUT_TOKENS=1000,
REPRESENTATION_BATCH_TARGET_INPUT_TOKENS=2048,
)
def _configured_with_timeout(timeout: object) -> ConfiguredModelSettings:
return ConfiguredModelSettings.model_validate(
{
"model": "gpt-5.4-mini",
"transport": "openai",
"overrides": {"provider_params": {"timeout": timeout}},
}
)
@pytest.mark.parametrize("timeout", [30, 42.5, "42.5", " 60 "])
def test_provider_timeout_is_normalized_at_config_load(timeout: object) -> None:
settings = _configured_with_timeout(timeout)
normalized = settings.overrides.provider_params["timeout"]
assert isinstance(normalized, float)
assert normalized == float(str(timeout).strip())
@pytest.mark.parametrize(
"timeout",
["slow", "", 0, -1, True, float("nan"), float("inf"), "nan", "inf", None, [30]],
)
def test_provider_timeout_is_rejected_at_config_load(timeout: object) -> None:
with pytest.raises(
ValueError, match=r"provider_params\.timeout must be a positive number"
):
_configured_with_timeout(timeout)
def test_provider_timeout_on_fallback_overrides_is_validated_at_config_load() -> None:
with pytest.raises(
ValueError, match=r"provider_params\.timeout must be a positive number"
):
ConfiguredModelSettings.model_validate(
{
"model": "gpt-5.4-mini",
"transport": "openai",
"fallback": {
"model": "gpt-4.1",
"transport": "openai",
"overrides": {"provider_params": {"timeout": "slow"}},
},
}
)