* fix(deriver): truncate oversize observations so one cannot drop the batch simple_batch_embed raised ValueError when any input exceeded the per-input token cap, which failed the entire deriver save when a single observation was over-length. Add on_oversize="truncate": oversize inputs are embedded from a token-capped prefix (re-encoded until it fits, with a warning), preserving one vector per input. Default stays "raise" so existing callers are unchanged. RepresentationManager opts into truncate. Also add a live embedding test that fails on main (raise / missing kwarg) and passes once a mixed short+oversize batch survives. Refs #569 * fix(deriver): surface failure when all observer saves fail When every observer's save_representation failed (e.g. embedding retries exhausted under a sustained 429), the deriver logged the error and returned normally, so the queue marked the work unit processed with zero documents saved. Collect per-observer errors and, after telemetry is emitted, raise RepresentationSaveError when no observer succeeded. Partial failures stay processed (saved observers must not be discarded) and are recorded via an additive failed_observer_count on RepresentationCompletedEvent. Refs #728 * fix(embedding): guarantee truncation progress and truncate on re-embed The retry slice in _truncate_to_token_limit always recomputed the same keep count, so a slice whose re-encode grew past the cap could oscillate. Decrement keep after each unsuccessful retry. Document re-embed in the reconciler used the default on_oversize="raise", so one oversize document failed every other document in the batch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: drop ticket ids and shrink comments to one sentence Comments and docstrings describe current behavior, not the PR that introduced them. Ticket numbers stay in the commit/PR. * chore: annotate RepresentationSaveError and assert truncate on re-embed * fix(embedding): truncate on conclusion create paths and document BPE loop Storage callers in create_observations (API + agent tools) now pass on_oversize="truncate" so a single oversize item cannot drop the batch. Docstring on _truncate_to_token_limit notes why decode/re-encode is load-bearing. --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| README.md | ||
| __init__.py | ||
| conftest.py | ||
| embedding_matrix.py | ||
| model_matrix.py | ||
| test_live_anthropic.py | ||
| test_live_embeddings.py | ||
| test_live_gemini.py | ||
| test_live_openai.py | ||
| test_live_structured_output_unions.py | ||
| test_live_timeouts.py | ||
| test_live_tools_structured_output.py | ||
README.md
Live LLM Tests
These tests call real provider APIs and are disabled by default.
Run them with:
uv run pytest tests/live_llm -n 0 --live-llm --no-header -q
Required API key env vars:
LLM_ANTHROPIC_API_KEYLLM_OPENAI_API_KEYLLM_GEMINI_API_KEY
Model-family env vars:
LIVE_LLM_ANTHROPIC_45_PLUS_MODELSLIVE_LLM_OPENAI_GPT4_MODELSLIVE_LLM_OPENAI_GPT5_MODELSLIVE_LLM_OPENAI_OPENROUTER_NON_REASONING_MODELS(OpenAI-transport → OpenRouter-served non-reasoning models)LIVE_LLM_GEMINI_25_MODELSLIVE_LLM_GEMINI_30_MODELSLIVE_LLM_GEMINI_31_MODELS
Embedding-model env vars:
LIVE_EMBEDDING_GEMINI_MODELS(default:gemini-embedding-001,gemini-embedding-2; addgemini-embedding-2-previewto cover the preview twin)LIVE_EMBEDDING_OPENAI_MODELS(default:text-embedding-3-small)LIVE_EMBEDDING_OPENAI_COMPATIBLE_MODELS(no default → skipped) — OpenAI transport pointed at a third-party OpenAI-compatible provider. Also readsOPENROUTER_API_KEY,LIVE_EMBEDDING_OPENAI_COMPATIBLE_BASE_URL(defaulthttps://openrouter.ai/api/v1),LIVE_EMBEDDING_OPENAI_COMPATIBLE_DIMENSIONS(default3072) andLIVE_EMBEDDING_OPENAI_COMPATIBLE_SEND_DIMENSIONS(default on; set to0for a provider that rejects OpenAI'sdimensionsparam)
export OPENROUTER_API_KEY="sk-or-v1-..."
export LIVE_EMBEDDING_OPENAI_COMPATIBLE_MODELS="google/gemini-embedding-001"
Each model env var accepts a comma-separated list of bare model ids or provider-qualified ids.
Examples:
export LIVE_LLM_ANTHROPIC_45_PLUS_MODELS="claude-sonnet-4-5,claude-sonnet-4-6"
export LIVE_LLM_OPENAI_GPT4_MODELS="gpt-4.1"
export LIVE_LLM_OPENAI_GPT5_MODELS="gpt-5,gpt-5.4,gpt-5.4-mini"
export LIVE_LLM_OPENAI_OPENROUTER_NON_REASONING_MODELS="inception/mercury-2"
export LIVE_LLM_GEMINI_25_MODELS="gemini-2.5-flash,gemini-2.5-pro"
export LIVE_LLM_GEMINI_30_MODELS="gemini-3-flash-preview"
export LIVE_LLM_GEMINI_31_MODELS="gemini-3.1-pro-preview"
OpenRouter-routed models require additional env for the proxy endpoint:
export OPENROUTER_API_KEY="sk-or-v1-..."
# Per-feature config example:
# DERIVER_MODEL_CONFIG__TRANSPORT=openai
# DERIVER_MODEL_CONFIG__MODEL=inception/mercury-2
# DERIVER_MODEL_CONFIG__OVERRIDES__BASE_URL=https://openrouter.ai/api/v1
# DERIVER_MODEL_CONFIG__OVERRIDES__API_KEY_ENV=OPENROUTER_API_KEY
Coverage by provider:
- Anthropic: structured output path, prompt caching metrics, thinking blocks, multi-turn tool replay
- OpenAI GPT-4 class: structured outputs, prompt caching
- OpenAI GPT-5 class (incl. gpt-5.x point-releases): structured outputs, prompt caching,
reasoning_effort,max_completion_tokensrouting - OpenAI transport → OpenRouter non-reasoning models (e.g.
inception/mercury-2): non-chat / diffusion architectures must stay onmax_tokens, noreasoning_effort, tool-calling parameter-schema compatibility is the canary for exotic OR-served providers - Gemini 2.5/3.0 classes: structured outputs, cached-content reuse, thought signatures, multi-turn tool replay
- Gemini 3.1 class: thinking and tool replay coverage by default; structured-output/caching coverage should only be added once Google documents support for that path
- Embeddings (
test_live_embeddings.py): single embed, batched embed, batch-vs-single alignment, chunk-to-id mapping, and oversize-truncate survival (on_oversize="truncate") for every configured embedding model.gemini-embedding-2*is the reason this exists — those models collapse a list of bare strings into one document (#745), and only a live call catches it. Also covers first-classEmbeddingModelConfig.timeoutplumbing (one representative model per transport): configured timeout lands on the SDK client, and a near-zero timeout aborts before the provider answers - OpenAI-compatible embedding providers (e.g. OpenRouter's
google/gemini-embedding-001): the #932 surface. Those providers reject a base64 embedding request outright (HTTP 400) or answer HTTP 200 with empty data, so the whole matrix fails withoutencoding_format="float". Real OpenAI accepts base64 happily, so only a third-party provider catches it. Note that OpenRouter load-balances across upstreams, so the base64 failure is per-attempt rather than guaranteed: a retry can land on an endpoint that accepts it.test_live_openai_float_encoding_matches_base64covers the other side, that the float switch must not move vectors on real OpenAI