The declared floor, the local pin, the lockfile, the image, and CI had all
drifted apart: pyproject said >=3.10, .python-version said 3.11, uv.lock
resolved >=3.11, the Dockerfile ships 3.13, and unified-tests hardcoded 3.12.
Set the server's floor to 3.13 to match the only interpreter that actually
ships, and pin CI to the same version. Three workflows passed
python-version-file: pyproject.toml, which makes setup-python read
requires-python as a range and install the newest match -- run 33108027458
resolved to CPython 3.14.7, so CI tested a Python nobody deploys and the 3.10
floor was never exercised. Pointing them at .python-version installs an exact
version instead.
That promotion makes .python-version load-bearing for CI, so add it to the
path filters in unittest.yml and live-llm-tests.yml. Those keyed on
pyproject.toml for the interpreter before; without this, editing the pin alone
would change which Python CI runs on without retriggering the suites that run
on it. unified-tests.yml is left alone -- it filters on src/** and tests/**
and never keyed on pyproject.toml either.
Re-resolving the lockfile drops async-timeout, tomli, and overrides, which
were backport shims only needed below 3.13.
sdks/python (honcho-ai, >=3.8) and honcho-cli (>=3.11) are left alone; they
publish to PyPI, so those floors are consumer-facing.
Ruff's inferred target-version moves py310 -> py313, which surfaces ~200 new
findings tree-wide (mostly UP017, datetime.timezone.utc -> datetime.UTC).
Ruff runs only via pre-commit, not in any workflow, so CI is unaffected; the
sweep is left for a follow-up rather than adding unrelated churn here.
Fixes DEV-2415
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gate job only runs on `pull_request: labeled`, so it is skipped on
push. A skipped ancestor propagates down the needs chain unless a job
opts out, which `unified-tests` never did — so the suite has been
skipped on every merge to main while still burning a Fly machine.
* fix(llm): support combined tool calling and structured output across backends
- OpenAI: parse() 500s on non-strict function tools; route tool-carrying
structured requests through create() with an explicit json_schema
response_format (mirrors the streaming path)
- Anthropic: skip the '{' JSON prefill when tools are present so tool_use
blocks stay reachable; make the schema instruction conditional and rely
on parse + repair
- Gemini: native response_schema + function calling is rejected before
Gemini 3; with tools present, inject a schema instruction into the final
turn instead and rely on parse + repair
- All backends: tool-call turns carry no consumable content, so skip
structured-output parsing on them
Extracted from the dialectic structured-output branch (DEV-1652) so the
transport layer can land independently.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(live_llm): exercise combined tools + structured output per provider
Two-turn live flow per backend: a forced tool-call turn (structured
parsing must be skipped) followed by a replay turn that must return a
schema-conforming answer with tools still attached. Asserts the
provider-specific request shaping: no parse() for OpenAI (500s on
non-strict tools), no '{' prefill for Anthropic, no native
response_schema for Gemini.
Verified against live OpenAI (gpt-4.1, gpt-5, gpt-5.4, gpt-5.4-mini)
and Gemini (gemini-2.5-flash).
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: some needed unrelated test failures
* ci: add label-triggered live LLM test workflow
Adding the run-live-llm label to a PR (or workflow_dispatch) runs
tests/live_llm/ against real provider APIs — the only place the
--live-llm suite runs in CI. Reuses the unified-tests environment and
its Secrets Manager staging-dotenv resolution for provider keys; runs
on ubuntu-latest (no Fly runner, no Docker — the suite only touches the
LLM backends). Pins LIVE_LLM_ANTHROPIC_45_PLUS_MODELS=claude-sonnet-4-5
since the Anthropic family has no default models and would otherwise
silently collect empty.
Opt-in by design: live model behavior is variable, so this is a signal,
not a required check.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: run live LLM tests on main pushes touching the transport
Mirrors unified-tests' push trigger, scoped to paths that can affect
the live suite (src/llm/, config, the tests, deps, and the workflow
itself) so provider API calls aren't spent on unrelated changes.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: disable auth in live LLM test environment
The staging dotenv sets AUTH_USE_AUTH=true without a usable JWT secret,
and src/config.py validates the pair at import time — the same reason
unified-tests overrides it. This suite never runs the API server.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(live_llm): fix gpt-5.4 reasoning_effort and gemini replay-turn flake
- test_live_openai: gpt-5.4 dropped 'minimal' from the reasoning_effort
vocabulary, so the gpt5 caching test 400'd — and the OpenAI backend's
BadRequestError terminal swallowed it into an empty CompletionResult.
Pick the effort per model generation.
- test_live_tools_structured_output: use tool_choice='auto' on the
replay turn, matching the production dialectic loop (which never
forces 'none') — NONE mode is what provoked gemini-2.5-flash's empty
candidates. Drop the temperature pin so retries actually resample,
and treat a repeat tool call as a retryable attempt.
Verified live: full suite green, gemini 4/4 consecutive passes.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: fail live LLM run when no staging secret was loaded
If the latest-tag fetch fails and no second tag exists, the fallback
step is skipped rather than failed, and the job would proceed without
provider keys — every test then skips via require_provider_key and the
run goes green. Guard on both fetch outcomes so that path fails loudly.
DEV-2035
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(live-llm-tests-GHA): remove extra comments
* ci(CODEOWNERS): introduce CODEOWNERS and gate GHA heavy test runs behind being a CODEOWNER
* ci(GHA-live-LLM-tests): consolidate common GHA steps
* test(test_live_openai): fix reasoning level adjustment for gpt-5
* test(live-llm-tests): temporary removal of gate to test the workflow
* test(live-llm-tests): revert removal of gate
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* chore: use git tags to fetch secrets for unified test
* fix: test failure
* fix: override AUTH_USE_AUTH and SENTRY_ENABLED
* fix: upload traces
* chore: rm run on PR
* fix: rm bucket from logs
* feat: run unified tests in CI
* fix: attempt use aws secrets manager
* fix: temp add verification workflow
* fix: CodeRabbit comments
* fix: remove debugging step
* fix: only run on main
* fix: add UNIFIED_TEST_LOG_LEVEL env var; default to WARNING