hermes-agent/tests/run_agent
Ben Barclay f52feed1ef
fix(azure-foundry): scope Responses reasoning suppression to post-tool turns (#84320)
Azure Foundry's OpenAI-compatible Responses surface rejects the post-tool
follow-up payload with HTTP 400 `invalid_payload` when a replayed encrypted
`reasoning` item is sent alongside `function_call` / `function_call_output`.
The initial function-call request and ordinary multi-turn continuity are both
accepted, so the failure only appears after the first tool executes.

Detect the Foundry endpoint in `ResponsesApiTransport.build_kwargs` and drop
only the encrypted reasoning replay on that follow-up turn, leaving
function_call / function_call_output continuity intact.

Salvage of #59981, rebuilt on current main. Same root cause and fix direction
as the original, which was correct; this version resolves three defects:

- No `chat_completion_helpers.py` change. main already forwards `provider`
  and `base_url` to the Responses transport, so the original's re-added
  arguments produced `SyntaxError: keyword argument repeated: provider` on
  merge. Dropping the hunk removed the syntax error and the conflict.

- Host matching uses `utils.base_url_host_matches`, not a substring test.
  `".services.ai.azure.com" in base_url` also matches URLs carrying the
  domain in a path or query segment, which would silently disable reasoning
  replay on an unrelated provider.

- The post-tool predicate tests the trailing messages, not the whole history.
  Scanning for any tool call plus any tool result made it sticky: one tool
  call early in a conversation suppressed reasoning on every later turn.

- Tool calls pair on `call_id` as well as `id`. Responses histories carry the
  function call id in `call_id` while `id` holds the response item id
  (`fc_...`). Identity is resolved via the converter's own
  `_split_responses_tool_id`, covering composite `"call_x|fc_y"` ids and bare
  `fc_` ids on both sides of the pairing.

Tests: 27 cases across the transport and the live `build_api_kwargs` bridge,
including six parametrized tool-call id shapes, non-Foundry host lookalikes,
the sticky-history guard, parallel tool results, and an unpaired tool result.
Each guard was confirmed to catch its defect by reverting the fix.

Verified with `scripts/run_tests.sh tests/agent/ tests/run_agent/`:
532 files, 5602 tests passed, 0 failed.

Not verified against a live Azure Foundry endpoint — no credentials. The
original HTTP 400 reproduction and post-fix Foundry Project / Azure Container
Apps harness runs are @AshuJoshi's, from #59981. This change is verified at
the payload-construction layer only.

Closes #59981.

Co-authored-by: Ashu Joshi <AshuJoshi@users.noreply.github.com>
2026-08-14 12:41:26 +10:00
..
__init__.py
conftest.py
repro_48013_image_shrink_brick.py
test_413_compression.py
test_860_dedup.py
test_1630_context_overflow_loop.py
test_18028_content_policy_blocked.py
test_24996_fallback_exhaustion_cooldown.py
test_28161_anthropic_stream_pool_cleanup.py
test_31273_402_not_retried.py
test_32646_fallback_429_after_timeout.py
test_63425_credential_pool_auto_detect.py
test_66267_multimodal_interim.py
test_70773_shared_client_fd_corruption.py
test_81641_text_turn_incremental_persistence.py
test_agent_guardrails.py
test_anthropic_mid_tool_call_drop.py
test_anthropic_prompt_cache_policy.py fix: address review feedback from #85512 2026-08-14 00:59:38 +05:30
test_anthropic_response_header_capture.py
test_anthropic_third_party_oauth_guard.py
test_anthropic_truncation_continuation.py
test_api_max_retries_config.py
test_async_httpx_del_neuter.py
test_auth_provider_failover.py
test_authorization_gate.py
test_background_review.py
test_background_review_cache_parity.py
test_background_review_cost_controls.py
test_background_review_summary.py
test_background_review_toolset_restriction.py
test_callable_api_key.py
test_codex_app_server_compaction.py
test_codex_app_server_integration.py
test_codex_app_server_lifecycle.py
test_codex_multimodal_tool_result.py
test_codex_no_tools_nonetype.py
test_codex_silent_hang_hint.py
test_codex_xai_oauth_recovery.py
test_commit_memory_session_context_engine.py
test_compress_focus_plugin_fallback.py
test_compression_abort_state_reset.py
test_compression_boundary.py
test_compression_boundary_hook.py
test_compression_feasibility.py
test_compression_lock_defer.py
test_compression_persistence.py
test_compression_trigger_excludes_reasoning.py
test_compressor_fallback_update.py
test_concurrent_interrupt.py
test_context_token_tracking.py
test_continuation_ceiling_wedge.py
test_conversation_fallback_state.py
test_copilot_native_vision_headers.py
test_create_openai_client_disables_sdk_retries.py
test_create_openai_client_kwargs_isolation.py
test_create_openai_client_proxy_env.py
test_create_openai_client_reuse.py
test_create_openai_client_ssl_verify.py
test_credential_pool_interrupt.py
test_credential_rotation_route_settings.py
test_credits_notices_toggle.py
test_custom_provider_extra_headers_client.py
test_deepseek_reasoning_content_echo.py
test_deepseek_v4_thinking_live.py
test_dict_tool_call_args.py
test_dropped_tool_call_recovery.py
test_empty_response_recovery_persistence.py
test_empty_terminal_reasoning_surface.py
test_env_credential_turn_refresh.py
test_exit_cleanup_interrupt.py
test_fallback_api_mode_preservation.py fix: explicit fallback api_mode always wins; clean up dead-code guard 2026-08-13 12:31:02 -07:00
test_fallback_credential_isolation.py
test_fallback_reasoning_override.py
test_file_mutation_verifier.py
test_fireworks_live.py
test_identity_flush.py
test_image_generate_parallel.py
test_image_rejection_fallback.py
test_image_shrink_recovery.py
test_in_place_compaction.py
test_infinite_compaction_loop.py
test_init_fallback_on_exhausted_pool.py
test_interactive_interrupt.py
test_interrupt_propagation.py
test_invalid_context_length_warning.py
test_iteration_budget_race.py
test_jsondecodeerror_retryable.py
test_last_reasoning_per_turn.py
test_lmstudio_load_mode.py
test_long_context_tier_429.py
test_malformed_tool_arguments.py
test_materialize_data_url_cleanup.py
test_memory_nudge_counter_hydration.py
test_memory_provider_init.py
test_memory_sync_interrupted.py
test_message_sequence_repair.py
test_moa_fanout_cadence.py
test_moa_loop_mode.py
test_moa_privacy_filter.py
test_moa_streaming.py
test_multimodal_tool_content_recovery.py
test_native_compaction.py fix(compression): harden native compaction rejection matcher + config coercion (#82777) 2026-08-13 03:04:31 -07:00
test_nonretryable_error_html_summary.py
test_notice_spine.py
test_nous_429_fallback_reentry.py
test_nous_fallback_unavailable.py
test_openai_client_lifecycle.py
test_overflow_overhead_aware_tokens.py
test_partial_stream_finish_reason.py
test_per_model_compression_threshold.py
test_per_model_threshold_init_ordering.py
test_percentage_clamp.py
test_plugin_context_engine_init.py
test_plugin_stream_hooks.py test(plugins): assert per-hook stream ordering, not cross-thread interleaving 2026-08-12 19:15:02 -07:00
test_post_tool_compression_attempt_cap.py
test_pre_compress_memory_context.py
test_preflight_compression_cap_e2e.py
test_primary_runtime_restore.py feat: server-side ui_meta on profiles.list/configure (#85440) 2026-08-13 10:07:39 -07:00
test_proactive_prune_loop_wiring.py
test_provider_attribution_headers.py fix: send Hermes Agent attribution headers to OpenCode Zen and Go 2026-08-13 02:03:40 -07:00
test_provider_fallback.py
test_provider_parity.py
test_repair_tool_call_arguments.py
test_repair_tool_call_name.py
test_request_client_reuse_abort_races.py
test_reset_aware_primary_restore.py
test_retry_status_buffer.py
test_review_prompt_class_first.py
test_run_agent.py test: register setup_mcp in the desktop_ui toolset + post-hook contracts 2026-08-13 01:06:51 -05:00
test_run_agent_codex_responses.py fix(azure-foundry): scope Responses reasoning suppression to post-tool turns (#84320) 2026-08-14 12:41:26 +10:00
test_run_agent_multimodal_prologue.py
test_sequential_chats_live.py
test_session_activity_persist.py
test_session_id_env.py
test_session_meta_filtering.py
test_session_reset_fix.py
test_session_source.py
test_start_order_gate.py
test_steer.py
test_stream_drop_logging.py
test_stream_interrupt_retry.py
test_stream_single_writer_65991.py
test_stream_stale_breaker_reset.py
test_stream_stale_circuit_breaker.py
test_streaming.py
test_streaming_tool_call_repair.py
test_strict_api_validation.py
test_strip_reasoning_tags_cli.py
test_summarize_api_error.py
test_switch_model_context.py
test_switch_model_fallback_prune.py
test_switch_model_pool_reload_52727.py
test_switch_model_reapplies_headers.py
test_switch_model_reasoning_override.py
test_switch_model_rollback.py
test_switch_model_stale_base_url.py
test_thinking_only_sanitizer.py fix(agent): hoist checkpoint carrier guard above the reasoning branches 2026-08-13 03:04:45 -07:00
test_thinking_prefill_trailing_turn.py
test_thinking_sig_recovery_persistence.py
test_tls_fd_recycle_corruption.py
test_token_persistence_non_cli.py
test_tool_arg_coercion.py
test_tool_batch_segmentation.py
test_tool_call_args_sanitizer.py
test_tool_call_guardrail_runtime.py
test_tool_call_incremental_persistence.py
test_tool_executor_contextvar_propagation.py
test_tool_name_db_persistence.py
test_turn_completion_explainer.py
test_unicode_ascii_codec.py
test_verification_continuation_budget.py
test_vision_aware_preprocessing.py
test_vision_tool_messages.py
test_wait_state_visibility.py