hermes-agent/tests/agent
brooklyn! 3e74f75e41
feat(agent): coding-context posture across CLI/TUI/desktop/ACP (#43316)
* feat(agent): coding-context posture with per-model edit-format tuning

Hermes detects when it's running in a coding context — an interactive
surface (CLI, TUI, ACP, desktop) sitting in a code workspace (git repo or
recognised project root) — and shifts into a coding posture. Outside that
(chat platforms, non-workspaces) nothing changes.

The posture is modelled as a frozen RuntimeMode selected from a small
ContextProfile registry (coding/general). A profile is data: the toolset to
collapse to, the operating brief to inject, and seams for model routing and
memory. Every domain reads the same resolved object instead of re-probing
git/config on its own:

- System prompt — RuntimeMode.system_blocks(): an operating brief (gather
  context before editing, edit through tools not chat, verify with terminal,
  cap retry loops) plus a live git/workspace snapshot, built once and baked
  into the stable prompt tier so per-conversation caching is preserved.
- Per-model edit-format tuning — the brief nudges each model family toward
  the patch mode it handles best: OpenAI/Codex toward mode='patch' (V4A
  multi-file diffs), Anthropic toward mode='replace' (string replacement).
  The model id rides on RuntimeMode; unknown families keep neutral wording.
- Skill index — non-coding skill categories are pruned from the prompt's
  skill index (discovery-only; skills_list/skill_view still reach the full
  catalog, with a disclosure note).
- Toolset — only under the opt-in 'focus' mode does the posture collapse to
  the coding toolset + enabled MCP servers; the default posture is
  prompt-only and never overrides configured toolsets.

Activation via agent.coding_context: auto (default), focus, on, off.
Subagents inherit the posture for free via toolset inheritance + the shared
prompt builder. Detection is not memoized so a long-lived gateway/TUI
process can't pin a stale posture across working directories.

* feat(agent): cover new-file authoring in the coding edit-format nudge

The per-model edit-format guidance only addressed editing existing code
(patch mode='patch' vs 'replace'), but authoring a brand-new file —
write_file, not patch — is a large fraction of real coding work and the
nudge was silent on it. Surfaced when building a single-file artifact where
the dominant operation was write_file and the steering offered no guidance.

Both family lines now lead with "author new files with write_file; for
edits to existing code prefer ...". Tests assert write_file appears in each
family's brief; unknown families still get neutral wording.

* docs(agent): correct memoization docstring + clarify TUI config-load asymmetry

* feat(agent): sharpen the coding posture — verify-loop facts, wider edit steering, $HOME guard

Tuning pass on the coding posture from dogfooding it as a harness:

- Workspace snapshot now hands the model its verify loop up front:
  detected manifests + package manager (lockfile sniff), the exact
  verify commands (package.json scripts, Makefile targets,
  scripts/run_tests.sh, pytest config), and which context files
  (AGENTS.md / CLAUDE.md / .cursorrules) exist at the root. Marker-only
  (non-git) projects get the snapshot too instead of nothing. The
  "verify before claiming done" brief line was the highest-value piece
  in evals — this turns it from advice into an executable loop instead
  of making the model rediscover the test command every session. Still
  stat-cheap, size-guarded reads, built once at prompt time.

- Edit-format steering covers the families Hermes actually serves:
  Gemini and open-weight coding models (DeepSeek, Qwen, Kimi, GLM,
  Grok, Hermes, Llama, Mistral, Devstral, MiniMax) steer to
  mode='replace' — their RL scaffolds use str_replace-style editors.
  Previously only GPT/Codex and Claude families got steering; the
  models Hermes users disproportionately run all fell to neutral.

- Operating brief gains four behaviors elite harnesses encode: batch
  independent reads/searches in one turn; fix root causes and the bug
  class (sibling call paths), not the reported site; no drive-by
  refactors/renames/reformatting; never read, print, or commit secrets.
  Plus a patch-failure escalation ladder: after the same region fails
  twice, rewrite the enclosing function/file with write_file instead of
  a third patch attempt.

- $HOME dotfiles guard: a git repo rooted exactly at the home directory
  (or a marker sitting in it, e.g. a global ~/AGENTS.md) is user config,
  not a code workspace — without the guard, every session anywhere under
  a dotfiles-managed home silently flipped to the coding posture. Real
  projects under such a home still detect via their own markers/repos;
  'on' mode bypasses the guard.
2026-06-10 23:06:44 -05:00
..
lsp
…
transports
…
__init__.py
…
test_anthropic_adapter.py
…
test_anthropic_keychain.py
…
test_anthropic_kwargs_sanitize.py
…
test_anthropic_mcp_prefix_strip.py
…
test_anthropic_oauth_pkce.py
…
test_anthropic_output_field_leak.py fix(anthropic): strip output-only SDK fields from replayed content blocks 2026-06-10 20:45:16 -07:00
test_anthropic_thinking_block_order.py refactor: keep anthropic_content_blocks in-memory only (no state.db column) 2026-06-10 20:45:16 -07:00
test_arcee_trinity_overrides.py
…
test_async_utils.py
…
test_auxiliary_client.py
…
test_auxiliary_client_anthropic_custom.py
…
test_auxiliary_client_azure_foundry.py
…
test_auxiliary_client_xai_oauth_recovery.py
…
test_auxiliary_config_bridge.py
…
test_auxiliary_main_first.py
…
test_auxiliary_named_custom_providers.py
…
test_auxiliary_transport_autodetect.py
…
test_auxiliary_user_default_headers.py
…
test_azure_identity_adapter.py
…
test_bedrock_1m_context.py
…
test_bedrock_adapter.py
…
test_bedrock_integration.py
…
test_cascading_interrupt_6600.py
…
test_codex_cloudflare_headers.py
…
test_codex_responses_adapter.py
…
test_codex_ttfb_watchdog.py
…
test_coding_context.py feat(agent): coding-context posture across CLI/TUI/desktop/ACP (#43316) 2026-06-10 23:06:44 -05:00
test_compress_focus.py
…
test_compression_concurrent_fork.py
…
test_compression_logging_session_context.py
…
test_compressor_historical_media.py
…
test_compressor_image_tokens.py
…
test_context_compressor.py
…
test_context_compressor_cross_session_guard.py
…
test_context_compressor_summary_continuity.py
…
test_context_compressor_temporal_anchoring.py
…
test_context_engine.py
…
test_context_engine_host_contract.py
…
test_context_references.py
…
test_copilot_acp_client.py
…
test_copilot_acp_deprecation.py
…
test_credential_pool.py
…
test_credential_pool_routing.py
…
test_credits_cold_start.py
…
test_credits_fixture_snapshot.py
…
test_credits_policy.py
…
test_credits_tracker.py
…
test_crossloop_client_cache.py
…
test_curator.py
…
test_curator_activity.py
…
test_curator_backup.py
…
test_curator_classification.py
…
test_curator_reports.py
…
test_custom_provider_extra_body.py
…
test_custom_providers_vision.py
…
test_deepseek_anthropic_thinking.py
…
test_direct_provider_url_detection.py
…
test_display.py feat(web): Parallel-backed web search & extract — free Search MCP when keyless, v1 REST when keyed 2026-06-10 19:54:38 -07:00
test_display_emoji.py
…
test_display_todo_progress.py
…
test_display_tool_failure.py
…
test_error_classifier.py
…
test_external_skills.py
…
test_external_skills_dirs_cache.py
…
test_file_safety.py
…
test_file_safety_container_mirror.py
…
test_file_safety_credentials.py
…
test_file_safety_cross_profile.py
…
test_file_safety_sandbox_mirror.py
…
test_gemini_cloudcode.py
…
test_gemini_fast_fallback.py
…
test_gemini_free_tier_gate.py
…
test_gemini_native_adapter.py
…
test_gemini_schema.py
…
test_i18n.py
…
test_image_gen_registry.py
…
test_image_routing.py
…
test_insights.py
…
test_jiter_preload.py
…
test_kimi_coding_anthropic_thinking.py
…
test_last_total_tokens.py
…
test_local_stream_timeout.py
…
test_markdown_tables.py
…
test_memory_async_sync.py
…
test_memory_provider.py
…
test_memory_session_switch.py
…
test_memory_user_id.py
…
test_minimax_auxiliary_url.py
…
test_minimax_provider.py
…
test_model_metadata.py
…
test_model_metadata_local_ctx.py
…
test_model_metadata_ssl.py
…
test_models_dev.py
…
test_moonshot_schema.py
…
test_non_stream_stale_timeout.py
…
test_nous_credits_gauge.py
…
test_nous_credits_snapshot.py
…
test_nous_oauth_401_guidance.py
…
test_nous_rate_guard.py
…
test_onboarding.py
…
test_openrouter_response_cache.py
…
test_plugin_llm.py
…
test_portal_tags.py
…
test_prompt_builder.py feat(agent): coding-context posture across CLI/TUI/desktop/ACP (#43316) 2026-06-10 23:06:44 -05:00
test_prompt_caching.py
…
test_proxy_and_url_validation.py
…
test_rate_limit_tracker.py
…
test_redact.py
…
test_resume_stale_active_task.py
…
test_runtime_cwd.py
…
test_save_url_image.py
…
test_set_runtime_main_custom_provider.py
…
test_shell_hooks.py
…
test_shell_hooks_consent.py
…
test_skill_bundles.py
…
test_skill_commands.py
…
test_skill_commands_reload.py
…
test_skill_utils.py
…
test_stream_read_timeout_floor.py fix(streaming): stop socket read timeout from preempting stale-stream detector (#43570) 2026-06-10 20:21:38 -05:00
test_streaming_context_scrubber.py
…
test_subagent_progress.py
…
test_subagent_stop_hook.py
…
test_subdirectory_hints.py
…
test_summary_prefix_semantics.py
…
test_system_prompt.py feat(agent): coding-context posture across CLI/TUI/desktop/ACP (#43316) 2026-06-10 23:06:44 -05:00
test_system_prompt_restore.py
…
test_think_scrubber.py
…
test_title_generator.py
…
test_tool_dispatch_helpers.py
…
test_tool_guardrails.py
…
test_tool_result_classification.py
…
test_transcription_registry.py
…
test_tts_registry.py
…
test_turn_context.py
…
test_turn_retry_state.py
…
test_unsupported_parameter_retry.py
…
test_unsupported_temperature_retry.py
…
test_usage_pricing.py
…
test_video_gen_registry.py
…
test_vision_resolved_args.py
…
test_vision_routing_31179.py
…