hermes-agent/scripts
kshitijk4poor 2ccfdb2db4 fix(agent): exempt parseable vLLM/LM Studio output-cap errors from compression-disabled guard
Salvage of #63862. is_output_cap_error() returns False for vLLM/LM Studio
error messages that contain 'prompt contains ... input tokens' (treated as
input-overflow signal). But parse_available_output_tokens_from_error() CAN
extract a valid available_tokens from those same messages. The
compression-disabled guard only checked is_output_cap_error(), so vLLM/LM
Studio users with compression off still got a terminal failure instead of
the max-tokens retry.

Fix: also exempt when parse_available_output_tokens_from_error() returns a
value — that function determines whether the retry path can actually handle
the error, so it's the right predicate for the exemption.

Added test: verify vLLM-format error with compression_disabled=False still
triggers the max-tokens retry path.

Co-authored-by: dmabry <dmabry@users.noreply.github.com>
2026-07-14 03:58:54 +05:30
..
ci fix(ci): add missing Δ Wait column to skipped job rows 2026-07-13 17:46:11 -04:00
lib
…
tests
…
whatsapp-bridge
…
LIVETEST_README.md
…
analyze_livetest.py
…
benchmark_browser_eval.py
…
build_model_catalog.py
…
build_skills_index.py
…
check-windows-footguns.py
…
check_subprocess_stdin.py
…
contributor_audit.py
…
dev-sandbox.sh
…
discord-voice-doctor.py
…
docker_config_migrate.py
…
docker_rebootstrap_nous_session.py
…
hermes-gateway
…
install.cmd
…
install.ps1
…
install.sh
…
install_psutil_android.py
…
keystroke_diagnostic.py
…
kill_modal.sh
…
lint_diff.py
…
profile-tui.py
…
release.py fix(agent): exempt parseable vLLM/LM Studio output-cap errors from compression-disabled guard 2026-07-14 03:58:54 +05:30
run_tests.sh feat(ci): python test speedups 2026-07-13 15:29:20 -04:00
run_tests_parallel.py feat(ci): python test speedups 2026-07-13 15:29:20 -04:00
sample_and_compress.py
…
tool_search_livetest.py
…