skill_view re-sent full skill content on every call: ~286k tokens of
verbatim repeat views in a 400k-msg production window (one session
loaded the same skill 9 times), and a single repeat view of a large
skill costs ~25k tokens.
Mirrors read_file's proven unchanged-stub pattern: a per-task cache
keyed on (resolved name, file_path) with an mtime+size fingerprint of
the served file. On a repeat view of an UNCHANGED file, return a short
stub pointing at the earlier result. This does NOT violate the
skills-are-loaded-fully rule — the stub only ever replaces content
that is already fully present earlier in the same conversation, and:
- any on-disk change (patch, external edit) invalidates the entry;
- context compression clears the cache (wired next to
reset_file_dedup in conversation_compression.py) so post-compression
re-views return full content;
- setup-needed views are never deduped (readiness can change without
the file changing);
- no task_id -> no dedup; caches are task-isolated; 200-entry cap.
Live E2E: repeat view of hermes-agent-dev 99,739 chars -> 374-char
stub.