The GpuPassthroughRemediationProvider probed GPU passthrough by exec'ing
`nvidia-smi` inside nomad_ollama. When Ollama was created CPU-only — most
commonly because the NVIDIA container runtime was registered with Docker
only AFTER the AI Assistant was first installed (DockerService attaches a
GPU DeviceRequest only when 'nvidia' is in docker.info().Runtimes at
install time) — the container ships no nvidia-smi, so the exec returned an
error string whose letters the alphabetic-output check mistook for
"passthrough healthy". That false negative suppressed all auto-remediation,
leaving Ollama silently stuck on CPU (models run at a fraction of a token
per second) until the user manually clicked "Fix: Reinstall AI Assistant".
Replace the nvidia-smi probe with Ollama's own "inference compute" startup
log line — the ground truth for whether a GPU backend actually loaded — via
a new, unit-tested classifyOllamaComputeBackend() helper. One signal covers
both the "created CPU-only" case and the original "DeviceRequests present
but the toolkit binding tore" case. Add a cooldown on gpu.autoRemediatedAt
so a GPU that genuinely cannot be accelerated (an architecture Ollama has
no kernels for) does not trigger a reinstall on every admin boot.
Co-Authored-By: Claude <noreply@anthropic.com>