Ollama's scheduler drops integrated GPUs unless OLLAMA_IGPU_ENABLE=1 is set. NOMAD sets HSA_OVERRIDE_GFX_VERSION for AMD but never this flag, so AMD APUs (780M/890M/8060S) silently fell back to CPU-only inference despite correct /dev/kfd and /dev/dri passthrough. Set OLLAMA_IGPU_ENABLE=1 whenever AMD acceleration is configured, on both the install and update provisioning paths. The flag is a no-op on discrete AMD cards, so it's safe to set unconditionally within the AMD branch. On the update path we also strip any prior value so containers provisioned before this change pick up the flag on their next update. Verified on NOMAD2 (Ryzen AI 9 HX 370 / Radeon 890M): before, Ollama logged "dropping integrated GPU" and ran on CPU; after, it reports the 890M as an iGPU ROCm inference device and models load at 100% GPU. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| controllers | ||
| exceptions | ||
| jobs | ||
| middleware | ||
| models | ||
| services | ||
| utils | ||
| validators | ||