Hermes aec2ea8143 feat: add llama-cpp-hermes service with ROCm 6.1 + gfx906 support
- Add custom llama.cpp Dockerfile with ROCm 6.1 + gfx906 (MI50) build
- Add llama-cpp-hermes service serving Hermes 4.3 on dual MI50 GPUs
- Strip GPU devices/ROCm env from ollama service (CPU-only for embeddings)

Hermes 4.3 runs at ~19 t/s on dual MI50s with 160K context.
2026-07-05 21:34:54 -04:00
2025-08-08 15:08:10 -04:00
Description
No description provided
1,019 KiB
Languages
Dockerfile 91.5%
Makefile 8.5%