Skip to content

add Local LLM Hosting guide - #38

Merged
dougburks merged 1 commit into
3/devfrom
docs/local-llm-hosting
Sep 4, 2026
Merged

dougburks merged 1 commit into
3/devfrom
docs/local-llm-hosting

Conversation

@TOoSmOotH

Copy link
Copy Markdown
Contributor

What

Adds docs/local-llm.md, a tested walkthrough for standing up a local OpenAI-compatible endpoint for Onion AI on an AMD Strix Halo machine, plus a link to it from onion-ai.md and a nav entry under Security Onion Pro.

Why

The Hosting Local Models section of onion-ai.md offered only "use LM Studio" and "You need at least 96GB of VRAM" with no actual procedure. That VRAM figure is right for the large models listed above it but misleading as general guidance, so it is now scoped to those models and points at the new page.

Validation

This was built and run end to end on an AMD Strix Halo box (Ryzen AI Developer Platform, gfx1151, ROCm 7.13, 128GB unified memory). Every command and number in the page came from that machine — nothing is inferred.

Findings worth calling out, all of which are now documented:

  • The llama.cpp package ships its own llama-server systemd unit and /etc/default/llama-server config, so no custom unit file is needed.
  • The packaged _llama-server account cannot reach the GPU without being added to the render and video groups. Interactive shells work regardless, because logind grants a per-user ACL on /dev/kfd and /dev/dri/renderD128. The failure mode is nasty: the service starts, reports healthy, and silently falls back to unusably slow CPU inference. This gets a prominent warning.
  • Gemma 4 is a reasoning model. It populates reasoning_content and leaves content empty until it finishes thinking, so too low a token limit produces an empty response that reads as a broken endpoint.
  • Tool calling works and returns well-formed tool_calls, which matters because the assistant depends on it.
  • Q8_0 at the full 262,144 token window uses only ~32 GB, and a fact planted at the start of a 193,139 token prompt was retrieved correctly at the end — the window is usable, not merely allocatable.

Measured: ~41 tok/s generation on short context, ~22 tok/s at ~193k, bulk prefill ~850 tok/s falling to ~228 tok/s averaged across a 193k prompt.

vLLM

vLLM was evaluated first and rejected. gfx1151 is not in its supported GPU list, there is no ROCm wheel, and every working reference builds from source with patches (stub amdsmi, force the arch into CMakeLists.txt, disable AITER MoE) inside a container. That patch set rots as vLLM moves and isn't something worth publishing in customer-facing docs. The packaged llama.cpp path is vendor-supported and survives upgrades.

Note for reviewers

Gemma 4 26B A4B is deliberately not added to the "have been tested with OnionAI" list in onion-ai.md. It serves correctly and emits valid tool calls, but the reference box is a bare AMD dev platform with no grid pointed at it, so it hasn't been exercised against SOC itself. Happy to add it once someone runs it against a real grid.

mkdocs build passes with no new warnings.

🤖 Generated with Claude Code

https://claude.ai/code/session_011MvtwjmU5wWs4rkPwKifNB

Adds a tested walkthrough for serving a local OpenAI-compatible endpoint
for Onion AI on an AMD Strix Halo machine, using the packaged
GPU-accelerated llama.cpp rather than a hand-built stack.

The Hosting Local Models section of onion-ai.md previously offered only
"use LM Studio" and "you need at least 96GB of VRAM" with no procedure.
That VRAM figure is accurate for the large models listed above it but
misleading in general, so it is now scoped to those models and points at
the new page.

Every command and figure in the page was executed on the reference
hardware. Notable findings baked into the guide:

- The llama.cpp package ships its own llama-server systemd unit and
  /etc/default/llama-server config, so no custom unit is needed.
- The packaged _llama-server account cannot reach the GPU without being
  added to the render and video groups. Interactive shells work anyway
  because logind grants a per-user ACL on /dev/kfd and /dev/dri, so the
  failure is easy to miss: the service starts, looks healthy, and
  silently falls back to unusably slow CPU inference.
- Gemma 4 is a reasoning model and returns reasoning_content, leaving
  content empty until it finishes thinking. Too low a token limit yields
  an empty response, which reads as a broken endpoint.
- Q8_0 with the full 262144 token window uses only ~32 GB, and a fact at
  the start of a 193,139 token prompt was retrieved correctly.

vLLM was evaluated and rejected: gfx1151 is not in its supported GPU
list, there is no ROCm wheel, and every working reference builds from
source with patches inside a container. That is not a procedure worth
publishing.

Gemma 4 26B A4B is deliberately not added to the tested-models list in
onion-ai.md. It serves correctly and emits well-formed tool calls, but
it has not been exercised against an actual grid.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011MvtwjmU5wWs4rkPwKifNB
@dougburks
dougburks merged commit cafb46b into 3/dev Sep 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants