Repository navigation
add Local LLM Hosting guide - #38
Merged
Merged
Conversation
Adds a tested walkthrough for serving a local OpenAI-compatible endpoint for Onion AI on an AMD Strix Halo machine, using the packaged GPU-accelerated llama.cpp rather than a hand-built stack. The Hosting Local Models section of onion-ai.md previously offered only "use LM Studio" and "you need at least 96GB of VRAM" with no procedure. That VRAM figure is accurate for the large models listed above it but misleading in general, so it is now scoped to those models and points at the new page. Every command and figure in the page was executed on the reference hardware. Notable findings baked into the guide: - The llama.cpp package ships its own llama-server systemd unit and /etc/default/llama-server config, so no custom unit is needed. - The packaged _llama-server account cannot reach the GPU without being added to the render and video groups. Interactive shells work anyway because logind grants a per-user ACL on /dev/kfd and /dev/dri, so the failure is easy to miss: the service starts, looks healthy, and silently falls back to unusably slow CPU inference. - Gemma 4 is a reasoning model and returns reasoning_content, leaving content empty until it finishes thinking. Too low a token limit yields an empty response, which reads as a broken endpoint. - Q8_0 with the full 262144 token window uses only ~32 GB, and a fact at the start of a 193,139 token prompt was retrieved correctly. vLLM was evaluated and rejected: gfx1151 is not in its supported GPU list, there is no ROCm wheel, and every working reference builds from source with patches inside a container. That is not a procedure worth publishing. Gemma 4 26B A4B is deliberately not added to the tested-models list in onion-ai.md. It serves correctly and emits well-formed tool calls, but it has not been exercised against an actual grid. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011MvtwjmU5wWs4rkPwKifNB
dougburks
approved these changes
Sep 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds
docs/local-llm.md, a tested walkthrough for standing up a local OpenAI-compatible endpoint for Onion AI on an AMD Strix Halo machine, plus a link to it fromonion-ai.mdand a nav entry under Security Onion Pro.Why
The Hosting Local Models section of
onion-ai.mdoffered only "use LM Studio" and "You need at least 96GB of VRAM" with no actual procedure. That VRAM figure is right for the large models listed above it but misleading as general guidance, so it is now scoped to those models and points at the new page.Validation
This was built and run end to end on an AMD Strix Halo box (Ryzen AI Developer Platform, gfx1151, ROCm 7.13, 128GB unified memory). Every command and number in the page came from that machine — nothing is inferred.
Findings worth calling out, all of which are now documented:
llama.cpppackage ships its ownllama-serversystemd unit and/etc/default/llama-serverconfig, so no custom unit file is needed._llama-serveraccount cannot reach the GPU without being added to therenderandvideogroups. Interactive shells work regardless, becauselogindgrants a per-user ACL on/dev/kfdand/dev/dri/renderD128. The failure mode is nasty: the service starts, reports healthy, and silently falls back to unusably slow CPU inference. This gets a prominent warning.reasoning_contentand leavescontentempty until it finishes thinking, so too low a token limit produces an empty response that reads as a broken endpoint.tool_calls, which matters because the assistant depends on it.Q8_0at the full 262,144 token window uses only ~32 GB, and a fact planted at the start of a 193,139 token prompt was retrieved correctly at the end — the window is usable, not merely allocatable.Measured: ~41 tok/s generation on short context, ~22 tok/s at ~193k, bulk prefill ~850 tok/s falling to ~228 tok/s averaged across a 193k prompt.
vLLM
vLLM was evaluated first and rejected. gfx1151 is not in its supported GPU list, there is no ROCm wheel, and every working reference builds from source with patches (stub
amdsmi, force the arch intoCMakeLists.txt, disable AITER MoE) inside a container. That patch set rots as vLLM moves and isn't something worth publishing in customer-facing docs. The packaged llama.cpp path is vendor-supported and survives upgrades.Note for reviewers
Gemma 4 26B A4B is deliberately not added to the "have been tested with OnionAI" list in
onion-ai.md. It serves correctly and emits valid tool calls, but the reference box is a bare AMD dev platform with no grid pointed at it, so it hasn't been exercised against SOC itself. Happy to add it once someone runs it against a real grid.mkdocs buildpasses with no new warnings.🤖 Generated with Claude Code
https://claude.ai/code/session_011MvtwjmU5wWs4rkPwKifNB