🤖 AI Tools
llama.cpp GGUF Quantization and llama-server Launch Planner: Quant Choice by VRAM, Context Length and KV Cache Budget, GPU Layer Offload, Parallel Slots, Chat Template, and OpenAI Compatible Endpoint Tests
Plan a local model server on your own GPU with llama.cpp: choose a GGUF quantization that fits your VRAM, budget the KV cache for the context length and number of parallel users, set GPU offload, flash attention, and cache type flags, load the right chat template, and test the OpenAI compatible endpoint before pointing apps at it.
0Reviews
Prompt
Act as a local inference engineer who deploys llama.cpp llama-server on workstations and small servers for teams, and who sizes every launch from the model architecture and the GPU's free memory instead of guessing until it stops crashing. Inputs: - The model and its architecture details from the model card or GGUF metadata (parameters, layers, attention heads, KV heads, head dimension, trained context): [ModelArch] - GPU model, total VRAM, and what else uses it (desktop, other apps); CPU and system RAM: [Hardware] - Context length needed per request and how many concurrent users or requests: [Workload] - Quality priority (coding, extraction, chat) and acceptable speed: [QualityNeeds] - llama.cpp build or release tag and how it was installed: [BuildVersion] - Who connects and from where (localhost only, LAN, an app using the OpenAI SDK): [ClientSetup] - Output format: [Format] Generate: 1. A quantization shortlist (for example Q8_0, Q6_K, Q5_K_M, Q4_K_M, IQ4_XS) with approximate file size for this parameter count, the expected quality tradeoff in plain words, and which one fits QualityNeeds. 2. KV cache math: bytes per token = 2 x layers x KV heads x head dimension x bytes per element, then total for the context in Workload, for f16 and for q8_0 cache types. Use the KV head count, not the attention head count, for grouped query attention models. 3. A VRAM budget table: weights plus KV cache plus a compute buffer allowance against free VRAM, with the recommended quant and context, and what to change first if it does not fit (cache type, context, quant, then partial offload). 4. Parallel slots: how the total context is divided across --parallel slots in most builds, so per user context is clear, and when to raise context instead. 5. The llama-server command with each flag explained: model path, --host and --port, context size, --parallel, GPU layers, flash attention, cache types, --jinja for the model's chat template, an API key, and threads for any CPU part. 6. Endpoint tests with curl: /health, /v1/models, and a /v1/chat/completions request, plus the base_url and key settings for an OpenAI SDK client. 7. Tuning notes: how to read the startup log for offloaded layers and buffer sizes, and the symptom of each common mistake (out of memory at load, gibberish from a wrong template, slow speed from partial offload). Constraints: - Flag names change between releases; tell the user to confirm against llama-server --help for their BuildVersion. - Do not publish benchmark speeds you cannot know; give sizes and math, and label estimates. - Never suggest exposing the server on a public interface without an API key and a firewall. No em dashes.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
ModelArch: Llama 3.1 8B Instruct GGUF; 32 layers, 32 attention heads, 8 KV heads, head dimension 128, trained context 128K Hardware: RTX 4070 12 GB, Windows desktop uses about 1 GB; Ryzen 7, 32 GB RAM Workload: 4 teammates at once, each up to 8K tokens of context for document Q and A QualityNeeds: accurate extraction from contracts, speed secondary BuildVersion: recent prebuilt release with CUDA ClientSetup: LAN, an internal Python app using the OpenAI SDK Format: tables, command, tests
Example Output
1. Quant shortlist (8B, approximate file sizes)
| Quant | Size | Notes |
| Q8_0 | about 8.5 GB | Closest to full quality, tight with 32K cache |
| Q6_K | about 6.6 GB | Very close to Q8_0 for extraction |
| Q5_K_M | about 5.7 GB | Good balance, most headroom for cache |
| Q4_K_M | about 4.9 GB | Fine for chat, more small errors on exact extraction |
Pick Q6_K for contract extraction; fall back to Q5_K_M if VRAM is tight.
2. KV cache
Per token, f16: 2 x 32 layers x 8 KV heads x 128 x 2 bytes = 131,072 bytes = 128 KiB
4 users x 8K = 32,768 tokens
f16: 32,768 x 128 KiB = 4.0 GiB
q8_0: roughly 53 percent of f16, about 2.1 GiB
3. VRAM budget (12 GB card, about 11 GB free)
| Plan | Weights | KV | Buffers (est.) | Total | Fits |
| Q6_K + f16 KV | 6.6 | 4.3 | 0.8 | 11.7 GB | No |
| Q6_K + q8_0 KV | 6.6 | 2.3 | 0.8 | 9.7 GB | Yes |
| Q5_K_M + q8_0 KV | 5.7 | 2.3 | 0.8 | 8.8 GB | Yes, more headroom |
Recommended: Q6_K with q8_0 K and V cache, all layers on GPU.
4. Parallel slots
With -c 32768 and --parallel 4, most builds give each slot 8,192 tokens. If one teammate needs 16K, raise -c to 65536 (KV doubles to about 4.3 GiB at q8_0) or lower --parallel.
5. Command
llama-server ^
-m D:\models\Meta-Llama-3.1-8B-Instruct-Q6_K.gguf ^
--host 0.0.0.0 --port 8080 ^
-c 32768 --parallel 4 ^
-ngl 99 ^
--flash-attn on ^
--cache-type-k q8_0 --cache-type-v q8_0 ^
--jinja ^
--api-key %LLAMA_API_KEY%
- -ngl 99 offloads every layer; the log should report all layers offloaded to GPU.
- A quantized V cache needs flash attention; older builds take a bare -fa instead of --flash-attn on.
- --jinja uses the chat template stored in the GGUF, which avoids wrong formatting.
- 0.0.0.0 listens on the LAN; allow port 8080 only from the office subnet in Windows Firewall.
6. Tests
curl http://192.168.1.50:8080/health
curl -H "Authorization: Bearer %LLAMA_API_KEY%" http://192.168.1.50:8080/v1/models
curl http://192.168.1.50:8080/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer %LLAMA_API_KEY%" -d "{\"messages\":[{\"role\":\"user\",\"content\":\"Reply with OK\"}]}"
Python: OpenAI(base_url="http://192.168.1.50:8080/v1", api_key=os.environ["LLAMA_API_KEY"]); the model name can be any string.
7. Tuning
- Out of memory at load: switch to Q5_K_M before cutting context.
- Rambling or role tags in replies: template problem; confirm --jinja.
- Slow tokens and low GPU use: some layers stayed on CPU; check the offload line.