How to Use the llama.cpp GGUF and llama-server Planner Prompt to Fit a Local Model on Your GPU
Plan a local model server on your own GPU with llama.cpp: choose a GGUF quantization that fits your VRAM, budget the KV cache for the context length and number of parallel users, set GPU offload, flash attention, and cache type flags, load the right chat template, and test the OpenAI compatible endpoint before pointing apps at it.

Running a local model with llama.cpp is easy until the server runs out of memory at load, replies with odd formatting, or slows down because half the model ended up on the CPU. Most of these problems come from guessing the quantization, context length, and flags instead of doing the math. The llama.cpp GGUF Quantization and llama-server Launch Planner: Quant Choice by VRAM, Context Length and KV Cache Budget, GPU Layer Offload, Parallel Slots, Chat Template, and OpenAI Compatible Endpoint Tests prompt sizes your launch from the model architecture and your GPU, writes the llama-server command, and gives you tests for the OpenAI compatible endpoint.
What the prompt produces
- A quantization shortlist with approximate file sizes and quality tradeoffs in plain words.
- KV cache math for your context length, for both f16 and q8_0 cache types.
- A VRAM budget table that adds weights, cache, and buffers against your free memory.
- A parallel slots explanation so you know how much context each user gets.
- The llama-server command with every flag explained.
- Endpoint tests with curl and settings for an OpenAI SDK client.
- Tuning notes that match common symptoms to their causes.
How to fill the inputs
ModelArch lists the model and its architecture details: parameters, layers, attention heads, KV heads, head dimension, and trained context. These come from the model card or GGUF metadata. The KV head count matters most for cache size.
Hardware covers the GPU, its total VRAM, what else uses it, and your CPU and system RAM.
Workload sets the context needed per request and how many users or requests run at once.
QualityNeeds describes the task, such as extraction, coding, or chat, and how much speed matters.
BuildVersion is the llama.cpp release you installed. Flag names change between releases.
ClientSetup says who connects and from where, such as a local app or a team on the LAN.
Reading the example output
The example plans a server for an 8B instruct model on a 12 GB GPU for four teammates doing document questions:
- The shortlist compares four quantizations and picks a higher quality one for contract extraction, with a smaller fallback.
- The KV cache math is shown in full. It uses the eight KV heads, not the thirty two attention heads, which makes the cache much smaller.
- The budget table shows which plans fit. The f16 cache does not fit with the chosen quantization, but a q8_0 cache does.
- Parallel slots are explained. Four slots with a 32K context give each user about 8K tokens.
- The command explains each flag, including offloading all layers, enabling flash attention for the quantized cache, and using the GGUF's chat template.
- The tests check health, models, and a chat request, and show the base URL for the Python client.
Tips for better results
- Check free VRAM with nothing else running, then subtract what your desktop needs.
- Read the startup log to confirm every layer is offloaded.
- Change the cache type before dropping to a smaller quantization.
- Keep the API key in an environment variable, not in scripts.
- Confirm every flag against llama-server help for your build.
Mistakes to avoid
- Do not use the attention head count in KV cache math for grouped query attention models.
- Do not expose the server to the internet without an API key and a firewall.
- Do not skip the chat template. A wrong template produces rambling output.
- Do not trust speed numbers from other machines. Measure on yours.
Who it is for
Developers running private models for a team, IT staff setting up internal AI tools, hobbyists with a gaming GPU, privacy focused businesses, and anyone moving from a hosted API to a local server.
Related PromptDig links
Open the llama.cpp GGUF Quantization and llama-server Launch Planner: Quant Choice by VRAM, Context Length and KV Cache Budget, GPU Layer Offload, Parallel Slots, Chat Template, and OpenAI Compatible Endpoint Tests prompt and enter your model and GPU. For more AI tool prompts, Browse more prompts. If you have a local AI prompt that works, Share a prompt.