🤖 AI Tools
vLLM OpenAI Compatible Server Deployment Planner: vllm serve Flags, max-model-len and gpu-memory-utilization, KV Cache Memory Math, Tensor Parallel Size, Prefix Caching, Quantized Checkpoints, Docker Run, API Key, and a Load Test Plan
Plan a self hosted vLLM endpoint that fits the GPU you actually have: model weight and KV cache memory math, the vllm serve command with each flag justified, tensor parallel choice, when a quantized checkpoint is worth it, a Docker run line, smoke tests against the OpenAI compatible API, a load test, and the metrics to watch once real traffic arrives.
0Reviews
Prompt
Act as an ML infrastructure engineer who deploys open weight models with vLLM in production, sizes every deployment with memory math before touching a GPU, and treats the startup log as the source of truth. Inputs: - Model repo id, parameter count, and architecture details if known (layers, KV heads, head dim): [ModelChoice] - GPUs: type, count, memory each, and whether they share a node: [GPUInventory] - Workload: typical prompt and output lengths, longest context needed, expected concurrent users, latency target: [WorkloadProfile] - Serving setup: bare metal, Docker, or Kubernetes, and how clients will call it: [ServingSetup] - Constraints such as license gating, data residency, or a required quantization: [ModelConstraints] - Output format: [Format] Generate: 1. Weight memory: parameters times bytes per parameter for the chosen dtype (about 2 bytes for bf16, about 0.5 to 0.6 for 4 bit AWQ or GPTQ checkpoints). 2. KV cache math: bytes per token = 2 x layers x KV heads x head dim x bytes per element; then tokens that fit in the budget left after weights and overhead at the chosen gpu-memory-utilization; then how many full length sequences that is at max-model-len. 3. A decision on tensor-parallel-size from GPUInventory, and whether a quantized checkpoint or fp8 KV cache is needed to meet WorkloadProfile. 4. The vllm serve command with each flag explained: model, served-model-name, max-model-len, gpu-memory-utilization, tensor-parallel-size, max-num-seqs, enable-prefix-caching if prompts share long prefixes, api-key, host, and port. 5. A Docker run line for the vllm/vllm-openai image with GPU access, ipc=host, the Hugging Face cache mounted, and HF_TOKEN for gated models; tell the user to pin an image tag. 6. Smoke tests with curl: /health, /v1/models, one /v1/chat/completions call, and one streaming call. 7. A load test plan: request mix from WorkloadProfile, the benchmark tool shipped with the installed vLLM version, and the numbers to record (time to first token, output tokens per second, failures). 8. What to read in the startup log (KV cache size and maximum concurrency lines) and in /metrics (running and waiting requests, KV cache usage) and what each pattern means. 9. Mark any flag name or metric that changed across vLLM versions as "check vllm serve --help for your version". Constraints: - Show the arithmetic. Commands copy and paste ready. No em dashes.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
ModelChoice: meta-llama/Llama-3.1-8B-Instruct, 8B params, 32 layers, 8 KV heads, head dim 128 GPUInventory: one NVIDIA L4, 24 GB WorkloadProfile: internal support assistant, prompts about 1,500 tokens including a shared 1,000 token system prompt, outputs about 300 tokens, longest context 8k, up to 10 concurrent users ServingSetup: Docker on a single cloud VM, clients use the OpenAI Python SDK ModelConstraints: gated model, must stay in our VPC Format: math, command, docker line, tests, monitoring
Example Output
1) WEIGHTS
8.0B params x 2 bytes (bf16) = about 16 GB.
2) KV CACHE
Per token: 2 x 32 layers x 8 KV heads x 128 x 2 bytes = 131,072 bytes = 128 KiB.
Budget at 0.90 utilization: 24 x 0.90 = 21.6 GB. Minus 16 GB weights and about 1.5 GB for activations and CUDA graphs leaves about 4 GB.
4 GiB / 128 KiB = about 32,000 tokens of KV cache.
At max-model-len 8192 that is 4 full length sequences at once. Real requests average about 1,800 tokens, so roughly 15 to 17 typical requests fit, which covers 10 users.
3) DECISIONS
- tensor-parallel-size 1 (single GPU).
- bf16 is enough for 10 users. If you need 16k context or 30 users, switch to an AWQ 4 bit checkpoint (about 5.5 GB of weights, roughly 3x the KV room) and test answer quality first.
- Prefix caching on: every request shares the same 1,000 token system prompt.
4) COMMAND
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--served-model-name support-8b \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--max-num-seqs 16 \
--enable-prefix-caching \
--api-key "$VLLM_API_KEY" \
--host 0.0.0.0 --port 8000
5) DOCKER
docker run -d --name vllm --gpus all --ipc=host -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e HF_TOKEN="$HF_TOKEN" -e VLLM_API_KEY="$VLLM_API_KEY" \
vllm/vllm-openai:<pinned-tag> \
--model meta-llama/Llama-3.1-8B-Instruct --served-model-name support-8b \
--max-model-len 8192 --gpu-memory-utilization 0.90 --max-num-seqs 16 \
--enable-prefix-caching --api-key "$VLLM_API_KEY"
Accept the Llama license on Hugging Face first. Expose port 8000 only inside the VPC.
6) SMOKE TESTS
curl -s localhost:8000/health
curl -s localhost:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl -s localhost:8000/v1/chat/completions -H "Authorization: Bearer $VLLM_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"support-8b","messages":[{"role":"user","content":"Reset steps for a locked account?"}],"max_tokens":200}'
Repeat with "stream": true and confirm chunks arrive.
SDK clients: base_url="http://<vm-ip>:8000/v1", api_key=VLLM_API_KEY, model="support-8b".
7) LOAD TEST
Use the serving benchmark that ships with your vLLM version: 1,500 token prompts, 300 token outputs, 10 then 20 concurrent. Record time to first token p50 and p95, output tokens per second, and errors.
8) WATCH
- Startup log: KV cache size in tokens and the maximum concurrency line. If it is far below 32,000 tokens, lower max-model-len or quantize.
- /metrics: waiting requests above 0 for minutes means you are out of KV room; KV cache usage near 100% with waiting requests confirms it.
9) CHECK
Benchmark command and metric names vary by release: check vllm serve --help and the docs for your version.