Back to Discover

#self hosted ai

1 prompt found

vLLM OpenAI Compatible Server Deployment Planner: vllm serve Flags, max-model-len and gpu-memory-utilization, KV Cache Memory Math, Tensor Parallel Size, Prefix Caching, Quantized Checkpoints, Docker Run, API Key, and a Load Test Plan
🤖 AI Tools

vLLM OpenAI Compatible Server Deployment Planner: vllm serve Flags, max-model-len and gpu-memory-utilization, KV Cache Memory Math, Tensor Parallel Size, Prefix Caching, Quantized Checkpoints, Docker Run, API Key, and a Load Test Plan

Ppromptstudio·Oct 8, 2026
No rating

Plan a self hosted vLLM endpoint that fits the GPU you actually have: model weight and KV cache memory math, the vllm serve command with each flag justified, tensor parallel choice, when a quantized checkpoint is worth it, a Docker run line, smoke tests against the OpenAI compatible API, a load test, and the metrics to watch once real traffic arrives.

Act as an ML infrastructure engineer who deploys open weight models with vLLM in production, sizes every deployment with memory math before touching a GPU, and treats the startup log as the source of truth. Inputs: - Model repo id, parameter count, and architecture details if known (layers, KV heads, head dim): [ModelChoice] - GPUs: type, count, memory each, and whether they share a node: [GPUInventory] - Workload: typical prompt and output lengths, longest context needed, expected concurrent users, latency target: [WorkloadProfile] - Serving setup: bare metal, Docker, or Kubernetes, and how clients will call it: [ServingSetup] - Constraints such as license gating, data residency, or a required quantization: [ModelConstraints] - Output format: [Format] Generate: 1. Weight memory: parameters times bytes per parameter for the chosen dtype (about 2 bytes for bf16, about 0.5 to 0.6 for 4 bit AWQ or GPTQ checkpoints). 2. KV cache math: bytes per token = 2 x layers x KV heads x head dim x bytes per element; then tokens that fit in the budget left after weights and overhead at the chosen gpu-memory-utilization; then how many full length sequences that is at max-model-len. 3. A decision on tensor-parallel-size from GPUInventory, and whether a quantized checkpoint or fp8 KV cache is needed to meet WorkloadProfile. 4. The vllm serve command with each flag explained: model, served-model-name, max-model-len, gpu-memory-utilization, tensor-parallel-size, max-num-seqs, enable-prefix-caching if prompts share long prefixes, api-key, host, and port. 5. A Docker run line for the vllm/vllm-openai image with GPU access, ipc=host, the Hugging Face cache mounted, and HF_TOKEN for gated models; tell the user to pin an image tag. 6. Smoke tests with curl: /health, /v1/models, one /v1/chat/completions call, and one streaming call. 7. A load test plan: request mix from WorkloadProfile, the benchmark tool shipped with the installed vLLM version, and the numbers to record (time to first token, output tokens per second, failures). 8. What to read in the startup log (KV cache size and maximum concurrency lines) and in /metrics (running and waiting requests, KV cache usage) and what each pattern means. 9. Mark any flag name or metric that changed across vLLM versions as "check vllm serve --help for your version". Constraints: - Show the arithmetic. Commands copy and paste ready. No em dashes.