Open-weight LLMs · on-prem sizing
On-Prem LLM Sizing Matrix
Where the current open-weight models fit and roughly how fast they generate, from a 16 GB gaming card to an 8-GPU Blackwell node. Memory capacity decides whether a model fits. Memory bandwidth decides how fast it writes. Mixture-of-experts models split the two: total parameters set the memory, active parameters set the speed.
October 2026Estimates, not benchmarks
Expect ±30–50%
| Model active of total · needs | Desktop GPU | Workstation & unified memory | Datacenter | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
RTX 5060 Ti16 GB448 GB/s~$450 | RTX 3090 / 409024 GB1.01 TB/s$0.8–2k | RTX 509032 GB1.79 TB/s$2.5–4k | 2× RTX 3090/409048 GB2 × 1.01 TB/s$1.6–4k | DGX Spark / Strix Halo128 GB273 GB/s$2.5–4k | RTX PRO 600096 GB1.79 TB/s~$8–9k | Mac Studio M3 Ultra512 GB819 GB/s~$10k | 4× RTX PRO 6000384 GB4 × 1.79 TB/s~$45k | 1× H100 80GB80 GB3.35 TB/s$25–31k | 1× H200141 GB4.8 TB/s$30–40k | 8× H100 node640 GB8 × 3.35 TB/s$250–320k | 8× H200 node1.1 TB8 × 4.8 TB/s$320–420k | 8× B200 node1.4 TB8 × 8 TB/s$400–500k | 8× B300 node2.3 TB8 × 8 TB/son quote | |
| 1T+ parameters | ||||||||||||||
| Kimi K3104B of 2.8T · MoEMXFP4 1.6 TB | ||||||||||||||
| Qwen3.8-2.4T-A95B95B of 2.4T · MoEFP8 2.6 TB · Q4 1.5 TB | ||||||||||||||
| DeepSeek V4 Pro49B of 1.6T · MoEFP4/FP8 923 GB | ||||||||||||||
| Kimi K2.632B of 1.1T · MoEINT4 615 GB | ||||||||||||||
| 250–800B parameters | ||||||||||||||
| GLM-5.240B of 744B · MoEFP8 783 GB · Q4 454 GB | ||||||||||||||
| DeepSeek V4.1 Flash16B of 749B · MoEFP4/FP8 527 GB · Q4 455 GB | ||||||||||||||
| Mistral Large 341B of 675B · MoEFP8 716 GB · Q4 412 GB | ||||||||||||||
| Nemotron 3 Ultra55B of 550B · MoEFP8 582 GB · Q4 334 GB | ||||||||||||||
| MiniMax M323B of 428B · MoEMXFP8 455 GB · Q4 262 GB | ||||||||||||||
| Qwen3.5-397B-A17B17B of 397B · MoEFP8 436 GB · Q4 242 GB | ||||||||||||||
| DeepSeek V4 Flash13B of 284B · MoEFP4/FP8 166 GB | ||||||||||||||
| 100–130B parameters | ||||||||||||||
| Mistral Small 46.5B of 119B · MoEFP8 133 GB · Q4 74 GB | ||||||||||||||
| Nemotron 3 Super12B of 120B · MoEFP8 133 GB · Q4 74 GB | ||||||||||||||
| gpt-oss-120b5.1B of 117B · MoEMXFP4 69 GB | ||||||||||||||
| Llama 4 Scout17B of 109B · MoEFP8 122 GB · Q4 69 GB | ||||||||||||||
| Up to 35B parameters | ||||||||||||||
| Qwen3.6-35B-A3B3B of 35B · MoEFP8 40 GB · Q4 23 GB | ||||||||||||||
| Gemma 4 31B31B of 31B · denseFP8 38 GB · Q4 23 GB | ||||||||||||||
| Qwen3.8-27B27B of 27B · denseFP8 33 GB · Q4 19 GB | ||||||||||||||
| Gemma 4 26B-A4B4B of 26B · MoEFP8 31 GB · Q4 19 GB | ||||||||||||||
| gpt-oss-20b3.6B of 21B · MoEMXFP4 15 GB | ||||||||||||||
| Gemma 4 12B12B of 12B · denseFP8 16 GB · Q4 9.8 GB | ||||||||||||||
Cells show estimated decode speed for one user in tokens/s, with the precision used below. Context: 32K tokens. Click any cell for the breakdown.
gpt-oss-120b on DGX Spark / Strix Halo
Fits in memory at MXFP4.
OpenAI · 117B total, 5.1B active · ships MXFP4Apache 2.0. Native MXFP4, about 65 GB: the classic model for 128 GB boxes and single 80–96 GB cards.
DGX Spark / Strix Halo · 128 GB · $2.5–4k128 GB unified memory at 256–273 GB/s: good for MoE models, very slow for large dense ones.
Memory footprint against hardware capacity
Each bar runs from the 4-bit footprint to the serving-precision footprint (up to 8-bit), including KV cache for 32K context × 1 user. The ring marks ~2.5-bit. Dashed lines are usable memory per hardware tier. Log scale.
Calibration against measured numbers
The same formulas, run on configurations with published single-machine measurements. Estimates landing within roughly ±30% of measurements is the accuracy to expect across the matrix.
| Configuration | Hardware | Measured | This model | Agreement | Source |
|---|---|---|---|---|---|
| Qwen3 32B · Q4_K_M | RTX 5090 | 61 tok/s | 57 tok/s | −7% | Hardware Corner (llama.cpp) |
| Qwen2.5 32B · Q4_K_M | RTX 3090 / 4090 | 42 tok/s | 34 tok/s | −19% | Kunal Ganglani (llama.cpp) |
| Llama 3.1 8B · Q4_K_M | RTX 5090 | 220 tok/s | 190 tok/s | −13% | Kunal Ganglani (llama.cpp) |
| gpt-oss-120b · MXFP4 | DGX Spark / Strix Halo | 61 tok/s | 54 tok/s | −12% | llama.cpp maintainer, via LocalAIMaster |
| Llama 3.1 70B · FP8 (dense) | DGX Spark / Strix Halo | 2.7 tok/s | 2.7 tok/s | −0% | LMSYS, via LocalAIMaster |
| Llama 3.1 70B · FP8 · 64 users | 1× H100 80GB | 460–984 tok/s | 696 tok/s agg. | within range | vLLM: Morph (460) and Prem AI (984) |
How the estimates work
Memory needed
total params × bits ÷ 8 for weights, plus KV cache (per-model estimate × context × users), plus 3% runtime overhead and 1 GB per GPU. Usable memory is 93% of GPU memory, 110 GB on 128 GB unified boxes, and 470 GB on a 512 GB Mac with the wired-memory limit raised.
Decode speed
Each step reads the active weights once (for MoE with several users, the union of experts they touch) plus the KV cache. Step time is bytes ÷ effective bandwidth or FLOPs ÷ effective compute, whichever is larger, plus a per-layer overhead. Effective bandwidth is 75% of peak (70% on Apple).
Why 8 GPUs don't make one user 8× faster
Every layer pays a fixed cost for kernel launches, MoE routing and, across GPUs, an all-reduce: about 0.03 ms (dense) or 0.1 ms (MoE) per layer, plus 0.06–0.15 ms on multi-GPU. On a node this floor dominates, so a single user sees tens of tokens per second. The node earns its price on concurrency and on capacity.
Multi-GPU and offload
Tensor-parallel scaling is taken as 60% on two PCIe GPUs, 50% on four, and 75% on an NVLink node. With offload on, weights that don't fit in VRAM sit in system RAM read at about 60 GB/s; this is how MoE models run on a single desktop card.
Concurrency
Aggregate throughput is users ÷ step time, discounted for scheduling and prefill interference (about 2% per extra user). KV cache memory scales with users, so long contexts for many users run out of memory first.
What it ignores
Prompt processing time (time to first token), KV-cache quantization (FP8 KV halves that memory), quality loss from quantization, and runtime-specific kernels. Best fit never goes above 8-bit because FP8 serving is close to lossless. Layer counts and KV sizes for several 2026 models are estimated from their architecture families; Mistral Large 3 and Nemotron 3 Ultra are assumed to ship FP8.
Sources
Model specs
- Kimi K3 (Vast.ai), K3 hardware (Yotta Labs)
- Qwen3.8-2.4T-A95B (vLLM recipes), Qwen3.8-27B (The Decoder)
- DeepSeek V4 Pro / Flash (NYU Shanghai RITS), V4 Pro GA weights (AI Weekly)
- DeepSeek V4.1 Flash (Yotta Labs), V4 Flash sizes (LocalAIMaster)
- Kimi K2.6 local guide, GLM-5.2 guide (Codersera)
- MiniMax M3 (Artificial Analysis), Nemotron 3 Ultra (Decrypt)
- Mistral Small 4 (SGLang), Mistral Large 3 (AI Weekly)
- Gemma 4 / Qwen3.6 (Unsloth)
Measured speeds
Prices
- H200 pricing (Akash)
- Data center GPU prices (IntuitionLabs)
- Spark vs Strix Halo vs Mac (LocalAIMaster)
- Consumer and workstation prices: rough street estimates, USD.
Compiled 10 October 2026. Open-weight model releases move monthly; treat every figure as a starting point and benchmark on your own stack before buying.
← Back to the blog