Open-weight LLMs · on-prem sizing

On-Prem LLM Sizing Matrix

Where the current open-weight models fit and roughly how fast they generate, from a 16 GB gaming card to an 8-GPU Blackwell node. Memory capacity decides whether a model fits. Memory bandwidth decides how fast it writes. Mixture-of-experts models split the two: total parameters set the memory, active parameters set the speed.

October 2026Estimates, not benchmarks
Expect ±30–50%

Precision
Context per user
Concurrent users
<55–1515–4040–100100+tokens/stight fit (<10% headroom)runs with RAM offloadonly at ~2.5-bit (quality loss)doesn't fit2 nodesneeds that many nodes
Model
active of total · needs
Desktop GPUWorkstation & unified memoryDatacenter
RTX 5060 Ti16 GB448 GB/s~$450
RTX 3090 / 409024 GB1.01 TB/s$0.8–2k
RTX 509032 GB1.79 TB/s$2.5–4k
2× RTX 3090/409048 GB2 × 1.01 TB/s$1.6–4k
DGX Spark / Strix Halo128 GB273 GB/s$2.5–4k
RTX PRO 600096 GB1.79 TB/s~$8–9k
Mac Studio M3 Ultra512 GB819 GB/s~$10k
4× RTX PRO 6000384 GB4 × 1.79 TB/s~$45k
1× H100 80GB80 GB3.35 TB/s$25–31k
1× H200141 GB4.8 TB/s$30–40k
8× H100 node640 GB8 × 3.35 TB/s$250–320k
8× H200 node1.1 TB8 × 4.8 TB/s$320–420k
8× B200 node1.4 TB8 × 8 TB/s$400–500k
8× B300 node2.3 TB8 × 8 TB/son quote
1T+ parameters
Kimi K3104B of 2.8T · MoEMXFP4 1.6 TB
Qwen3.8-2.4T-A95B95B of 2.4T · MoEFP8 2.6 TB · Q4 1.5 TB
DeepSeek V4 Pro49B of 1.6T · MoEFP4/FP8 923 GB
Kimi K2.632B of 1.1T · MoEINT4 615 GB
250–800B parameters
GLM-5.240B of 744B · MoEFP8 783 GB · Q4 454 GB
DeepSeek V4.1 Flash16B of 749B · MoEFP4/FP8 527 GB · Q4 455 GB
Mistral Large 341B of 675B · MoEFP8 716 GB · Q4 412 GB
Nemotron 3 Ultra55B of 550B · MoEFP8 582 GB · Q4 334 GB
MiniMax M323B of 428B · MoEMXFP8 455 GB · Q4 262 GB
Qwen3.5-397B-A17B17B of 397B · MoEFP8 436 GB · Q4 242 GB
DeepSeek V4 Flash13B of 284B · MoEFP4/FP8 166 GB
100–130B parameters
Mistral Small 46.5B of 119B · MoEFP8 133 GB · Q4 74 GB
Nemotron 3 Super12B of 120B · MoEFP8 133 GB · Q4 74 GB
gpt-oss-120b5.1B of 117B · MoEMXFP4 69 GB
Llama 4 Scout17B of 109B · MoEFP8 122 GB · Q4 69 GB
Up to 35B parameters
Qwen3.6-35B-A3B3B of 35B · MoEFP8 40 GB · Q4 23 GB
Gemma 4 31B31B of 31B · denseFP8 38 GB · Q4 23 GB
Qwen3.8-27B27B of 27B · denseFP8 33 GB · Q4 19 GB
Gemma 4 26B-A4B4B of 26B · MoEFP8 31 GB · Q4 19 GB
gpt-oss-20b3.6B of 21B · MoEMXFP4 15 GB
Gemma 4 12B12B of 12B · denseFP8 16 GB · Q4 9.8 GB

Cells show estimated decode speed for one user in tokens/s, with the precision used below. Context: 32K tokens. Click any cell for the breakdown.

gpt-oss-120b on DGX Spark / Strix Halo

Fits in memory at MXFP4.

47tok/s, one user
bandwidth-bound decode
Weights 65 GB + KV cache 1.1 GB (32K × 1) + runtime 2.9 GB = 69 GB · usable 110 GB

OpenAI · 117B total, 5.1B active · ships MXFP4Apache 2.0. Native MXFP4, about 65 GB: the classic model for 128 GB boxes and single 80–96 GB cards.

DGX Spark / Strix Halo · 128 GB · $2.5–4k128 GB unified memory at 256–273 GB/s: good for MoE models, very slow for large dense ones.

Memory footprint against hardware capacity

Each bar runs from the 4-bit footprint to the serving-precision footprint (up to 8-bit), including KV cache for 32K context × 1 user. The ring marks ~2.5-bit. Dashed lines are usable memory per hardware tier. Log scale.

4 GB8 GB16 GB32 GB64 GB128 GB256 GB512 GB1 TB2 TB4 TB5060 Ti409050902× 4090H100PRO 6000SparkH2004× PROMac 5128× H1008× H2008× B2008× B300Kimi K31.6 TBQwen3.8-2.4T-A95B2.6 TBDeepSeek V4 Pro923 GBKimi K2.6615 GBGLM-5.2783 GBDeepSeek V4.1 Flash527 GBMistral Large 3716 GBNemotron 3 Ultra582 GBMiniMax M3455 GBQwen3.5-397B-A17B436 GBDeepSeek V4 Flash166 GBMistral Small 4133 GBNemotron 3 Super133 GBgpt-oss-120b69 GBLlama 4 Scout122 GBQwen3.6-35B-A3B40 GBGemma 4 31B38 GBQwen3.8-27B33 GBGemma 4 26B-A4B31 GBgpt-oss-20b15 GBGemma 4 12B16 GB

Calibration against measured numbers

The same formulas, run on configurations with published single-machine measurements. Estimates landing within roughly ±30% of measurements is the accuracy to expect across the matrix.

ConfigurationHardwareMeasuredThis modelAgreementSource
Qwen3 32B · Q4_K_MRTX 509061 tok/s57 tok/s−7%Hardware Corner (llama.cpp)
Qwen2.5 32B · Q4_K_MRTX 3090 / 409042 tok/s34 tok/s−19%Kunal Ganglani (llama.cpp)
Llama 3.1 8B · Q4_K_MRTX 5090220 tok/s190 tok/s−13%Kunal Ganglani (llama.cpp)
gpt-oss-120b · MXFP4DGX Spark / Strix Halo61 tok/s54 tok/s−12%llama.cpp maintainer, via LocalAIMaster
Llama 3.1 70B · FP8 (dense)DGX Spark / Strix Halo2.7 tok/s2.7 tok/s−0%LMSYS, via LocalAIMaster
Llama 3.1 70B · FP8 · 64 users1× H100 80GB460–984 tok/s696 tok/s agg.within rangevLLM: Morph (460) and Prem AI (984)

How the estimates work

Memory needed

total params × bits ÷ 8 for weights, plus KV cache (per-model estimate × context × users), plus 3% runtime overhead and 1 GB per GPU. Usable memory is 93% of GPU memory, 110 GB on 128 GB unified boxes, and 470 GB on a 512 GB Mac with the wired-memory limit raised.

Decode speed

Each step reads the active weights once (for MoE with several users, the union of experts they touch) plus the KV cache. Step time is bytes ÷ effective bandwidth or FLOPs ÷ effective compute, whichever is larger, plus a per-layer overhead. Effective bandwidth is 75% of peak (70% on Apple).

Why 8 GPUs don't make one user 8× faster

Every layer pays a fixed cost for kernel launches, MoE routing and, across GPUs, an all-reduce: about 0.03 ms (dense) or 0.1 ms (MoE) per layer, plus 0.06–0.15 ms on multi-GPU. On a node this floor dominates, so a single user sees tens of tokens per second. The node earns its price on concurrency and on capacity.

Multi-GPU and offload

Tensor-parallel scaling is taken as 60% on two PCIe GPUs, 50% on four, and 75% on an NVLink node. With offload on, weights that don't fit in VRAM sit in system RAM read at about 60 GB/s; this is how MoE models run on a single desktop card.

Concurrency

Aggregate throughput is users ÷ step time, discounted for scheduling and prefill interference (about 2% per extra user). KV cache memory scales with users, so long contexts for many users run out of memory first.

What it ignores

Prompt processing time (time to first token), KV-cache quantization (FP8 KV halves that memory), quality loss from quantization, and runtime-specific kernels. Best fit never goes above 8-bit because FP8 serving is close to lossless. Layer counts and KV sizes for several 2026 models are estimated from their architecture families; Mistral Large 3 and Nemotron 3 Ultra are assumed to ship FP8.

Sources

Compiled 10 October 2026. Open-weight model releases move monthly; treat every figure as a starting point and benchmark on your own stack before buying.

← Back to the blog

Privacy and analytics

Only with your consent, PostHog (EU) and Cloudflare measure pages visited and their order, visit duration, referral, approximate country, browser, device, and performance.

More details

We collect the paths of pages visited and their order, visit times and duration, the referring site, approximate country (derived from the IP address), and basic browser and device information. Cloudflare also measures page performance.

We do not record form contents, individual clicks, or session videos. PostHog does not create a personal profile and keeps an anonymous identifier only for this browser tab. We save your choice in the browser for future visits.

If you disagree, you can leave the site without activating this analytics.

Disagree and leave site