What it takes to run the biggest open models

The weights are free to download. The hardware is not. Kimi-K2-Instruct-0905 has 1026B parameters and needs about 632 GB of memory even squeezed down to Q4 — roughly 9 H100 cards, or $243,000 of silicon before you have powered any of it on.

Nobody is running these at a desk — that isn't the point. It's worth seeing the scale of what “open” actually means at the top end, and how far it sits from the models on the rest of this section. What your own computer can run →

For scale: the largest computer you can simply order — a Mac Studio M3 Ultra, 512GB at about $9,500 — holds 512GB. That is still 1.2× too small for Kimi-K2-Instruct-0905.

The models, by size

ModelParametersMemory @ Q4H100sRough cost
Kimi-K2-Instruct-0905Moonshot AI1026B632 GB9$243,000
MiMo-V2.5-ProXiaomi1023B635 GB9$243,000
DeepSeek-V3.1DeepSeek685B416 GB6$162,000
RakutenAI-3.0Rakuten671B415 GB6$162,000
ERNIE-4.5-VL-424B-A47B-PTBaidu AI · 47B active424B260 GB4$108,000
Qwen3.5-397B-A17BAlibaba / Qwen · 17B active403B244 GB4$108,000
Ornith-1.0-397BDeepreinforce · 79B active397B246 GB4$108,000
ERNIE-4.5-300B-A47B-PTBaidu AI · 47B active300B184 GB3$81,000
MiniMax-M2.7MiniMax · 46B active229B144 GB2$54,000
MiniMax-M2.5MiniMax · 46B active229B143 GB2$54,000

Memory is at Q4_K_M with 4K context — the smallest anyone sensibly runs these at; at full precision they are four times larger again. Card counts divide memory by the H100's 80GB with a 10% allowance, so they are a floor: real deployments lose more to activations, long-context caches and sharding. Prices are indicative street prices for the silicon alone, excluding the server, networking, power and cooling around it.

Kimi-K2-Instruct-0905 on different hardware

The same model, and what it would take on each generation of accelerator.

AcceleratorMemory eachCards neededRough costEst. speed
A100 80GBNVIDIA published specification: 80 GB HBM2e, 2,039 GB/s80 GB9$135,000~2.1 tok/s
H100 80GBNVIDIA published specification: 80 GB HBM3, 3.35 TB/s80 GB9$243,000~3.5 tok/s
H200 141GBNVIDIA published specification: 141 GB HBM3e, 4.8 TB/s141 GB5$160,000~5.0 tok/s
B200 192GBNVIDIA published specification: 192 GB HBM3e, 8 TB/s192 GB4$160,000~8.4 tok/s
Mac Studio M3 Ultra, 512GBApple's largest unified-memory configuration: 512 GB at 819 GB/s512 GBDoes not fit$9,500

Speeds assume a single card's bandwidth and are optimistic: splitting a model across several cards adds communication cost that a single-card figure ignores.

What about GPT, Claude and Gemini?

We can't tell you. OpenAI, Anthropic and Google don't publish parameter counts or release weights for their frontier models, so any figure for them would be a guess dressed up as a fact. Every number on this page comes from a published file size or an official specification, and we'd rather leave a gap than fill it with invention.

That difference is itself the point of the page. With an open-weight model you can measure exactly what it costs to run, because the weights are right there. With a closed one you get a price per token and no way to check what sits behind it.

Verified against HuggingFace on 2026-08-02What your computer can run →Why bandwidth sets the speed →1.07 GB per GiB